← Back to explorer

Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Shows that reasoning distillation recipes like s1 and LIMO -- training a model on the same thousand or fewer worked solutions many times over -- leave reasoning fragile to later training stages, even stages with nothing to do with reasoning. Qwen3.5-9B-Base fine-tuned on its own correct competition-math solutions solved ~95% of held-out problems whether drilled on a few hundred solutions ~8 times each or shown many more once. But one pass of ordinary instruction tuning left the once-shown model intact while the drilled one fell to 86.0%, and harsher later stages took it to 59.3% or below. A third model that revisited the drilled problems equally often with a fresh solution each time was unharmed, so the damage comes from seeing the same texts again, not from few problems. The break recurs with a stronger model's traces, across further runs, models, and tasks. It is cheap to undo: reasoning is suppressed rather than erased -- five updates of reasoning training (or brief training on the reasoning format with almost no mathematics) bring almost all of it back. Fresh solutions and replaying 6.25% of original solutions in gentler later stages also prevent the damage; sharpening alone does not explain it.

Keywords

reasoning · post-training · instruction tuning · distillation · s1 · LIMO · Qwen

Topics

reasoning, post-training, fine-tuning, distillation

Research notes

  • Discovery: shared directly as an arXiv link (arxivb.org mirror): https://arxivb.org/abs/2609.33559
  • Single-author paper by Ely Sheikh, submitted to arXiv 2026-09-27, cs.LG, CC-BY 4.0.
  • Practical takeaway for post-training pipelines: avoid multi-epoch small-data reasoning distillation without replay; mitigations include fresh solutions per visit, ~6% solution replay in later stages, or a brief reasoning-format refresh.
  • Also relevant to the collection's reasoning-distillation entries (s1, LIMO-style recipes).