Fine Until Fine-Tuned: Repeated Solutions Make Reasoning Fragile
Organizations: Aix-Marseille University
Abstract
Recipes such as s1 and LIMO teach a model to reason with little data by showing it the same thousand or fewer worked solutions many times over. Judged when that training ends, the repetition looks harmless. But reasoning models are often trained again, and we find that repetition leaves their reasoning fragile to that next stage, even when the stage has nothing to do with reasoning. We fine-tuned Qwen3.5-9B-Base on its own correct solutions to competition math problems, either drilling a few hundred of them about eight times each or showing many more once; with the same amount of training, both solve about 95% of held-out problems. A single pass of ordinary instruction tuning leaves the once-trained model where it was, while the drilled one falls to 86.0%, and harsher later stages take it to 59.3% or below. A third model that visited the drilled problems just as often, with a new solution at every visit, was unharmed, so the damage comes from seeing the same texts again rather than from having few problems. The break recurs with a stronger model's traces, in further training runs and on other models and tasks. It is also cheap to undo: the reasoning is suppressed rather than erased, and five updates of reasoning training bring almost all of it back, as does brief training on the reasoning format with almost no mathematics. Fresh solutions prevented the damage, and so did replaying 6.25% of the original solutions in a gentler later stage, so our claim concerns later training without such replay. Sharpening alone does not explain the break, since a model sharpened three-quarters as much without repetition was unharmed. On a skill the base model could not perform within a token budget, repetition mainly cost learning.
Figures & tables
| Stress tests | ||||||
|---|---|---|---|---|---|---|
| Training texts | Model | Passes | Handoff | One epoch | Instructions | Answer-only |
| None | Base | – | 94.7 | 94.5 | 95.6 | 91.7 |
| Own solutions | Drilled | 1.1 | 95.4 | – | 94.5 | 94.3 |
| Drilled | 3.3 | 95.1 | – | 94.0 | 92.8 | |
| Drilled | 7.7 | 95.4 | 86.0 | 59.3 | 22.5 | |
| Once-trained | 1.0 | 94.6 | 94.8 | 94.7 | 94.8 | |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Arm | Trained on | Updates | Passes |
|---|---|---|---|---|
| Math | D | 579 of the base model’s correct solutions | 20–140 | 1.1–7.7 |
| O | 4,480 of the base model’s correct solutions | 60, 140 | 0.43, 1.0 | |
| P | D’s 579 problems, a fresh correct solution at each visit | 140 | 1.0 | |
| D2 | 579 other solutions from O’s set | 140 | 7.7 | |
| D-r128 | D’s solutions, rank-128 adapters | 140 | 7.7 | |
| D-T | gpt-oss-120b traces for 575 problems | 20, 140 | 1.1, 7.8 |
| At handoff | Forced after | At cap | ||||||
|---|---|---|---|---|---|---|---|---|
| Task | Model | Passes | Free | Forced | (%) | Instr. tuning | Answer-only | after AO |
| Math, 9B | Base | – | 74.7 | 94.7 | 0 | 95.6 | 91.7 | 7.0 |
| D-u20 | 1.1 | 95.0 | 95.4 | 0.1 | 94.5 | 94.3 | 5.8 | |
| D-u60 | 3.3 | 97.3 | 95.1 | 3.4 | 94.0 | 92.8 | 4.2 | |
| D-u140 | 7.7 | 95.2 | 95.4 | 24 | 59.3 | 22.5 | 2.6 | |
| O-u140 | 1.0 | 94.1 | 94.6 | 0.1 | 94.7 | 94.8 | 4.5 | |
| 35B-A3B | Nemotron | |||||
| State | H | IT | AO | H | IT | AO |
| MATH-500, levels 3–5 | ||||||
| Base | 96.0 | 48.1 | 93.7 | 93.6 | 83.6 | 94.1 |
| D-u20 | 96.0 | 93.4 | 96.0 | 93.1 | 88.8 | 90.4 |
| D-u60 | 96.3 | 91.6 | 79.6 | 95.6 | 69.8 | 92.4 |
| D-u140 | 94.9 | 50.8 | 21.5 | 93.4 | 22.3 | 54.2 |
| Crediting | Median | Under 100 | Ends | Closes | ||||
|---|---|---|---|---|---|---|---|---|
| Model and stage | Correct | integers | Unboxed | tokens | tokens | its turn | block | At cap |
| D-u140, answer-only | 22.5 | 32.2 | 40.3 | 103 | 49 | 97 | 0.9 | 2.6 |
| D-u140, instruction tuning | 59.3 | 67.2 | 20.1 | 451 | 7.5 | 78 | 29 | 5.9 |
| O-u140, answer-only | 94.8 | 95.0 | 3.6 | 1,294 | 0.0 | 0.0 | 96 | 4.5 |
| Base, answer-only | 91.7 | 92.2 | 5.7 | 1,474 | 1.0 | 11 | 84 | 7.0 |
| Instr. tuning | Answer-only | |
| Excess over O-u140 | ||
| First pass | 36.2 [31.9, 40.4] | 73.1 [69.3, 76.7] |
| Budget rule | 21.3 [16.9, 25.7] | 68.3 [64.3, 72.2] |
| Answer rule | 26.5 [22.3, 30.5] | 58.0 [53.7, 62.2] |
| D-u140, forced accuracy | ||
| First pass | 59.3 | 22.5 |
| State | First pass | Budget rule | Answer rule |
|---|---|---|---|
| Base | 94.7 | 92.8 | 94.7 |
| Base, instr. tuning | 95.6 | 93.3 | 95.7 |
| Base, answer-only | 91.7 | 82.5 | 92.3 |
| D-u140 | 95.4 | 93.7 | 95.4 |
| D-u140, instr. tuning | 59.3 | 71.7 | 69.0 |
| D-u140, answer-only | 22.5 | 26.1 | 37.6 |
| Comparison | Later stage | Excess | Interval |
|---|---|---|---|
| Drilled vs. once-trained | instructions, one epoch | 9.6 | [6.9, 12.4] |
| Fresh solutions vs. once-trained | instructions, one epoch | 0.2 | [ 1.9, 1.6] |
| Drilled vs. once-trained | instructions, stress test | 36.2 | [31.9, 40.4] |
| answer-only, stress test | 73.1 | [69.3, 76.7] | |
| Drilled vs. fresh solutions | instructions, stress test | 36.4 | [32.4, 40.4] |
| answer-only, stress test | 72.5 | [68.8, 76.0] |
| At handoff | After twenty answer-only updates | |||||||
| Arm | Forced, 4,096 | Forced, 16,384 | Free, 4,096 | 4,096 | 16,384 | At cap | Strict | |
| F-u140 | 61.7 | 81.5 | 55.7 | 4.2 | 11.7 | 87.5 | 4.9 | 0.139 |
| FA1 | 51.0 | 75.3 | 62.2 | 25.8 | 48.2 | 49.0 | 13.0 | 0.0085 |
| FA025 | 53.6 | 78.6 | 61.7 | 7.6 | 27.9 | 71.1 | 10.2 | 0.022 |
| FN | 58.3 | 85.7 | 55.2 | 4.9 | 25.3 | 72.4 | 6.5 | 0.013 |
| U2 | 45.3 | 70.6 | 22.4 | 0.3 | 6.5 | 92.2 | 2.1 | 0.154 |
| State | 4–7 | 8–15 | 16–31 | 32–63 | 64–127 | 128–255 | 256–511 | 512–1,024 |
|---|---|---|---|---|---|---|---|---|
| T-u30 | 0.0057 | 0.0003 | 0.0010 | 0.0017 | 0.0004 | 0.0006 | 0.0003 | 0.0002 |
| S-u230 | 0.018 | 0.0040 | 0.0033 | 0.0044 | 0.0044 | 0.0039 | 0.0033 | 0.0032 |
| U-u20 | 0.0059 | 0.0026 | 0.0019 | 0.0030 | 0.0037 | 0.0029 | 0.0026 | 0.0027 |
| F-u20 | 0.0017 | 0.0021 | 0.0003 | 0.0025 | 0.0032 | 0.0030 | 0.0026 | 0.0024 |
| U-u60 | 0.025 | 0.0020 | 0.0008 | 0.019 | 0.020 | 0.019 | 0.017 | 0.013 |
| F-u60 | 0.015 | 0.0004 | 0.0072 | 0.0092 | 0.014 | 0.020 | 0.016 | 0.016 |
| State | Passes | , full | , 1,024 | (%) | Kept |
|---|---|---|---|---|---|
| T-u30 | – | 0.0001 | 0.0004 | 1.3 | 0.98 |
| S-u230 | 1.15 | 0.0034 | 0.0035 | 2.2 | 0.91 |
| U-u20 | 1.1 | 0.0027 | 0.0027 | 1.9 | 0.89 |
| F-u20 | 1.1 | 0.0024 | 0.0025 | 1.9 | 0.90 |
| U-u60 | 3.3 | 0.012 | 0.015 | 3.7 | 0.74 |
| F-u60 | 3.3 | 0.013 | 0.016 | 3.8 | 0.74 |
| Origin | Update 1 | Update 2 | Update 3 |
|---|---|---|---|
| M0 | 10.45 | 8.06 | 9.31 |
| T-u30 | 10.46 | 7.89 | 9.10 |
| S-u230 | 10.43 | 7.74 | 10.10 |
| U-u20 | 10.89 | 8.05 | 9.08 |
| F-u20 | 10.97 | 7.86 | 9.11 |
| U-u60 | 10.56 | 8.52 | 9.38 |
| State | Resumed | Strict | Share resumed | Free |
|---|---|---|---|---|
| M0 | 54.9 | 13.8 | 52 | 13.5 |
| T-u30 | 62.2 | 22.1 | 48 | 44.0 |
| S-u230 | 57.0 | 14.6 | 51 | 37.8 |
| U-u20 | 59.9 | 10.4 | 61 | 43.8 |
| F-u20 | 58.6 | 22.9 | 43 | 51.8 |
| U-u60 | 64.8 | 14.1 | 62 | 46.1 |
| State | Answer-only, | Answer-only, | Instr. tuning |
|---|---|---|---|
| M0 | 45.1 / 9.9 | – | 53.6 / 26.3 |
| T-u30 | 59.4 / 22.9 | – | – |
| S-u230 | 57.6 / 8.6 | – | – |
| U-u20 | 58.1 / 25.3 | 61.5 / 10.4 | 58.9 / 26.8 |
| F-u20 | 62.5 / 22.1 | 60.9 / 21.6 | 57.8 / 37.0 |
| U-u60 | 46.4 / 7.6 | – | – |
| State | Handoff | After twenty answer-only updates |
|---|---|---|
| M0 | 77.6 / 19.3 (67) | 73.4 / 18.8 (61) |
| T-u30 | 78.1 / 22.9 (63) | 74.7 / 34.6 (45) |
| S-u230 | 77.9 / 26.3 (59) | 75.8 / 10.9 (81) |
| U-u140 | 75.5 / 17.4 (66) | 6.0 / 0.8 (35) |
| F-u140 | 81.5 / 68.5 (14) | 11.7 / 4.9 (20) |
| At handoff | After twenty answer-only updates | |||||||
| State | 4,096 | 8,192 | 16,384 | 4,096 | 8,192 | 16,384 | Over | Cap |
| M0 | 50.0 | 65.1 | 77.6 | 50.8 | 64.8 | 73.4 | 46 | 22.1 |
| T-u30 | 58.6 | 71.6 | 78.1 | 55.5 | 65.6 | 74.7 | 41 | 18.0 |
| S-u230 | 54.2 | 68.8 | 77.9 | 54.2 | 65.6 | 75.8 | 41 | 16.1 |
| U-u140 | 52.9 | 65.4 | 75.5 | 1.8 | 2.9 | 6.0 | 98 | 93.2 |
| F-u140 | 60.4 | 73.4 | 81.5 | 4.2 | 7.3 | 11.7 | 95 | 87.5 |
| Model and stage | 2,048 | 4,096 | 8,192 | 16,384 | Capped | Unboxed | Mean tokens |
|---|---|---|---|---|---|---|---|
| Base, handoff | 67.2 | 79.6 | 88.3 | 94.7 | 5.1 | 3.8 | 3,340 |
| Base, instruction tuning | 71.8 | 83.7 | 91.1 | 95.6 | 2.5 | 1.9 | 2,712 |
| Base, answer-only | 67.9 | 78.5 | 86.5 | 91.7 | 7.0 | 5.7 | 3,608 |
| O-u140, handoff | 68.3 | 80.1 | 88.8 | 94.6 | 5.4 | 3.8 | 3,370 |
| O-u140, instruction tuning | 67.5 | 81.3 | 89.3 | 94.7 | 4.8 | 4.1 | 3,196 |
| O-u140, answer-only | 67.9 | 79.5 | 89.6 | 94.8 | 4.5 | 3.6 | 3,162 |
| Removed | Problems |
| Answer not a single integer | 12,493 |
| Solution boxes several different answers | 2,614 |
| Duplicate copies disagree | 107 |
| Figures | 331 |
| Prompt over 512 tokens | 28 |
| Contamination hits | 91 |
| Source | Forced | Free | Direct | Median tokens |
|---|---|---|---|---|
| MATH-500 levels 1–2 | 1.00 | 1.00 | 0.58 | 850 |
| MATH-500 level 3 | 0.92 | 0.83 | 0.42 | 1,326 |
| MATH-500 level 4 | 0.88 | 0.71 | 0.17 | 954 |
| MATH-500 level 5 | 0.62 | 0.58 | 0.08 | 2,928 |
| OlympiadBench (integer) | 0.46 | 0.42 | 0.00 | 6,144 (cap) |
| NuminaMath-TIR (integer) | 0.73 | 0.67 | 0.19 | 1,410 |
| Model | Handoff | After | Loss | Interval |
|---|---|---|---|---|
| Base, new adapter | 94.7 | 94.5 | 0.2 | [ 1.2, 1.6] |
| Drilled | 95.4 | 86.0 | 9.4 | [6.8, 12.1] |
| Once-trained | 94.6 | 94.8 | 0.2 | [ 1.5, 1.0] |
| Fresh solutions | 94.6 | 95.0 | 0.5 | [ 1.6, 0.7] |
| Training texts | Model | Handoff | After |
|---|---|---|---|
| Own solutions | Drilled | 95.4 | 94.3 |
| Once-trained | 94.6 | 94.7 | |
| gpt-oss-120b traces | Drilled | 95.0 | 95.5 |
| Once-trained | 96.3 | 95.8 |
| Origin | 1 | 2 | 5 | 20 | Habit, | ||
|---|---|---|---|---|---|---|---|
| D + IT | 59.3 | 84.7 | 90.3 | 93.9 | 94.3 | 0.98 [0.94, 1.01] | 94.0 (0.98) |
| D + AO | 22.5 | 51.0 | 79.5 | 92.8 | 93.3 | 0.97 [0.95, 1.00] | 92.3 (0.97) |
| D-T + IT | 33.7 | 76.4 | 87.1 | 93.1 | 93.3 | 0.98 [0.94, 1.00] | – |
| D-T + AO | 1.1 | 21.6 | 70.5 | 91.5 | 93.0 | 0.96 [0.94, 0.99] | – |
| O + IT | 94.7 | 95.2 | 94.2 | 94.3 | 94.6 | – | – |
| Model | Handoff | Instructions | Answer-only | |
|---|---|---|---|---|
| Sharpened | 0.082 | 93.1 | 95.0 | 92.4 |
| Once-trained | 0.0004 | 94.6 | 94.7 | 94.8 |
| Drilled | 0.111 | 95.4 | 59.3 | 22.5 |
| Run | Model | Handoff | IT | Gentle |
|---|---|---|---|---|
| Second | Drilled | 94.8 | 32.7 | 80.9 |
| Once-trained | 96.2 | 95.1 | 95.4 | |
| Fresh solutions | 95.6 | 94.6 | 94.0 | |
| Drilled’s excess | 61.1 [57.0, 65.2] | 13.1 [10.0, 16.3] | ||
| Third | Drilled | 94.6 | 45.9 | 76.1 |
| Once-trained | 94.9 | 95.1 | 94.2 |
| Model | Handoff | Gentle | Stress test | Loss, gentle |
|---|---|---|---|---|
| Base | 0.2 | – | – | – |
| Drilled (D-S) | 37.9 | 26.6 | 0.1 | 11.3 [7.8, 14.8] |
| Once-trained (O-S) | 66.3 | 50.8 | 0.4 | 15.5 [11.7, 19.3] |
| Fresh solutions (P-S) | 58.1 | 38.7 | 0.4 | 19.4 [15.4, 23.3] |
| ID | Prediction | Threshold | Observed | Outcome | Claim |
|---|---|---|---|---|---|
| RS1 | One epoch breaks D | excess 5, CI | 9.6 | held | break |
| RS2 | One epoch spares P | P O within 3 | 0.2 | held | cause |
| RP1 | Replay mitigates, own solutions | reduction 50% | 91%, prevents | held | remedy |
| RL1 | Relearning is fast | of D + IT 0.5 | 0.98 | held | competence |
| RL2 | The ceiling control holds | O + IT within 2 at every | at most 0.6 | held | competence |
| SH1 | Sharpening alone does less than half | AO excess of O-sharp 36.5 | 0.9 | held | signature |
| ID | Prediction | Threshold | Observed | Outcome | Claim |
|---|---|---|---|---|---|
| Math pilot | |||||
| MP1 | Drilling over-sharpens | (D) (O) 5% of the base’s NLL, CI | 24% against 0.1% | held | signature |
| MP2 | rises with passes | D-u20 D-u60 D-u140 | 0.1, 3.4, 24% | held | signature |
| MP3 | The drilled model loses more, each stage | more than D-u20 and O, CIs | 36.2 (IT), 73.1 (AO) over O | held | break |
| MP5 | Free accuracy falls more than forced | base and O, both stages | all four | held | decision |
| Fresh solutions (P) | |||||
| ID | Prediction | Threshold | Observed | Outcome | Claim |
|---|---|---|---|---|---|
| Depth profile and cap test | |||||
| DP1 | Repetition, not training amount, moves the reasoning | (S) below (U) and (F), CIs | 0.0035 against 0.107, 0.161 | held | signature |
| DP2 | Every lasting below every drilled | separation | highest lasting 0.0004, lowest drilled 0.061 | held | signature |
| DP3 | rises with passes in U and F | monotone | monotone | held | signature |
| DP4 | ranks the loss | Spearman 0.7 | 0.93 | held | signature |
| CT1 | Competence lost, not slowed | U, F 15% within 16,384 tokens | 6.0, 11.7 | held | break |