Solving Without Stopping: On-Policy Distillation at Small Scale
Organizations: University of Luxembourg · Seafill Open-Source Community · Universit´e Paris-Saclay
Abstract
On-policy distillation, where a student learns from a stronger teacher's feedback on its own outputs, is a common way to pass reasoning to smaller models. We analyze what it transfers at small scale, distilling Qwen3-8B into Qwen3 4B, 1.7B and 0.6B students, in thinking mode (reason at length, then end the reasoning and answer) and, for comparison, in non-thinking mode (no separate reasoning phase). Long reasoning needs two abilities, solving a problem and knowing when it is solved, and we find that distillation transfers the first, but in thinking mode not the second. Solving improves at every size, up to two ceilings, which we measure comprehensively across both modes and all student sizes: a student's single attempt never exceeds what it could already reach in many attempts before training, and the smaller the student, the further it stays below the teacher. Stopping is where the modes part. In non-thinking mode every student keeps stopping; in thinking mode students stop ending their reasoning early in training, and the smaller the student, the less of this ability survives: the teacher signals a stop almost only where a student already ends its reasoning, so distillation teaches no new stops; it only keeps the student's existing stops that land on a right answer, and a weak student has few such stops. The smallest students often reach the right value but do not commit to it: they either rarely mark it or mark it and write past it. Together, these results describe how small students behave under on-policy distillation, and a diagnostic that separates answer marking, correctness and stopping.
Figures & tables
| Start | After distillation | After as a share of | ||||
| Student | start | teacher | ||||
| Non-thinking OPD (teacher: 0.595, 0.869) | ||||||
| 4B | 0.427 | 0.736 | 0.549 | 0.791 | 0.75 | 0.922 |
| 1.7B | 0.294 | 0.646 | 0.440 | 0.725 | 0.68 | 0.739 |
| 0.6B | 0.048 | 0.512 | 0.328 | 0.593 | 0.64 | 0.551 |
| Thinking OPD, external SFT (teacher: 0.857, 0.955) | ||||||
| All samples | Truncated samples | |||||||
| Model | Truncated | With a marked answer | Correct value present | Chance | present vs. chance | First at | Chars. after it | Loop |
| Teacher (thinking) | 0.045 | 0.173 | 0.184 | 0.063 | 0.48 | 23,492 | 0.077 | |
| External-SFT start, before and after thinking-mode distillation | ||||||||
| 4B start | 0.274 | 0.035 | 0.067 | 0.028 | 0.08 | 33,181 | 0.115 | |
| after | 0.284 | 0.066 | 0.211 | 0.049 | 0.21 | 29,796 | 0.084 | |
| 1.7B start | 0.387 | 0.047 | 0.126 | 0.024 | 0.07 | 25,693 | 0.163 | |
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Student start | Teacher | Budget | Steps | Evaluated | Used in |
| Non-thinking route | |||||
| 4B pretrained-only | non-thinking | 7,168 | 212 | 212 | Section 3 |
| 1.7B pretrained-only | non-thinking | 7,168 | 212 | 212 | Section 3 |
| 0.6B pretrained-only | non-thinking | 7,168 | 212 | 212 | Section 3 |
| 4B pretrained-only, curriculum | non-thinking | 7,168 | 99 | – | Figure 1 |
| 1.7B pretrained-only, curriculum | non-thinking | 7,168 | 99 | – | Figure 1 |
| On-policy distillation (verl) | |
| prompts per step; updates per step | 128; 4 mini-batches of 32 |
| optimizer | AdamW, learning rate , 10 warm-up steps, then constant; weight decay 0.01; gradient clip 1.0 |
| advantage | clipped per-token reverse KL, Eq. (1), clip at ; no task reward |
| loss aggregation | sum over tokens, mean over sequences, divided by a constant of 7,168 |
| rollouts | 1 per prompt, temperature 1.0, budget 7,168 tokens (controls: 1,024 and 16,384) |
| length | 2 epochs 212 steps; thinking mode evaluated at step 142 |
| AIME 2026 | AMC23 | GSM8K-200 | Mean | ||||||
| Model | |||||||||
| Teacher (8B) | thinking | 0.674 | 0.900 | 0.960 | 1.000 | 0.937 | 0.965 | 0.857 | 0.955 |
| non-thinking | 0.163 | 0.667 | 0.694 | 0.975 | 0.929 | 0.965 | 0.595 | 0.869 | |
| Non-thinking route: from the pretrained-only model, non-thinking teacher | |||||||||
| 4B | start | 0.083 | 0.333 | 0.438 | 0.900 | 0.761 | 0.975 | 0.427 | 0.736 |
| after | 0.141 | 0.467 | 0.587 | 0.925 | 0.919 | 0.980 | 0.549 | 0.791 | |
| Model | ||||||||
| Teacher (thinking) | 0.857 | 0.898 | 0.923 | 0.938 | 0.948 | 0.954 | 0.955 | |
| Teacher (non-thinking) | 0.595 | 0.651 | 0.703 | 0.751 | 0.796 | 0.840 | 0.869 | |
| 4B pretrained-only | 0.427 | 0.528 | 0.597 | 0.651 | 0.693 | 0.723 | 0.736 | |
| + non-thinking route | 0.549 | 0.605 | 0.657 | 0.703 | 0.741 | 0.773 | 0.791 | 3 |
| 4B external-SFT start | 0.532 | 0.597 | 0.653 | 0.701 | 0.735 | 0.757 | 0.768 | |
| + thinking-mode distillation | 0.651 | 0.709 | 0.746 | 0.777 | 0.807 | 0.838 | 0.855 | 4 |
| Never solved at the start | Solved in fewer than half | Solved in at least half | |||||||||
| Route | Size | problems | after | contrib. | problems | start after | contrib. | problems | start after | contrib. | enter / leave |
| non-thinking | 4B | 29 | 0.025 | 42 | 0.219 0.437 | 199 | 0.819 0.967 | 8 / 2 | |||
| 1.7B | 37 | 0.017 | 86 | 0.250 0.523 | 147 | 0.747 0.946 | 11 / 4 | ||||
| 0.6B | 91 | 0.183 | 175 | 0.098 0.736 | 4 | 0.633 1.000 | 49 / 7 | ||||
| external SFT | 4B | 25 | 0.046 | 36 | 0.204 0.496 | 209 | 0.952 0.953 | 9 / 4 | |||
| 1.7B | 35 | 0.002 | 61 | 0.200 0.268 | 174 | 0.862 0.869 | 3 / 11 | ||||
| Route | Size | AIME 2026 | AMC23 | GSM8K-200 | Mean |
| (a) distilled p@1 minus starting p@n | |||||
| non-thinking route | 4B | [-0.328, -0.075] | [-0.421, -0.210] | [-0.084, -0.032] | [-0.246, -0.133] |
| 1.7B | [-0.125, -0.002] | [-0.538, -0.296] | [-0.188, -0.114] | [-0.255, -0.160] | |
| 0.6B | [-0.283, -0.032] | [-0.473, -0.222] | [-0.106, -0.008] | [-0.247, -0.126] | |
| external-SFT start | 4B | [-0.176, +0.023] | [-0.294, -0.112] | [-0.114, -0.056] | [-0.165, -0.074] |
| 1.7B | [-0.311, -0.063] | [-0.550, -0.282] | [-0.270, -0.187] | [-0.335, -0.212] | |
| Route | Size | marked-answer rate | answer accuracy | lenient p@1 | lenient p@n |
| non-thinking route | 4B | [+0.088, +0.143] | [+0.041, +0.088] | [+0.097, +0.149] | [+0.014, +0.102] |
| 1.7B | [+0.106, +0.161] | [+0.054, +0.111] | [+0.118, +0.177] | [+0.024, +0.139] | |
| 0.6B | [+0.598, +0.667] | [+0.065, +0.118] | [+0.249, +0.310] | [+0.008, +0.151] | |
| external-SFT start | 4B | [-0.043, +0.033] | [+0.142, +0.243] | [+0.083, +0.158] | [+0.031, +0.147] |
| 1.7B | [-0.180, -0.110] | [+0.072, +0.145] | [+0.013, +0.057] | [-0.084, +0.008] | |
| 0.6B | [-0.290, -0.210] | [+0.089, +0.142] | [+0.035, +0.077] | [-0.137, -0.007] |
| Model | committed p@1 start end | change | end minus chance | strict p@1 end | truncation start end |
| Teacher (non-thinking) | 0.499 | – | [+0.200, +0.300] | 0.497 | 0.026 |
| 4B | 0.157 0.397 | [+0.198, +0.282] | [+0.099, +0.195] | 0.361 | 0.462 0.125 |
| 1.7B | 0.191 0.275 | [+0.050, +0.117] | [-0.011, +0.062] | 0.270 | 0.181 0.111 |
| 0.6B | 0.049 0.234 | [+0.155, +0.215] | [-0.047, +0.017] | 0.233 | 0.581 0.154 |
| Run | steps 1–5 | 6–10 | 11–20 | 21–30 | first step below |
| External-SFT start, 7,168 tokens | |||||
| 4B, run 1 | 0.617 | 0.211 | 0.155 | 0.189 | 6 |
| 4B, run 2 | 0.633 | 0.208 | 0.170 | 0.212 | 6 |
| 1.7B | 0.445 | 0.128 | 0.071 | 0.051 | 6 |
| 0.6B, run 1 | 0.605 | 0.145 | 0.058 | 0.042 | 6 |
| 0.6B, run 2 | 0.583 | 0.113 | 0.055 | 0.052 | 6 |
| Run | stop reliability | closure, steps 1–4 | late closure | kept |
| 4B external-SFT start, run 1 | 0.521 | 0.648 | 0.226 | 0.349 |
| 4B external-SFT start, run 2 | 0.509 | 0.652 | 0.226 | 0.347 |
| 1.7B external-SFT start | 0.128 | 0.459 | 0.064 | 0.139 |
| 0.6B external-SFT start, run 1 | 0.042 | 0.611 | 0.023 | 0.038 |
| 0.6B external-SFT start, run 2 | 0.033 | 0.586 | 0.026 | 0.044 |
| 4B teacher-SFT start | 0.259 | 0.219 | 0.211 | 0.963 |
| Samples the student stopped | Truncated samples | Before training | |||||
| Student | at the stop | Stop-token advantage | Ever | Median of max | Stop reliability | ||
| 4B external-SFT start | 83 | 0.976 | 36 | 0.000 | 0.521 | ||
| 1.7B external-SFT start | 57 | 0.764 | 40 | 0.175 | 0.128 | ||
| 0.6B external-SFT start | 75 | 0.673 | 23 | 0.043 | 0.042 | ||
| Teacher (thinking), own samples | 0.885 | ||||||
| median rank | ||||||
| Position in the prefix | before | after | 95% interval | before after | ||
| Prefix set 1 | ||||||
| where the starting model closed | 205 | [ , ] | 0 0 | |||
| just after a correct marked answer | 71 | [ , ] | 64,401 67,323 | |||
| just after a wrong marked answer | 135 | [ , ] | 48,935 38,821 | |||
| at token 500 | 256 | [ , ] | 79,233 80,691 | |||
| closed | closed and correct | closed and wrong | |
| before training | 0.797 | 0.305 | 0.492 |
| null model, | 0.439 | 0.229 | 0.211 |
| after 20 thinking-mode steps | 0.438 | 0.311 | 0.127 |
| AIME 2026 | AMC23 | GSM8K-200 | |||||||
| Model | none | wrong | right | none | wrong | right | none | wrong | right |
| Teacher (thinking) | 0.092 | 0.234 | 0.674 | 0.014 | 0.026 | 0.960 | 0.007 | 0.056 | 0.937 |
| 4B external-SFT start | 0.537 | 0.315 | 0.148 | 0.242 | 0.210 | 0.549 | 0.020 | 0.081 | 0.899 |
| + distillation | 0.535 | 0.135 | 0.329 | 0.230 | 0.042 | 0.728 | 0.050 | 0.054 | 0.896 |
| 4B teacher-SFT start | 0.819 | 0.065 | 0.116 | 0.565 | 0.023 | 0.412 | 0.516 | 0.023 | 0.461 |
| + distillation | 0.531 | 0.153 | 0.316 | 0.228 | 0.044 | 0.728 | 0.038 | 0.059 | 0.903 |
| Model | Benchmark | unfinished | marked an answer | eligible | correct value present | chance | repeated ending |
| Teacher (thinking) | AIME 2026 | 154 | 0.143 | 143 | 0.147 | 0.070 | 0.058 |
| AMC23 | 36 | 0.250 | 1 | 0.000 | 0.000 | 0.111 | |
| GSM8K-200 | 58 | 0.207 | 30 | 0.367 | 0.033 | 0.103 | |
| 4B, distilled | AIME 2026 | 798 | 0.034 | 666 | 0.141 | 0.051 | 0.060 |
| AMC23 | 460 | 0.039 | 61 | 0.197 | 0.000 | 0.041 | |
| GSM8K-200 | 380 | 0.166 | 108 | 0.648 | 0.065 | 0.184 |
| all | solved 0/4 | solved 1–2/4 | solved 3–4/4 | |||||
| Model | recovered | recovered | recovered | recovered | ||||
| before training | 104 | 0.067 | 84 | 0.036 | 19 | 0.158 | 1 | 1.000 |
| after 20 thinking-mode steps | 288 | 0.101 | 206 | 0.049 | 60 | 0.183 | 22 | 0.364 |