Overcoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Organizations: North South University · KAIST · Andria Labs
Abstract
On-policy self-distillation (OPSD) trains a student to match a privileged teacher distribution along its own sampled trajectory. Standard OPSD applies this supervision to unverified student rollouts while conditioning the teacher on privileged context, typically a reference solution. We separate these roles in a factorial analysis and find that scaffold correctness has a stronger effect on downstream accuracy than context correctness. Unverified scaffolds create an imitation gap because the teacher can use information unavailable to the student. This gap shrinks with model scale, yet OPSD continues to supervise mostly unverified trajectories. In contrast, verified scaffolds remain effective even when the teacher is conditioned on the student's own unsuccessful rollout. Based on this finding, we introduce OASIS, which retains the OPSD objective but supervises mostly verified by label on-policy trajectories and replaces written solutions with unverified model-generated attempts as the teacher context. OASIS therefore requires only final-answer labels. Across Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT 2025, OASIS improves over the base model by 3.2--3.8 points on average, while OPSD's gain falls from 3.05 points at 1.7B to 0.14 at 8B. At 8B, OASIS improves over OPSD by 3.05 points, showing that verified on-policy scaffolds preserve the effectiveness of self-distillation as models scale.
Figures & tables
| Qwen3-1.7B | Qwen3-4B | ||||
| Prefix | Context | Recovers | Recovers | ||
| Failed | Gold reference | 0.77 | +0.40 | 0.67 | +0.17 |
| Answer only | 0.61 | +0.23 | 0.59 | +0.09 | |
| Gold, answer removed | 0.59 | +0.22 | 0.57 | +0.08 | |
| Gold, answer removed ∗ | 0.51 | +0.13 | 0.50 | +0.01 | |
| Unrelated reference | 0.32 | 0.37 | |||
| Model | Method | AIME24 | AIME25 | HMMT25 | Average |
|---|---|---|---|---|---|
| Qwen3-1.7B | Base | 51.50 | 36.70 | 23.10 | 37.10 |
| SFT | 48.40 | 36.70 | 22.70 | 35.93 | |
| GRPO | 51.10 | 38.30 | 23.70 | 37.70 | |
| OPSD | 53.88 | 40.80 | 25.76 | 40.15 | |
| AVSD | 51.85 | 38.61 | 26.39 | 38.95 | |
| OASIS (ours) | 54.17 | 41.11 | 26.94 | 40.74 |
| Benchmark | Metric | OPSD | OASIS |
|---|---|---|---|
| AIME 2024 | Avg@12 | ||
| Pass@12 | |||
| Maj@12 | |||
| AIME 2025 | Avg@12 | 67.9 | |
| Pass@12 | 83.3 | ||
| Maj@12 | 73.3 |
| Training budget | Accuracy | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Model | Method | Seen | Trained | Rollouts | Written solution? | AIME24 | AIME25 | HMMT25 | Avg. |
| Qwen3-4B | OPSD † | 12.8K | 12.8K | 12.8K | Yes | 76.40 | 68.30 | 46.10 | 63.60 |
| AVSD † | 12.8K | 12.8K | 12.8K | Yes | 76.20 | 68.30 | 44.20 | 62.90 | |
| OASIS | 6.4K | 3.2K | 51.2K | No | 76.66 | 69.17 | 45.27 | 63.70 | |
| Qwen3-8B | OPSD † | 12.8K | 12.8K | 12.8K | Yes | 77.80 | 70.80 | 45.80 | 64.80 |
| AVSD † | 12.8K | 12.8K | 12.8K | Yes | 75.40 | 69.60 | 47.10 | 64.03 | |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Qwen3-1.7B | Qwen3-4B | |||
| OPSD | OASIS | OPSD | OASIS | |
| Problems receiving gradient | 100% | 51.7% | 100% | 50.8% |
| Selected scaffolds reaching the answer | 28.3% | 100% | 30.6% | 100% |
| Candidate rollouts hitting the length cap | 62.0% | 58.0% | 64.4% | 63.1% |
| Mean supervised length (tokens) | 859 | 603 | 883 | 646 |
| Generated tokens | 2.76M | 43.5M | 2.80M | 44.8M |
| Quantity at the supervised state | Verified rollouts | Failed rollouts |
|---|---|---|
| Share of OPSD’s training scaffolds | 28.3% | 71.7% |
| KL per token | 0.274 | 0.181 |
| Teacher entropy | 0.297 | 0.461 |
| Teacher probability on the sampled token | 0.834 | 0.775 |
| Teacher top-1 sampled token | 0.862 | 0.814 |
| Tokens with a clipped entry ( ) | 0.288 | 0.323 |
| Scaffold | Supervised length | Avg@12 | Pass@12 | Maj@12 |
|---|---|---|---|---|
| Shortest verified † | 603 | |||
| Longest verified | 909 |
| Generated tokens | Avg@12 | Pass@12 | Maj@12 | |
|---|---|---|---|---|
| 4 | 22.8M | |||
| 8 † | 43.5M |
| Hyperparameter | Qwen3-1.7B | Qwen3-4B | Qwen3-8B |
|---|---|---|---|
| Optimization & Batching | |||
| Optimizer steps | 100 | 100 | 100 |
| Problems per step | 64 | 64 | 64 |
| Problems trained per step | |||
| Rollouts per problem ( , OASIS) | 8 | 8 | 8 |
| Learning rate | , linear decay, gradient clip | ||