Authors: ZheXu Wang, Mao-Lin Luo, Yankun Hong, Zi-Hao Zhou, Bo Ye, Jian Zhao, Xialiang Tong, Min-Ling Zhang, +1 more
Organizations: School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Huawei Noah’s Ark Lab · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.
Figures & tables
Figure 1: (a) Implicit token-level advantages under forward and reverse KL. PT and PS denote the probabilities assigned to a candidate token by the teacher and student, respectively; A(⋅) denotes the corresponding advantage function. (b,d) Teacher completion accuracy (TCA) after resuming generation from student-generated reasoning prefixes of increasing length. Results are averaged over OpenThoughts-Math30K, DeepScaleR, MATH-500, and DeepMath datasets. (c,e) Cumulative average teacher entropy along correct and incorrect student trajectories.
Qwen3-1.7B
Qwen3-4B
Qwen3-8B
Method
AIME24
AIME25
HMMT25
Avg.
AIME24
AIME25
HMMT25
Avg.
AIME24
AIME25
HMMT25
Avg.
Base
51.5
36.7
23.1
37.1
74.9
66.4
42.2
61.1
75.8
65.6
43.9
61.7
SFT
48.4
36.3
22.7
35.8
70.2
62.3
43.4
58.6
72.3
64.2
42.9
59.8
GRPO
51.1
38.3
23.7
37.7
75.6
68.1
44.4
62.7
76.4
68.9
46.7
64.0
OPSD
57.2
41.1
28.8
42.3
75.6
67.9
44.5
62.6
77.8
67.5
45.8
63.7
EOPD
52.2
39.1
26.9
39.4
75.7
66.5
43.9
62.0
77.5
70.0
46.6
64.7
Table 1: Main results on AIME 2024/2025 and HMMT 2025 across Qwen3-1.7B, Qwen3-4B and Qwen3-8B. Scores are Avg@12 and Average is the unweighted mean across benchmarks.
Figure 2: Ablation results of divergence selection strategy and distillation position on AIME24, AIME25 and HMMT25 benchmarks.
Figure 3: Train with different Nt and evaluate on multiple datasets.
Table 5
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
View
User-message template
Student
Problem: {problem}
Please reason step by step, and put your final answer within \boxed{}.
Teacher
Problem: {problem}
Here is a reference solution to this problem:
=== Reference Solution Begin ===
{reference completion}
Appendix
Table 4: User-message templates for the Student view without privileged information and the reference-conditioned Teacher view.
Figure 4: TCA and cumulative average teacher entropy on Qwen3-1.7B.
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Yi Ding, Ruqi Zhang
Department of Computer Science, Purdue University, USA
We study on-policy self-distillation (OPSD), where a language model improves its reasoning ability by distilling privileged teacher distributions along its own on-policy trajectories. Despite its promise, OPSD can suffer from training instability due to a pattern mismatch between teacher and student responses. Self-reflected teacher responses may introduce reflection-induced biases and response templates that miscalibrate token-level supervision, ultimately harming the student's reasoning ability. To mitigate this issue, we propose OGLS-SD, an outcome-guided logit-steering framework that leverages verifiable outcome rewards to calibrate privileged teacher logits. Specifically, OGLS-SD contrasts teacher logits induced by successful and failed on-policy trajectories, constructing an outcome-discriminative steering direction for token-level guidance. Experiments on mathematical reasoning benchmarks show that OGLS-SD stabilizes self-distillation and improves performance over standard OPSD and other variants.
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E2-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E2-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E2-OPSD remains simple, requiring no additional forward passes or networks.
Yifei Liu, Minghao Fang, Xinyu Gu +6
The Chinese University of Hong Kong · Zhejiang University · Shanghai Artificial Intelligence Laboratory +5