Organizations: The Chinese University of Hong Kong · Zhejiang University · Shanghai Artificial Intelligence Laboratory · Shanghai Jiaotong University · Visual Information Processing and Learning, ICT, CAS · Hunan University · Westlake University · Nanjing University
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E2-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E2-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E2-OPSD remains simple, requiring no additional forward passes or networks.
Figures & tables
Figure 1: Left: Entropy overshoot in Qwen3-1.7B. Right: The reference-conditioned teacher is less certain on student prefixes, suggesting a reasoning-path mismatch.
Figure 2: (a) Exemplar-Guided Teaching. Standard OPSD gives the teacher the current reference solution, letting it jump from the student prefix toward the answer. E 2 -OPSD instead uses a solved neighbor as a strategy guide, forcing the teacher to follow the student’s intermediate step and solve the current problem itself, turning answer leakage into strategy transfer. (b) Forward and reverse KL play complementary roles. Forward KL promotes exploration over teacher-supported alternatives, while reverse KL removes unsupported probability mass to concentrate the policy. Entropy-aware distillation switches between them.
Figure 3: Overview of E 2 -OPSD .
Math Reasoning (mean@16)
Model
Method
AMC23
AIME24
AIME25
HMMT25
Avg.
Qwen3-0.6B (Thinking)
Base
45.8
9.8
14.2
8.3
19.5
OPSD
45.9
10.9
14.2
9.2
20.0
SDPO
45.5
9.4
14.6
7.8
19.3
GRPO
46.1
9.8
16.4
8.4
20.2
RLSD
46.8
11.0
15.1
7.8
20.2
Table 1: Main results across model scales and types on math reasoning, reported as mean@16 over three seeds. Green subscripts show our gains over Base, computed from the displayed values. Best values are bold and second-best distinct values underlined. Full evaluation results are in Appendix B .
Logic: ZebraLogic
Coding: HumanEval
Model
Method
mean@16
pass@8
mean@16
pass@8
Base
28.8
56.5
52.6
78.4
Qwen3-0.6B (Thinking)
E 2 -OPSD
29.0 +0.2
56.5 +0.0
53.1 +0.5
79.0 +0.6
Base
62.5
78.3
83.8
93.2
Qwen3-1.7B (Thinking)
E 2 -OPSD
67.4 +4.9
83.8 +5.5
84.5 +0.7
93.8 +0.6
Base
79.1
92.4
84.0
89.9
Table 2: Out-of-domain evaluation on logical reasoning and code generation. E 2 -OPSD consistently improves both mean@16 and pass@8 across ZebraLogic and HumanEval, showing that its gains extend beyond in-domain mathematical reasoning.
Figure 4: Training dynamics. (a,b) Student and teacher token entropy under OPSD (dashed) and E 2 -OPSD (solid); the dotted line marks the first student–teacher entropy crossing under OPSD. (c) On Qwen3-1.7B Thinking, routing share denotes the fraction of response tokens assigned to forward or reverse KL by BER. The routing naturally shifts from forward-dominated exploration toward increased reverse-KL concentration as training proceeds.
Figure 5: Component ablation and mechanism probes. (a) Performance contribution of each design component; “+” denotes the component added on top of OPSD. (b) EGT improves teacher–student compatibility. The x-axis denotes teacher-to-student forward KL evaluated on student prefixes; the y-axis shows response length from each setting’s own rollouts (log scale). (c) AGD sharpens entropy-aware routing. HL denotes high-student/low-teacher entropy and LH the reverse.
Figure 6: State-dependent KL routing avoids fixed directional bias. KL-direction dynamics on Qwen3-1.7B Thinking with the teacher context fixed and AGD disabled. The x-axis shows optimizer steps; the left y-axis reports student and teacher token entropy, while the right y-axis of panel (d) reports validation mean@8. Panels compare (a) forward KL, (b) reverse KL, (c) an equal-weight forward/reverse KL mixture, and (d) BER. The trajectories reveal distinct distributional dynamics under fixed versus state-dependent KL routing.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Reasoning mode
Difficulty levels
Qwen3-0.6B
Thinking
1–3
Qwen3-1.7B
Thinking
5–7
Qwen3-4B-Instruct-2507
Instruct
6–8
Olmo-3-7B-Instruct
Instruct
7–10
Appendix
Table 3: Backbone configurations and DeepMath training difficulty bands.
Table 10: Out-of-domain evaluation on logical reasoning and code generation. ZebraLogic uses whole-puzzle correctness; best means are bold within each model and dataset.
Method
AMC23
AIME24
AIME25
HMMT25
Avg.
Base
75.50±0.28
46.04±1.78
34.58±0.29
21.88±0.78
44.50±0.27
OPSD
74.87±0.46
45.97±1.58
35.00±1.36
22.71±0.61
44.64±0.41
+ EGT
76.46±0.99
48.26±1.21
36.81±1.49
22.57±0.20
46.02±0.40
+ BER
77.26±0.38
52.01±1.13
37.15±0.94
24.03±0.35
47.61±0.64
+ BER + AGD
77.66±0.68
52.15±1.58
39.44±0.69
24.31±1.83
48.39±0.55
E 2 -OPSD (Ours)
79.02±0.41
52.29±0.45
39.24±0.26
25.00±0.90
48.89±0.32
Appendix
Table 11: Per-benchmark mean@16 with standard deviations for the component ablation on Qwen3-1.7B Thinking (Figure 5 (a)). Avg. equally weights the four benchmarks. Best means are bold and second-best distinct means underlined. AGD denotes adaptive gap damping.
Parameter
Value
Problems / responses per problem
256 / 8
Sampling temperature
1.0
Top- p / top- k
0.95 / 20
Maximum response length
16,384 tokens
Maximum total context length
40,960 tokens
Student prompt length limit
2,048 tokens
Appendix
Table 12: Diagnostic rollout settings. Indices i and g start at zero. The two teacher conditions share sampling seeds. Teacher generation budgets are also limited by the remaining context window.
Metric
Student
Reference teacher
EGT teacher
Mean response length (tokens)
9,139.7
835.6
2,089.5
Own-prefix entropy (nats)
0.2679
0.1900
0.2605
Forward KL to student (nats)
0
0.1549
0.1148
Appendix
Table 13: Teacher-context probe underlying Figure 5 (b). Lengths use 2,048 responses per condition; entropy and KL use 2,047 matched slots. Student zeros denote identical-prefix or self comparisons.
Bucket
Token share (%)
AGD off (%)
AGD on (%)
Change (pp)
HH
14.40
45.57
43.26
−2.31
HL
2.33
8.99
12.42
+3.43
LH
5.34
11.19
16.08
+4.89
LL
77.93
34.25
28.24
−6.01
Appendix
Table 14: AGD gradient probe on 2,048 fixed student responses. Token shares pool all response tokens; gradient shares use the response- and problem-level normalization described in Appendix C.1 . Changes are computed from the displayed gradient shares.
Subset
Positions
Mean overlap
Ot<0.5 (%)
All strongly damped positions
12,708,962
0.9848
0.380
with H(pt)≥ln2
766,682
0.8343
3.662
Appendix
Table 15: Overlap at strongly damped positions. Percentages are computed within each row.
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil +3
ACI PLC, Bangladesh · University of Dhaka, Bangladesh · BRAC University, Bangladesh +2
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with OP2SD (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, OP2SD improves over the base model, remains competitive with OPSD. The success of OP2SD implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
Mohamed bin Zayed University of Artificial Intelligence · Nagoya University · RIKEN AIP
On-policy self-distillation (OPSD) provides denser token-level supervision and better computational efficiency than Reinforcement Learning with Verifiable Rewards (RLVR). However, this denser supervision may introduce substantial noise and training instability. Existing improvements often rely on high-variance per-token statistics and introduce extra hyperparameters and trade-offs. Based on the advantage formulation in RLVR, we analyze the OPSD objective from the same perspective, incorporating outcome correctness signals. We find that vanilla OPSD imposes insufficient penalties and excessive rewards on incorrect trajectories because it applies a fixed divergence objective regardless of outcome correctness. Furthermore, the reliability of teacher supervision is associated with both trajectory outcome and the cumulative average teacher entropy along the rollout. Based on these observations, we propose Outcome-Guided On-Policy Self-Distillation (OG-OPSD), which dynamically adapts both the divergence objective and distillation position according to binary outcome rewards and the cumulative average teacher entropy. Extensive experiments show that OG-OPSD consistently improves the performance of vanilla OPSD and multiple strong baselines in mathematical reasoning, multimodal reasoning, and out-of-distribution tasks across Qwen3 models at 1.7B, 4B, and 8B scales, as well as Qwen3-VL-2B.
ZheXu Wang, Mao-Lin Luo, Yankun Hong +6
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Huawei Noah’s Ark Lab +2