Organizations: The Chinese University of Hong Kong · Zhejiang University · Shanghai Artificial Intelligence Laboratory · Shanghai Jiaotong University · Visual Information Processing and Learning, ICT, CAS · Hunan University · Westlake University · Nanjing University
On-policy self-distillation (OPSD) provides dense token-level supervision without a second model: one network acts as teacher with the reference solution and as student with only the problem. We identify a specific failure mode of this recipe. During training, student token entropy rises past the teacher's and remains elevated, a pattern we call entropy overshoot. We trace it to both sides of distillation. The reference-conditioned teacher is confident along its answer-directed reasoning path, but this confidence transfers poorly to student-generated prefixes, making its supervision overly tied to answer-specific cues rather than reusable reasoning patterns; meanwhile, the forward KL used by OPSD continually diffuses the student's predictive distribution without pulling it back. We introduce E2-OPSD to address both causes. Exemplar-guided teaching replaces the current answer with a retrieved solved neighboring problem, providing transferable reasoning guidance without revealing the destination and better matching student-reachable states. Entropy-aware distillation uses the student-teacher entropy gap to determine the direction and strength of each token's correction. E2-OPSD improves math reasoning by up to 4.3 points in mean@16 over OPSD, while out-of-domain evaluations show gains over the corresponding base models of up to 4.9 points in mean@16 and 5.5 points in pass@8. Despite these gains, E2-OPSD remains simple, requiring no additional forward passes or networks.
Figures & tables
Figure 1: Left: Entropy overshoot in Qwen3-1.7B. Right: The reference-conditioned teacher is less certain on student prefixes, suggesting a reasoning-path mismatch.
Figure 2: (a) Exemplar-Guided Teaching. Standard OPSD gives the teacher the current reference solution, letting it jump from the student prefix toward the answer. E 2 -OPSD instead uses a solved neighbor as a strategy guide, forcing the teacher to follow the student’s intermediate step and solve the current problem itself, turning answer leakage into strategy transfer. (b) Forward and reverse KL play complementary roles. Forward KL promotes exploration over teacher-supported alternatives, while reverse KL removes unsupported probability mass to concentrate the policy. Entropy-aware distillation switches between them.
Figure 3: Overview of E 2 -OPSD .
Math Reasoning (mean@16)
Model
Method
AMC23
AIME24
AIME25
HMMT25
Avg.
Qwen3-0.6B (Thinking)
Base
45.8
9.8
14.2
8.3
19.5
OPSD
45.9
10.9
14.2
9.2
20.0
SDPO
45.5
9.4
14.6
7.8
19.3
GRPO
46.1
9.8
16.4
8.4
20.2
RLSD
46.8
11.0
15.1
7.8
20.2
Table 1: Main results across model scales and types on math reasoning, reported as mean@16 over three seeds. Green subscripts show our gains over Base, computed from the displayed values. Best values are bold and second-best distinct values underlined. Full evaluation results are in Appendix B .
Logic: ZebraLogic
Coding: HumanEval
Model
Method
mean@16
pass@8
mean@16
pass@8
Base
28.8
56.5
52.6
78.4
Qwen3-0.6B (Thinking)
E 2 -OPSD
29.0 +0.2
56.5 +0.0
53.1 +0.5
79.0 +0.6
Base
62.5
78.3
83.8
93.2
Qwen3-1.7B (Thinking)
E 2 -OPSD
67.4 +4.9
83.8 +5.5
84.5 +0.7
93.8 +0.6
Base
79.1
92.4
84.0
89.9
Table 2: Out-of-domain evaluation on logical reasoning and code generation. E 2 -OPSD consistently improves both mean@16 and pass@8 across ZebraLogic and HumanEval, showing that its gains extend beyond in-domain mathematical reasoning.
Figure 4: Training dynamics. (a,b) Student and teacher token entropy under OPSD (dashed) and E 2 -OPSD (solid); the dotted line marks the first student–teacher entropy crossing under OPSD. (c) On Qwen3-1.7B Thinking, routing share denotes the fraction of response tokens assigned to forward or reverse KL by BER. The routing naturally shifts from forward-dominated exploration toward increased reverse-KL concentration as training proceeds.
Figure 5: Component ablation and mechanism probes. (a) Performance contribution of each design component; “+” denotes the component added on top of OPSD. (b) EGT improves teacher–student compatibility. The x-axis denotes teacher-to-student forward KL evaluated on student prefixes; the y-axis shows response length from each setting’s own rollouts (log scale). (c) AGD sharpens entropy-aware routing. HL denotes high-student/low-teacher entropy and LH the reverse.
Figure 6: State-dependent KL routing avoids fixed directional bias. KL-direction dynamics on Qwen3-1.7B Thinking with the teacher context fixed and AGD disabled. The x-axis shows optimizer steps; the left y-axis reports student and teacher token entropy, while the right y-axis of panel (d) reports validation mean@8. Panels compare (a) forward KL, (b) reverse KL, (c) an equal-weight forward/reverse KL mixture, and (d) BER. The trajectories reveal distinct distributional dynamics under fixed versus state-dependent KL routing.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Backbone
Reasoning mode
Difficulty levels
Qwen3-0.6B
Thinking
1–3
Qwen3-1.7B
Thinking
5–7
Qwen3-4B-Instruct-2507
Instruct
6–8
Olmo-3-7B-Instruct
Instruct
7–10
Appendix
Table 3: Backbone configurations and DeepMath training difficulty bands.
Table 10: Out-of-domain evaluation on logical reasoning and code generation. ZebraLogic uses whole-puzzle correctness; best means are bold within each model and dataset.
Method
AMC23
AIME24
AIME25
HMMT25
Avg.
Base
75.50±0.28
46.04±1.78
34.58±0.29
21.88±0.78
44.50±0.27
OPSD
74.87±0.46
45.97±1.58
35.00±1.36
22.71±0.61
44.64±0.41
+ EGT
76.46±0.99
48.26±1.21
36.81±1.49
22.57±0.20
46.02±0.40
+ BER
77.26±0.38
52.01±1.13
37.15±0.94
24.03±0.35
47.61±0.64
+ BER + AGD
77.66±0.68
52.15±1.58
39.44±0.69
24.31±1.83
48.39±0.55
E 2 -OPSD (Ours)
79.02±0.41
52.29±0.45
39.24±0.26
25.00±0.90
48.89±0.32
Appendix
Table 11: Per-benchmark mean@16 with standard deviations for the component ablation on Qwen3-1.7B Thinking (Figure 5 (a)). Avg. equally weights the four benchmarks. Best means are bold and second-best distinct means underlined. AGD denotes adaptive gap damping.
Parameter
Value
Problems / responses per problem
256 / 8
Sampling temperature
1.0
Top- p / top- k
0.95 / 20
Maximum response length
16,384 tokens
Maximum total context length
40,960 tokens
Student prompt length limit
2,048 tokens
Appendix
Table 12: Diagnostic rollout settings. Indices i and g start at zero. The two teacher conditions share sampling seeds. Teacher generation budgets are also limited by the remaining context window.
Metric
Student
Reference teacher
EGT teacher
Mean response length (tokens)
9,139.7
835.6
2,089.5
Own-prefix entropy (nats)
0.2679
0.1900
0.2605
Forward KL to student (nats)
0
0.1549
0.1148
Appendix
Table 13: Teacher-context probe underlying Figure 5 (b). Lengths use 2,048 responses per condition; entropy and KL use 2,047 matched slots. Student zeros denote identical-prefix or self comparisons.
Bucket
Token share (%)
AGD off (%)
AGD on (%)
Change (pp)
HH
14.40
45.57
43.26
−2.31
HL
2.33
8.99
12.42
+3.43
LH
5.34
11.19
16.08
+4.89
LL
77.93
34.25
28.24
−6.01
Appendix
Table 14: AGD gradient probe on 2,048 fixed student responses. Token shares pool all response tokens; gradient shares use the response- and problem-level normalization described in Appendix C.1 . Changes are computed from the displayed gradient shares.
Subset
Positions
Mean overlap
Ot<0.5 (%)
All strongly damped positions
12,708,962
0.9848
0.380
with H(pt)≥ln2
766,682
0.8343
3.662
Appendix
Table 15: Overlap at strongly damped positions. Percentages are computed within each row.
School of Computer Science and Engineering, Southeast University, Nanjing 210096, China · Key Laboratory of Computer Network and Information Integration (Southeast University), Ministry of Education, China · Huawei Noah’s Ark Lab +2