On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
Figures & tables
Figure 1: Comparison of OPSD and OPASD. (a) OPSD provides token-level supervision through teacher-student next-token distribution alignment. (b) OPASD provides both token- and attention-level supervision through joint next-token and projected attention distribution alignment.
Figure 2: Overview of On-Policy Attention Self-Distillation (OPASD). (1) Given a problem x and its verified solution y⋆ , the student samples an on-policy reasoning trajectory y^ from x alone. (2) The student and privileged teacher evaluate the same trajectory under their respective conditioning contexts, producing next-token distributions ptS,ptT and attention distributions AtS,AtT . (3) Because the teacher can attend to reference-solution positions unavailable to the student, its attention is restricted to the student-visible support and renormalized to obtain AtT . (4) OPASD jointly aligns the token and projected attention distributions, with gradients propagated only through the student.
Method
AIME24
AIME25
AIME26
HMMT25
Average
Qwen3-8B
Base
60.83
48.89
51.39
29.17
47.57
OPSD
59.16
51.97
53.33
33.33
49.45
OPASD
67.50
57.22
61.11
35.83
55.42
Qwen3-4B
Base
58.33
47.50
51.11
30.27
46.80
Table 1: Comparison of on-policy self-distillation methods across model scales. Avg@12 accuracy (%) of the instruction-tuned Base model, token-only OPSD, and OPASD on four competition-level mathematical reasoning benchmarks. OPASD achieves the highest performance across all three Qwen3 model sizes.
Objective
AIME24
AIME25
AIME26
HMMT25
Avg
FKL
39.72
33.05
33.05
21.11
31.73
RKL
40.83
32.22
35.83
20.56
32.36
JSD
41.94
32.50
35.83
21.11
32.85
Table 2: Effect of the attention divergence objective. JSD achieves the highest average accuracy across the three objectives.
λattn
AIME24
AIME25
AIME26
HMMT25
Avg
0.25
43.33
34.16
35.56
21.94
33.75
0.5
43.05
35.27
34.44
23.33
34.02
1.0
41.94
32.50
35.83
21.11
32.85
2.0
39.44
32.28
34.44
21.11
31.82
Table 3: Effect of attention loss weight. With the token-loss coefficient fixed at 1 , λattn=0.5 yields the highest average performance.
Layers
AIME24
AIME25
AIME26
HMMT25
Avg
1
43.05
35.27
34.44
23.33
34.02
2
41.13
34.57
33.36
21.11
32.54
3
39.44
30.83
33.88
20.00
31.04
Table 4: Effect of attention distillation depth. Using only the final transformer layer achieves the highest average performance compared with using the final two or three layers.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Seed / Statistic
AIME24
AIME25
AIME26
HMMT25
Average
Seed 1
63.27
50.83
55.33
34.72
51.04
Seed 42
63.33
50.27
55.56
33.89
50.76
Seed 123
63.54
50.63
54.87
34.36
50.85
Mean
63.38
50.58
55.25
34.32
50.88
Std
0.14
0.28
0.35
0.42
0.14
95% CI
±0.35
±0.70
±0.87
±1.03
±0.36
Appendix
Table 5: Robustness of OPASD across random seeds. Avg@12 accuracy (%) across three Qwen3-4B training runs with seeds 1, 42, and 123. All runs use the same experimental configuration. The final rows report the mean, sample standard deviation, and 95% confidence interval across 3 seeds.
Method
AIME24
AIME25
AIME26
HMMT25
Average
Redistribution
40.00
31.38
34.44
19.72
31.39
Shuffled Redistribution
42.22
31.67
33.89
20.83
32.15
Student-Support Projection
43.05
35.27
34.44
23.33
34.02
Appendix
Table 6: Handling teacher-only attention mass. Student-support projection achieves a higher average score than redistribution or shuffled redistribution.
Clipping Strategy
AIME24
AIME25
AIME26
HMMT25
Average
Point-wise Clipping
40.00
28.61
33.61
21.11
30.83
Per-token Clipping
41.94
32.50
35.83
21.11
32.85
No Clipping
40.00
33.33
34.72
20.83
32.22
Appendix
Table 7: Effect of token-loss clipping. Per-token clipping achieves the highest average performance compared with point-wise clipping and no clipping.
Training Objective
AIME24
AIME25
AIME26
HMMT25
Average
Token only (OPSD)
36.70
28.33
32.78
18.33
29.04
Attention only
40.56
30.27
33.61
19.44
30.97
Token + Attention (OPASD)
43.05
35.27
34.44
23.33
34.02
Appendix
Table 8: Comparison of distillation objectives on Qwen3-1.7B. Full OPASD achieves the highest average accuracy and outperforms both token-only OPSD and attention-only distillation.
Method
AIME24
AIME25
AIME26
HMMT25
Average
OPSD (point-wise clipping)
36.70
28.33
32.78
18.33
29.04
OPSD (per-token clipping)
35.83
29.56
32.80
18.16
29.09
OPASD (per-token clipping)
43.05
35.27
34.44
23.33
34.02
Appendix
Table 9: Effect of attention supervision on Qwen3-1.7B with per-token clipping held fixed. The original OPSD configuration is included for reference.
Parameter
OPSD
OPASD
Training data
OpenThoughts-Math-30K (OPSD split)
Teacher
Frozen teacher, conditioned on the reference solution
Student thinking mode
Disabled
Disabled
Teacher thinking mode
Enabled
Enabled
LoRA rank r
64
64
LoRA scaling α
128
128
Appendix
Table 10: Training hyperparameters for OPSD and OPASD.
Parameter
Value
Maximum new tokens
16,384
Thinking mode
Enabled
Temperature
0.6
Top- p
0.95
Top- k
20
Min- p
0.0
Appendix
Table 11: Evaluation hyperparameters.
Symbol
Meaning
First Use
Data, models, and trajectories
S , N
Reasoning dataset and number of problem–solution pairs
Sec. 3.1
xi , yi⋆
Problem and its verified reference solution
Sec. 3.1
pθ , pθ0
Student model with trainable parameters θ ; frozen initial model used by the teacher
Sec. 3.1
pS , pT
Student and privileged teacher policies
Eq. ( 1 )
y^ , T
Student-generated reasoning trajectory and its length
Eq. ( 2 )
Appendix
Table 12: Notation summary. Symbols used in the formulation and theoretical analysis of OPASD.
On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9% and LiveCodeBench v6 pass@12 reaches 70.9%, compared with 69.0% and 66.3% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
On-Policy Self-Distillation (OPSD) is commonly interpreted as the transfer of privileged information: a teacher observes the verified solution to the target problem and supervises the student's trajectory. However, this interpretation conflates two effects. The reference solution not only reveals the answer to the current instance but also changes the context under which the teacher provides token-level supervision. We investigate the role of target-specific privilege with OP2SD (On-Policy Self-Distillation from Other Problems), which replaces the paired reference with a problem and solution from a different example, while preserving the student rollout, teacher, and distillation objective. Across three models and three mathematics benchmarks, OP2SD improves over the base model, remains competitive with OPSD. The success of OP2SD implies that OPSD gains do not necessarily come from access to the reference solution, and that the teacher's context-induced behavior is an important factor.
Yuki Ichihara, Naoto Iwase, Mohammad Atif Quamar +1
Mohamed bin Zayed University of Artificial Intelligence · Nagoya University · RIKEN AIP