On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9% and LiveCodeBench v6 pass@12 reaches 70.9%, compared with 69.0% and 66.3% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.
Figures & tables
Figure 1: Overview of ACSD. (a) We extract a steering vector from the activation contrast between the model’s own trajectories that are correct within a generation budget and all remaining trajectories. (b) The vector steers hidden states at layer ℓ∗ to form the self-teacher. Teacher and student process the same problem and student-generated prefix, and the student learns by matching the teacher’s next-token distribution.
Figure 2: Variation among four discovery responses per problem, with 512 problems per model. (a) Fraction of problems with at least one correct response completed within the generation budget, computed cumulatively from the stored responses (Pass@4). Markers indicate ACSD’s budget B . (b) Fraction of eligible problems whose longest and shortest correct responses differ by at least the stated number of tokens. Eligibility requires at least two correct responses; eligible counts are 466, 469, 483, 476, and 476 in legend order.
DS-R1-0528-Qwen3-8B
DS-R1-Distill-Qwen-7B
OLMo-3-7B-Think
Method
AIME
AIME
HMMT
AIME
Avg.
AIME
AIME
HMMT
AIME
Avg.
AIME
AIME
HMMT
AIME
Avg.
24
25
25
26
24
25
25
26
24
25
25
26
Base
78.3
70.6
49.7
71.7
67.6
52.5
37.8
25.0
46.1
40.4
72.5
68.3
44.4
71.1
64.1
SFT
77.2
71.4
50.0
74.2
68.2
52.2
38.6
25.0
45.6
40.3
71.4
63.1
40.3
63.6
59.6
GRPO
79.2
73.9
55.8
74.2
70.8
56.1
41.4
26.4
48.3
43.1
74.7
68.6
46.9
71.4
65.4
OPSD
78.9
71.4
51.1
74.7
69.0
56.4
40.0
25.6
51.4
43.3
75.8
68.1
46.7
71.4
65.5
Table 1: Mathematical reasoning (Avg@12, %) and, bottom right, LiveCodeBench v6 code generation (pass@ k , %). Bold marks the best result per column within a model; averages use unrounded scores.
Figure 3: Direction ablations on DeepSeek-R1-0528-Qwen3-8B: mean Avg@12 over four mathematical benchmarks during the first 100 updates. (a) Trajectory contrasts. (b) Random, shuffled, and reversed directions. Controls share the layer and strength of the ACSD run in Table 1 ; dashed lines mark the base model.
Figure 4: Token-level updates on 256 held-out DeepSeek-R1-0528-Qwen3-8B trajectories under the training clipping rule. (a) Root-mean-square logit-update magnitude by response position. (b) ACSD and (c) OPSD signed updates for uncertainty estimation and final answers, as a percentage of the role’s total absolute update; both teachers share prefixes, marker groups, and color scale.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
B
Layer
α
Training window
DS-R1-0528-Qwen3-8B
8,192
18
0.30
4,096
DS-R1-Distill-Qwen-7B
4,096
14
0.30
4,096
OLMo-3-7B-Think
4,096
16
0.10
4,096
Qwen3-8B
4,096
13
0.05
4,096
Qwen3-4B
4,096
13
0.05
4,096
Appendix
Table 3: ACSD settings for the main mathematical reasoning experiments. B is used for trajectory labeling and activation extraction; α is the intervention strength relative to the current hidden-state norm.
Model family
Max tokens
Temperature
Top- p
Top- k
DeepSeek-R1
38,912
0.6
0.95
–
OLMo-3-Think
38,912
0.6
0.95
–
Qwen3
38,912
1.0
0.95
20
Appendix
Table 4: Decoding settings for mathematical reasoning benchmarks.
Figure 5: Recovery scan on DeepSeek-R1-0528-Qwen3-8B. Each cell reports the percentage of 512 continuations that pass answer verification. The black outline marks layer 18 and α=0.3 , used for distillation. Recovery without steering is 26.6%.
Figure 6: Distillation results for six intervention settings on DeepSeek-R1-0528-Qwen3-8B. The horizontal axis shows the best mean Avg@12 across four mathematical reasoning benchmarks. Labels also give recovery; the outline marks the main setting.
Figure label
Annotator label
Definition
Setup
initializing
Restating the task, setting up variables, or planning.
Deduction
deduction
Deriving a conclusion from the current assumptions.
Known facts
adding-knowledge
Recalling a theorem, formula, identity, or external fact.
Examples
example-testing
Checking a case, example, or sanity check.
Uncertainty
uncertainty- estimation
Expressing uncertainty, hedging, or assessing confidence.
Backtracking
backtracking
Abandoning, revising, or correcting an approach.
Appendix
Table 5: Reasoning roles used in the token-level audit.
Group
Examples
Hedging
probably, perhaps, maybe, possibly
Contrast
But, but, However, however
Revision
wait, Wait, Actually, actually, reconsider
Verification
check, verify
Progress
so, then, Then, implies
Conclusion
Therefore, Thus, hence, conclude
Appendix
Table 6: Marker groups used in the token-level update analysis.
Figure 7: Complete role and marker updates on 256 held-out trajectories. (a) ACSD. (b) OPSD. (c) ACSD minus OPSD normalized updates. Rows are marker groups; columns are reasoning roles.
Figure 8: Training curves for different windows on DeepSeek-R1-0528-Qwen3-8B. We report mean Avg@12 across four mathematical reasoning benchmarks over the first 100 updates. (a) ACSD. (b) OPSD. Lines connect evaluations; dashed lines mark the base model.
Model
Layer / α
Base
Steered
Δ
DS-R1-0528-Qwen3-8B
(18,0.30)
67.6
25.6
−42.0
DS-R1-Distill-Qwen-7B
(14,0.30)
40.4
13.8
−26.6
OLMo-3-7B-Think
(16,0.10)
64.1
62.2
−1.9
Qwen3-8B
(13,0.05)
63.3
62.1
−1.2
Qwen3-4B
(13,0.05)
62.3
62.4
+0.1
Appendix
Table 7: Direction intervention throughout generation. Values are average Avg@12 (%) across four mathematical reasoning benchmarks; Δ denotes Steered minus Base. Intervention settings come from Table 3 .
On-policy self-distillation trains reasoning models on their own trajectories using dense token distribution guidance from a privileged teacher with access to a verified solution. This supervision transfers what the teacher predicts without directly transferring where it attends within the preceding context. We introduce On-Policy Attention Self-Distillation (OPASD), which complements token-level supervision with solution-conditioned attention distillation. Because the privileged teacher can attend to verified solution tokens unavailable to the student, OPASD projects teacher attention onto student-visible positions and renormalizes the resulting distribution before alignment. Across three model sizes and four competition-level mathematics benchmarks, OPASD consistently outperforms token-only OPSD, improving average accuracy by 4.98 to 8.40 percentage points. OPASD also avoids the response-length inflation and performance degradation observed with token-only distillation, reducing generated rollout tokens by 73.9% and estimated model compute by 72.6% while training 1.53x faster. These results show that solution-conditioned attention provides a complementary supervision signal that makes on-policy self-distillation more accurate, stable, and compute-efficient.
Safaeid Hossain Arib, Rabeya Akter, Ismam Nur Swapnil +3
ACI PLC, Bangladesh · University of Dhaka, Bangladesh · BRAC University, Bangladesh +2
On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising direction for advancing reasoning capability without a stronger external teacher. Yet in math reasoning the gains are inconsistent, even when the same approach succeeds elsewhere. A pointwise mutual information analysis traces the failure to the privileged context itself: it inflates the teacher's confidence on tokens already implied by the solution (structural connectives, verifiable claims) and deflates it on deliberation tokens ("Wait", "Let", "Maybe") that drive multi-step search. We propose Anti-Self-Distillation (AntiSD), which ascends a divergence between student and teacher rather than descending it: this reverses the per-token sign and yields a naturally bounded advantage in one step. An entropy-triggered gate disables the term once the teacher entropy collapses, completing a drop-in replacement for default self-distillation. Across five models from 4B to 30B parameters on math reasoning benchmarks, AntiSD reaches the GRPO baseline's accuracy in 2 to 10x fewer training steps and improves final accuracy by up to 11.5 points. AntiSD opens a path to scalable self-improvement, where a language model bootstraps its own reasoning through its training signal.
Guobin Shen, Xiang Cheng, Chenxiao Zhao +4
1Xiaohongshu Inc. · Institute of Automation, Chinese Academy of Sciences
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu