On-policy self-distillation uses a model as its own teacher to provide dense supervision for reasoning, often through reference-solution conditioning. Providing privileged information does not by itself ensure effective token-level supervision throughout long responses. We introduce Activation-Conditioned Self-Distillation (ACSD), which extracts a steering vector by contrasting activations of self-generated trajectories that reach verified correct answers within a generation budget with those of all remaining trajectories. A frozen copy of the base model applies this vector at each prediction position, and the student learns from its next-token distributions on student-generated prefixes. Outcome verification is used for direction construction and calibration; distillation requires neither problem-specific reference text nor teacher parameter updates. The distilled student is used alone at inference. On each of five models, ACSD achieves the highest mean accuracy over four mathematical benchmarks among the evaluated methods. On DeepSeek-R1-0528-Qwen3-8B, mean mathematical accuracy reaches 71.9% and LiveCodeBench v6 pass@12 reaches 70.9%, compared with 69.0% and 66.3% for the reference-conditioned OPSD baseline. Contrasts among correct trajectories also support distillation, and extracted directions can be reused across mathematical training datasets. On fixed student trajectories, ACSD maintains more stable late-position logit-update magnitudes than OPSD.
Figures & tables
Figure 1: Overview of ACSD. (a) We extract a steering vector from the activation contrast between the model’s own trajectories that are correct within a generation budget and all remaining trajectories. (b) The vector steers hidden states at layer ℓ∗ to form the self-teacher. Teacher and student process the same problem and student-generated prefix, and the student learns by matching the teacher’s next-token distribution.
Figure 2: Variation among four discovery responses per problem, with 512 problems per model. (a) Fraction of problems with at least one correct response completed within the generation budget, computed cumulatively from the stored responses (Pass@4). Markers indicate ACSD’s budget B . (b) Fraction of eligible problems whose longest and shortest correct responses differ by at least the stated number of tokens. Eligibility requires at least two correct responses; eligible counts are 466, 469, 483, 476, and 476 in legend order.
DS-R1-0528-Qwen3-8B
DS-R1-Distill-Qwen-7B
OLMo-3-7B-Think
Method
AIME
AIME
HMMT
AIME
Avg.
AIME
AIME
HMMT
AIME
Avg.
AIME
AIME
HMMT
AIME
Avg.
24
25
25
26
24
25
25
26
24
25
25
26
Base
78.3
70.6
49.7
71.7
67.6
52.5
37.8
25.0
46.1
40.4
72.5
68.3
44.4
71.1
64.1
SFT
77.2
71.4
50.0
74.2
68.2
52.2
38.6
25.0
45.6
40.3
71.4
63.1
40.3
63.6
59.6
GRPO
79.2
73.9
55.8
74.2
70.8
56.1
41.4
26.4
48.3
43.1
74.7
68.6
46.9
71.4
65.4
OPSD
78.9
71.4
51.1
74.7
69.0
56.4
40.0
25.6
51.4
43.3
75.8
68.1
46.7
71.4
65.5
Table 1: Mathematical reasoning (Avg@12, %) and, bottom right, LiveCodeBench v6 code generation (pass@ k , %). Bold marks the best result per column within a model; averages use unrounded scores.
Figure 3: Direction ablations on DeepSeek-R1-0528-Qwen3-8B: mean Avg@12 over four mathematical benchmarks during the first 100 updates. (a) Trajectory contrasts. (b) Random, shuffled, and reversed directions. Controls share the layer and strength of the ACSD run in Table 1 ; dashed lines mark the base model.
Figure 4: Token-level updates on 256 held-out DeepSeek-R1-0528-Qwen3-8B trajectories under the training clipping rule. (a) Root-mean-square logit-update magnitude by response position. (b) ACSD and (c) OPSD signed updates for uncertainty estimation and final answers, as a percentage of the role’s total absolute update; both teachers share prefixes, marker groups, and color scale.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Model
B
Layer
α
Training window
DS-R1-0528-Qwen3-8B
8,192
18
0.30
4,096
DS-R1-Distill-Qwen-7B
4,096
14
0.30
4,096
OLMo-3-7B-Think
4,096
16
0.10
4,096
Qwen3-8B
4,096
13
0.05
4,096
Qwen3-4B
4,096
13
0.05
4,096
Appendix
Table 3: ACSD settings for the main mathematical reasoning experiments. B is used for trajectory labeling and activation extraction; α is the intervention strength relative to the current hidden-state norm.
Model family
Max tokens
Temperature
Top- p
Top- k
DeepSeek-R1
38,912
0.6
0.95
–
OLMo-3-Think
38,912
0.6
0.95
–
Qwen3
38,912
1.0
0.95
20
Appendix
Table 4: Decoding settings for mathematical reasoning benchmarks.
Figure 5: Recovery scan on DeepSeek-R1-0528-Qwen3-8B. Each cell reports the percentage of 512 continuations that pass answer verification. The black outline marks layer 18 and α=0.3 , used for distillation. Recovery without steering is 26.6%.
Figure 6: Distillation results for six intervention settings on DeepSeek-R1-0528-Qwen3-8B. The horizontal axis shows the best mean Avg@12 across four mathematical reasoning benchmarks. Labels also give recovery; the outline marks the main setting.
Figure label
Annotator label
Definition
Setup
initializing
Restating the task, setting up variables, or planning.
Deduction
deduction
Deriving a conclusion from the current assumptions.
Known facts
adding-knowledge
Recalling a theorem, formula, identity, or external fact.
Examples
example-testing
Checking a case, example, or sanity check.
Uncertainty
uncertainty- estimation
Expressing uncertainty, hedging, or assessing confidence.
Backtracking
backtracking
Abandoning, revising, or correcting an approach.
Appendix
Table 5: Reasoning roles used in the token-level audit.
Group
Examples
Hedging
probably, perhaps, maybe, possibly
Contrast
But, but, However, however
Revision
wait, Wait, Actually, actually, reconsider
Verification
check, verify
Progress
so, then, Then, implies
Conclusion
Therefore, Thus, hence, conclude
Appendix
Table 6: Marker groups used in the token-level update analysis.
Figure 7: Complete role and marker updates on 256 held-out trajectories. (a) ACSD. (b) OPSD. (c) ACSD minus OPSD normalized updates. Rows are marker groups; columns are reasoning roles.
Figure 8: Training curves for different windows on DeepSeek-R1-0528-Qwen3-8B. We report mean Avg@12 across four mathematical reasoning benchmarks over the first 100 updates. (a) ACSD. (b) OPSD. Lines connect evaluations; dashed lines mark the base model.
Model
Layer / α
Base
Steered
Δ
DS-R1-0528-Qwen3-8B
(18,0.30)
67.6
25.6
−42.0
DS-R1-Distill-Qwen-7B
(14,0.30)
40.4
13.8
−26.6
OLMo-3-7B-Think
(16,0.10)
64.1
62.2
−1.9
Qwen3-8B
(13,0.05)
63.3
62.1
−1.2
Qwen3-4B
(13,0.05)
62.3
62.4
+0.1
Appendix
Table 7: Direction intervention throughout generation. Values are average Avg@12 (%) across four mathematical reasoning benchmarks; Δ denotes Steered minus Base. Intervention settings come from Table 3 .
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu