Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Authors: Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou
Organizations: School of Computing and Data Science, The University of Hong Kong · Department of Computer Science and Technology, University of Cambridge · Shenzhen Loop Area Institute
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.
Figures & tables
Figure 1: Overview. (a) In OPD, the student writes a rollout and the teacher scores every token; the signal is strongest at decision tokens. (b) The swap readout places a student checkpoint in both student slots of the crosscoder, with the teacher’s activation fixed, and reads the features the checkpoint uses. (c) OPD reweights the features the student shares with the teacher, most at decision tokens, and acquires none of the teacher’s own; the warm-up moves the student partly along OPD’s direction and partly beyond it.
Figure 2: What the swap readout records for a feature on one input. A lit bulb marks a feature that fires, and a brighter bulb marks a stronger activation.
Setting
Base student model
Teacher model
Dimension (student / teacher)
JustRL
R1-Distill-Qwen-1.5B
JustRL-DeepSeek-1.5B
1536 / 1536
Skywork
R1-Distill-Qwen-1.5B
Skywork-OR1-Math-7B
1536 / 3584
R1-7B
R1-Distill-Qwen-1.5B
R1-Distill-Qwen-7B
1536 / 3584
Table 1: Detailed information of three OPD settings.
Figure 3: OPD reweights the student’s features and acquires none of the teacher’s. (a–c) Width-adjusted MAS of all features; dashed: MAS of 0.5 , above which one model dominates a feature. (d) Change in feature firing rates after OPD; dashed: ±20% .
Figure 4: The features OPD changes most fire on decision tokens. (a) Share of decision-token features among the k most-changed features; dashed: among all features. (b) Under JustRL, change in firing rate of the decision-token features among the most-changed ones, after OPD (arrows) and in the teacher (crosses). (c) Teacher–student KL at the steps that emit each category of decision token on the student’s rollouts, relative to the average step.
AIME24
AIME25
AMC23
Avg.
Recovered
Base student
4.6
2.1
22.5
9.7
–
OPD
7.9
4.6
40.3
17.6
30%
Warm-up + OPD
12.1
9.6
44.4
22.0
46%
Teacher
22.9
20.0
65.9
36.3
–
Table 2: The SFT warm-up makes OPD work better. avg@8 (%), the accuracy averaged over 8 samples per problem, of the student distilled from Qwen3-4B-Base-GRPO with and without the warm-up. Recovered: share of the teacher’s advantage over the base student that the student recovers.
Figure 5: SFT on teacher rollouts induces no model-specific features. NRN ( Shi et al., 2026 ) in two-model crosscoders: a model’s share of a feature’s decoder norm, near 1 if the feature is specific to that model and 1/2 (vertical line) if shared. (a,b) NRN of the teacher against the student before and after the warm-up. (c) NRN of the student after the warm-up against the student before it.
Figure 6: The SFT warm-up pre-reweights shared features in the direction of OPD. (a) Change in each feature’s firing rate after the warm-up against that after OPD; circles: binned means; dashed: y=x , where the warm-up would make all of OPD’s change; solid: fitted slopes for the warm-up and for a random shift. (b) Mean change of the features OPD raises and lowers most; stacked: the extra change OPD adds after the warm-up.
avg@8 (%)
Swap readout
Student
AIME24
AIME25
AMC23
Avg.
Δ
β
ρ
OPD
10.4
7.5
37.8
18.6
–
1.00
–
+ feature change
11.2
7.1
45.6
21.3
+2.7[+0.5,+6.0]
1.32
+0.93
+ shuffled change
7.5
4.6
35.6
15.9
−2.7[−5.4,+0.3]
0.95
−0.04
Warm-up + OPD
12.5
9.2
44.1
21.9
–
0.84
+0.69
− feature change
10.0
7.1
40.6
19.2
−2.7[−5.3,−0.4]
0.20
−0.48
Table 3: The reweighting found by the swap readout carries the warm-up’s benefit. avg@8 (%) of students with the warm-up’s feature change added at the crosscoder’s layer or removed from it, with the difference Δ from the unmodified student and its bootstrap interval, and their reweighting read with the swap readout. All rows are sampled in the same way, which differs from Table 2 (Appendix D.4 ).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
JustRL
Skywork
R1-7B
Qwen3
∥N∥F/∥S∥F (about 0.5 at initialization)
0.52
0.61
0.61
0.56
Cosine of the rows of WB and WO ( 0 at initialization)
−0.06
−0.24
−0.24
−0.16
∥δ∥/∥m∥
0.12
0.07
0.07
0.11
Own slot: NhB against 21ShB
1.04
1.28
1.29
1.16
Replaced slot: Nδ against 21Sδ
1.08
1.28
1.27
1.14
Joint encoding: Nδ against Sm
0.06
0.04
0.04
0.06
Appendix
Table 4: The part of the student encoders that training leaves undetermined. Ratios in the middle block compare the root mean square of the term in N with that of the corresponding term in S .
Hyperparameter
Value
Training prompts
DAPO-Math-17k, one epoch (279 steps)
Prompts per step × rollouts per prompt
64×4
Rollout sampling
temperature 1.0, top- p 1.0
Maximum prompt / response length
1,024 / 7,168 tokens
Teacher signal
log-ratio on the student’s top-16 tokens
KL penalty, entropy bonus, outcome reward
none
Appendix
Table 5: OPD hyperparameters, shared by all OPD runs in the paper.
Setting
Model
AIME24
AIME25
AMC23
Avg.
Recovered
–
Base student
29.3
23.8
71.7
41.6
–
JustRL
OPD student
46.9
36.0
86.9
56.6
80%
Teacher
52.6
38.4
90.4
60.5
–
Skywork
OPD student
39.1
30.0
79.0
49.4
26%
Teacher
67.2
52.1
93.9
71.1
–
R1-7B
OPD student
31.5
25.8
73.3
43.5
9%
Appendix
Table 6: Accuracy in the settings of Section 3 , avg@256 (%). All OPD students start from the same base student.
AIME24
AIME25
AMC23
Avg.
Recovered
Base student
4.6
2.1
22.5
9.7
–
Warm-up
9.6
6.3
40.9
18.9
35%
OPD
7.9
4.6
40.3
17.6
30%
Warm-up + OPD
12.1
9.6
44.4
22.0
46%
Teacher
22.9
20.0
65.9
36.3
–
Appendix
Table 7: Accuracy in the setting of Section 4 , avg@8 (%), including the student after the warm-up alone.
Hyperparameter
Value
Dictionary size
32,768 features
Sparsity
BatchTopK, k=50 active features per token on average over a batch
Inference
threshold estimated during training (from step 1,000, moving average with rate 0.999)
Batch size
4,096 token positions
Optimizer
Adam, β1=0.9 , β2=0.999
Learning rate
1.41×10−4 , 1,000 warm-up steps, no decay
Appendix
Table 8: Crosscoder hyperparameters, shared by all crosscoders in the paper.
FVE, joint encoding
FVE, swap readout
Setting
L0
Base
OPD
Teacher
Base
OPD
JustRL
50.5
0.816
0.813
0.811
0.817
0.814
Skywork
48.9
0.814
0.822
0.767
0.814
0.821
R1-7B
48.7
0.814
0.824
0.755
0.815
0.823
Appendix
Table 9: Reconstruction quality of the crosscoders of Section 3 on the held-out tokens. L0: mean number of active features per token under the joint encoding. FVE: fraction of the variance of a model’s activation explained by its reconstruction, for the base student (Base), the OPD student (OPD), and the teacher.
Figure 7: The swap readout on held-out passages, before and after OPD under JustRL. Each pair of lines reads one passage in the student before and after OPD for one decision-token feature; shading: the feature’s activation on each token; bold: the tokens on which the feature starts or stops firing. Right: the feature’s firing count before OPD, after OPD, and in the teacher.
Figure 8: Features that the Skywork and R1-7B teachers dominate. Strongest contexts of each feature under the joint encoding of the three models, away from the start of the sequence; shading: the feature’s activation on each token; bold: the strongest token. Right: the width-adjusted MAS of the teacher and of the student before and after OPD.
Figure 9: Decision-token features that OPD reweights under JustRL and Skywork. Strongest held-out contexts of each feature in the student before OPD, read with the swap readout; shading and bold as in Figure 8 . Right: the feature’s firing count before OPD, after OPD, and, under JustRL, in the teacher.
Figure 10: Features that the warm-up moves along OPD’s direction under Qwen3. Contexts and shading as in Figure 9 . Right: the feature’s firing count in the student before the warm-up and its change after OPD, after the warm-up, and after the warm-up and OPD.
Figure 11: Features that carry the warm-up’s reweighting outside OPD’s direction under Qwen3. Contexts and shading as in Figure 9 . Right: the feature’s firing count in the student before the warm-up and its change after OPD, after the warm-up, and after the warm-up and OPD.
On-policy distillation (OPD) and on-policy self-distillation (OPSD) have emerged as promising post-training methods for large language models, offering dense token-level supervision on trajectories sampled from the model's own policy. However, existing results on their effectiveness remain mixed: while OP(S)D has shown promise in system prompt and knowledge internalization, recent studies also report instability and degradation. In this work, we present a comprehensive empirical study of when OPD and OPSD work, when they fail, and why. We find that OPD on mathematical reasoning is highly sensitive to teacher choice and loss formulation, whereas OPSD fails in our tested settings due to test-time absence of instance-specific privileged information (PI). In contrast, OPSD is effective when PI represents a shared latent rule, such as a system prompt or alignment preference. We identify three failure mechanisms: (1) distribution mismatch between teacher and student caused by conditioning on student-generated prefixes, (2) optimization instability from biased TopK reverse-KL gradients, and (3) an OPSD-specific limitation where the student learns a PI-free policy that aggregates PI-conditioned teachers, which is insufficient when PI is instance-specific. We further show that stop-gradient TopK objectives, RLVR-adapted teachers, and SFT-stabilized students mitigate these failures.
Siqi Zhu, Xuyan Ye, Hongyu Lu +2
1UIUC · 2Renmin University of China · †Work done during an internship at UIUC. +1
On-policy distillation (OPD) has become a key paradigm in LLM post-training, yet its training dynamics remain poorly understood. We present a systematic study examining the role, pathologies, and regulations of OPD. We first clarify the role of OPD as an exploration catalyst: it steers the student toward correct reasoning paths via dense token-level guidance, without expanding capability ceiling. We confirm this by showing that prompt diversity matters more than per-problem sampling numbers, and critically, that the effectiveness of OPD hinges entirely on the quality of its guiding signal. This dependency exposes two pathologies that derail exploration. The Student-Teacher Mismatch occurs when a large teacher-student distributional gap causes the guiding signal to misalign with task correctness, steering exploration in counterproductive directions. Length Exploitation arises when the aggregated token-level objective creates length-dependent shortcuts, allowing the student to game the reward landscape through response truncation or redundant padding, exploring degenerate length modes rather than reasoning strategies. To tame these pathologies, we investigate lightweight signal regulations: advantage clipping and log-scale compression, ensuring exploration is guided by faithful signals. Experiments across seven benchmarks demonstrate that these regulations alleviate length exploitation and enable effective distillation, stably surpassing OPD variants and RLVR baselines, thereby confirming that well-regulated signal quality, rather than mere teacher scale, governs successful exploration in OPD.
Rui Wang, Hongru Wang, Yi Chen +4
†The Chinese University of Hong Kong · ‡Tencent AI Lab
On-policy distillation (OPD) trains a student on its own rollouts with token-level supervision from teacher models, but its effectiveness can depend strongly on the warm-up stage before OPD. In this paper, we demystify warm-up for OPD from both data and training perspectives. For data, we find that effective warm-up relies on teacher-compatible chain-of-thought supervision, and that even incorrect teacher rollouts can provide comparable benefits to correct ones. This suggests that warm-up primarily transfers a teacher-compatible thinking pattern rather than merely correct answers. For training, we show that low-rank adaptation (LoRA) with a near-saturation training duration better balances in-domain adaptation and out-of-distribution generalization than full-parameter SFT. Based on these findings, we propose Simple-OPD, a plug-and-play initialization method that warms up the student on teacher-generated CoT with LoRA before OPD. Experiments across diverse settings demonstrate the effectiveness and robustness of Simple-OPD.
Tao Liu, Taiqiang Wu, Mao Zheng +5
1Tsinghua University · 2The University of Hong Kong · 3LLM Department, Tencent