Understanding On-Policy Distillation: A Mechanistic Interpretability Perspective via Sparse Crosscoders
Authors: Zichao Yu, Qianshuo Ye, Xu Wang, Difan Zou
Organizations: School of Computing and Data Science, The University of Hong Kong · Department of Computer Science and Technology, University of Cambridge · Shenzhen Loop Area Institute
On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger teacher, yet what OPD actually distills into the student's internal representations remains unclear. We study this question with sparse crosscoders, which learn one feature dictionary shared by the student before and after OPD and the teacher. Standard crosscoder analyses, however, identify model-specific features but cannot tell how a model's use of its features changes, since all models are encoded into one set of feature activations. We therefore propose the swap readout, which reads each student checkpoint's feature activations on its own, measuring how training changes the student's use of each feature, even for checkpoints unseen by the crosscoder. Across three OPD settings, we find that OPD neither creates features nor passes on the teacher's own, and leaves the firing rates of over 98% of the student's frequently used features within 20%. We further examine the SFT warm-up on the teacher's rollouts that commonly precedes OPD and makes it more effective. Rather than adding features, the warm-up reweights the shared ones in two ways. First, it already raises and lowers many of the features that OPD later raises and lowers, doing part of OPD's work in advance. Second, it changes features that OPD alone would not, notably those for conversation format, reasoning style, and mathematical notation, and these changes persist through OPD. Imposing this reweighting on a directly distilled student's features, without changing its weights, brings its accuracy close to that of the warmed-up student, whereas the same change on shuffled features does not. Together, these findings suggest that OPD reweights existing features rather than acquiring new ones: the student learns from the teacher how to use the features they already share.
Figures & tables
Figure 1: Overview. (a) In OPD, the student writes a rollout and the teacher scores every token; the signal is strongest at decision tokens. (b) The swap readout places a student checkpoint in both student slots of the crosscoder, with the teacher’s activation fixed, and reads the features the checkpoint uses. (c) OPD reweights the features the student shares with the teacher, most at decision tokens, and acquires none of the teacher’s own; the warm-up moves the student partly along OPD’s direction and partly beyond it.
Figure 2: What the swap readout records for a feature on one input. A lit bulb marks a feature that fires, and a brighter bulb marks a stronger activation.
Setting
Base student model
Teacher model
Dimension (student / teacher)
JustRL
R1-Distill-Qwen-1.5B
JustRL-DeepSeek-1.5B
1536 / 1536
Skywork
R1-Distill-Qwen-1.5B
Skywork-OR1-Math-7B
1536 / 3584
R1-7B
R1-Distill-Qwen-1.5B
R1-Distill-Qwen-7B
1536 / 3584
Table 1: Detailed information of three OPD settings.
Figure 3: OPD reweights the student’s features and acquires none of the teacher’s. (a–c) Width-adjusted MAS of all features; dashed: MAS of 0.5 , above which one model dominates a feature. (d) Change in feature firing rates after OPD; dashed: ±20% .
Figure 4: The features OPD changes most fire on decision tokens. (a) Share of decision-token features among the k most-changed features; dashed: among all features. (b) Under JustRL, change in firing rate of the decision-token features among the most-changed ones, after OPD (arrows) and in the teacher (crosses). (c) Teacher–student KL at the steps that emit each category of decision token on the student’s rollouts, relative to the average step.
AIME24
AIME25
AMC23
Avg.
Recovered
Base student
4.6
2.1
22.5
9.7
–
OPD
7.9
4.6
40.3
17.6
30%
Warm-up + OPD
12.1
9.6
44.4
22.0
46%
Teacher
22.9
20.0
65.9
36.3
–
Table 2: The SFT warm-up makes OPD work better. avg@8 (%), the accuracy averaged over 8 samples per problem, of the student distilled from Qwen3-4B-Base-GRPO with and without the warm-up. Recovered: share of the teacher’s advantage over the base student that the student recovers.
Figure 5: SFT on teacher rollouts induces no model-specific features. NRN ( Shi et al., 2026 ) in two-model crosscoders: a model’s share of a feature’s decoder norm, near 1 if the feature is specific to that model and 1/2 (vertical line) if shared. (a,b) NRN of the teacher against the student before and after the warm-up. (c) NRN of the student after the warm-up against the student before it.
Figure 6: The SFT warm-up pre-reweights shared features in the direction of OPD. (a) Change in each feature’s firing rate after the warm-up against that after OPD; circles: binned means; dashed: y=x , where the warm-up would make all of OPD’s change; solid: fitted slopes for the warm-up and for a random shift. (b) Mean change of the features OPD raises and lowers most; stacked: the extra change OPD adds after the warm-up.
avg@8 (%)
Swap readout
Student
AIME24
AIME25
AMC23
Avg.
Δ
β
ρ
OPD
10.4
7.5
37.8
18.6
–
1.00
–
+ feature change
11.2
7.1
45.6
21.3
+2.7[+0.5,+6.0]
1.32
+0.93
+ shuffled change
7.5
4.6
35.6
15.9
−2.7[−5.4,+0.3]
0.95
−0.04
Warm-up + OPD
12.5
9.2
44.1
21.9
–
0.84
+0.69
− feature change
10.0
7.1
40.6
19.2
−2.7[−5.3,−0.4]
0.20
−0.48
Table 3: The reweighting found by the swap readout carries the warm-up’s benefit. avg@8 (%) of students with the warm-up’s feature change added at the crosscoder’s layer or removed from it, with the difference Δ from the unmodified student and its bootstrap interval, and their reweighting read with the swap readout. All rows are sampled in the same way, which differs from Table 2 (Appendix D.4 ).
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
JustRL
Skywork
R1-7B
Qwen3
∥N∥F/∥S∥F (about 0.5 at initialization)
0.52
0.61
0.61
0.56
Cosine of the rows of WB and WO ( 0 at initialization)
−0.06
−0.24
−0.24
−0.16
∥δ∥/∥m∥
0.12
0.07
0.07
0.11
Own slot: NhB against 21ShB
1.04
1.28
1.29
1.16
Replaced slot: Nδ against 21Sδ
1.08
1.28
1.27
1.14
Joint encoding: Nδ against Sm
0.06
0.04
0.04
0.06
Appendix
Table 4: The part of the student encoders that training leaves undetermined. Ratios in the middle block compare the root mean square of the term in N with that of the corresponding term in S .
Hyperparameter
Value
Training prompts
DAPO-Math-17k, one epoch (279 steps)
Prompts per step × rollouts per prompt
64×4
Rollout sampling
temperature 1.0, top- p 1.0
Maximum prompt / response length
1,024 / 7,168 tokens
Teacher signal
log-ratio on the student’s top-16 tokens
KL penalty, entropy bonus, outcome reward
none
Appendix
Table 5: OPD hyperparameters, shared by all OPD runs in the paper.
Setting
Model
AIME24
AIME25
AMC23
Avg.
Recovered
–
Base student
29.3
23.8
71.7
41.6
–
JustRL
OPD student
46.9
36.0
86.9
56.6
80%
Teacher
52.6
38.4
90.4
60.5
–
Skywork
OPD student
39.1
30.0
79.0
49.4
26%
Teacher
67.2
52.1
93.9
71.1
–
R1-7B
OPD student
31.5
25.8
73.3
43.5
9%
Appendix
Table 6: Accuracy in the settings of Section 3 , avg@256 (%). All OPD students start from the same base student.
AIME24
AIME25
AMC23
Avg.
Recovered
Base student
4.6
2.1
22.5
9.7
–
Warm-up
9.6
6.3
40.9
18.9
35%
OPD
7.9
4.6
40.3
17.6
30%
Warm-up + OPD
12.1
9.6
44.4
22.0
46%
Teacher
22.9
20.0
65.9
36.3
–
Appendix
Table 7: Accuracy in the setting of Section 4 , avg@8 (%), including the student after the warm-up alone.
Hyperparameter
Value
Dictionary size
32,768 features
Sparsity
BatchTopK, k=50 active features per token on average over a batch
Inference
threshold estimated during training (from step 1,000, moving average with rate 0.999)
Batch size
4,096 token positions
Optimizer
Adam, β1=0.9 , β2=0.999
Learning rate
1.41×10−4 , 1,000 warm-up steps, no decay
Appendix
Table 8: Crosscoder hyperparameters, shared by all crosscoders in the paper.
FVE, joint encoding
FVE, swap readout
Setting
L0
Base
OPD
Teacher
Base
OPD
JustRL
50.5
0.816
0.813
0.811
0.817
0.814
Skywork
48.9
0.814
0.822
0.767
0.814
0.821
R1-7B
48.7
0.814
0.824
0.755
0.815
0.823
Appendix
Table 9: Reconstruction quality of the crosscoders of Section 3 on the held-out tokens. L0: mean number of active features per token under the joint encoding. FVE: fraction of the variance of a model’s activation explained by its reconstruction, for the base student (Base), the OPD student (OPD), and the teacher.
Figure 7: The swap readout on held-out passages, before and after OPD under JustRL. Each pair of lines reads one passage in the student before and after OPD for one decision-token feature; shading: the feature’s activation on each token; bold: the tokens on which the feature starts or stops firing. Right: the feature’s firing count before OPD, after OPD, and in the teacher.
Figure 8: Features that the Skywork and R1-7B teachers dominate. Strongest contexts of each feature under the joint encoding of the three models, away from the start of the sequence; shading: the feature’s activation on each token; bold: the strongest token. Right: the width-adjusted MAS of the teacher and of the student before and after OPD.
Figure 9: Decision-token features that OPD reweights under JustRL and Skywork. Strongest held-out contexts of each feature in the student before OPD, read with the swap readout; shading and bold as in Figure 8 . Right: the feature’s firing count before OPD, after OPD, and, under JustRL, in the teacher.
Figure 10: Features that the warm-up moves along OPD’s direction under Qwen3. Contexts and shading as in Figure 9 . Right: the feature’s firing count in the student before the warm-up and its change after OPD, after the warm-up, and after the warm-up and OPD.
Figure 11: Features that carry the warm-up’s reweighting outside OPD’s direction under Qwen3. Contexts and shading as in Figure 9 . Right: the feature’s firing count in the student before the warm-up and its change after OPD, after the warm-up, and after the warm-up and OPD.