On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.
Figures & tables
Figure 1 : Overview of the setup. (a) Random lookup tables represent facts; chains and trees specify how their query results are combined. (b) Starting from S0 , teachers add exposure to held-out facts ( ΔF ), branching trees ( ΔS ), or both, with frozen S0 as the baseline. (c) A fixed teacher supervises student rollouts with reverse KL. (d) Evaluation separates factual recall on held-out facts from compositional skill over initial facts.
(a) Student and teacher capabilities
Model
Factual knowledge
Compositional skill
S0
Initial facts
Chain execution
T∅=S0
Initial facts
Chain execution
Tf
Initial + held-out facts
Chain execution
Ts
Initial facts
Chain and branching-tree execution
Tfs
Initial + held-out facts
Chain and branching-tree execution
Table 1 : Student and teacher capabilities and distillation conditions.
Figure 2 : Student performance after on-policy distillation under each teacher. Top: greedy final-answer accuracy; bottom: final-answer Pass@32 (at least one of 32 sampled completions correct). Numbers above each group show the unweighted mean across four models. All values are percentages.
Figure 3 : Effects of prefix source and KL direction. Rows show Qwen3-4B and Gemma-2-2B. Factual knowledge is evaluated after distillation from Tf ; compositional skill after distillation from Ts . Accuracy measures greedy final answers; NLL is reported for reference completions in bits per completion. Reference prefixes come from teacher demonstrations.
Model
Role / Condition
Knowledge-2026
Math-2026
Qwen3-4B
Initial student S0
11.25
17.41
Knowledge teacher TK
84.55
13.93
Reasoning teacher TR
12.24
43.78
Distilled from TK
9.64
19.40
Distilled from TR
10.75
23.88
Table 2 : Greedy accuracy (%) before and after on-policy distillation on real-world benchmarks. Shaded cells indicate each teacher’s target capability.
Figure 4 : Greedy execution traces from Qwen3-4B students after distillation. (a) Seven-step chain over held-out facts: reverse KL maintains correct execution syntax but outputs incorrect lookup values ( red ), whereas forward KL restores correct lookups. (b) Unseen tree over initial facts: forward KL recalls all facts correctly but flattens the tree into a linear sweep ( red ), whereas reverse KL from Ts executes the valid postorder. Red highlights incorrect values or misrouted operands; green marks correct final answers.
(a) Parameter Metric
Qwen3-4B
Gemma-2-2B
OLMo3-7B
Update Norm: Tf (factual)
2.01–2.86
1.33–1.52
3.08–3.89
Update Norm: Ts (compositional)
1.27–1.55
1.13–1.60
2.50–3.21
Update Norm: T∅ (control)
0.30–0.32
0.20–0.24
0.48–0.55
Cosine Similarity: Tf vs. Ts
0.004–0.018
0.011–0.036
0.004–0.046
Cosine Similarity: Rev. vs. Fwd. KL ( Ts )
0.73–0.78
0.64–0.82
0.68–0.80
Table 3: LoRA weight updates relative to S0 . (a) Update magnitude and directional alignment across models. (b) Layer-wise norm distribution across depth quartiles in Qwen3-4B. Full breakdowns across all projections and embedding drift appear in Appendix E (Table 6 ).
Teacher Tf
Teacher Ts
Objective
Fact-Single
Comp-Unseen
Fact-Single
Comp-Unseen
S0 (no distillation)
0.1
1.8
0.1
1.8
Reverse KL
10.9
0.7
0.1
76.8
Forward KL
54.1
0.5
0.1
76.4
GKD ( β=0.5 )
9.1
0.7
0.1
70.6
ToDi
64.3
0.5
0.1
86.4
Table 4: Greedy final-answer accuracy (%) on Qwen3-4B with alternative distillation objectives. All distillation runs use student prefixes.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Factual Knowledge
Compositional Skill
Model
Profile
Fact-Single
Fact-Chain
Comp- Shared
Comp- TeacherOnly
Comp- Unseen
Qwen3-0.6B
S0 (student)
0.00
0.00
0.15
2.10
0.20
Tf
100
100
0.20
2.05
0.45
Ts
0.00
0.00
100
100
99.95
Tfs
100
100
100
99.95
100
Gemma-2-2B
S0 (student)
0.00
0.00
0.40
2.20
0.45
Appendix
Table 5 : Greedy final-answer accuracy (%) of S0 and the specialized teachers on factual knowledge and compositional skill tests. Test definitions appear in Appendix B.3 . All tests use exact match of the final four-digit answer. Green shading highlights high accuracy.
Figure 5 : Training curves for on-policy distillation in the synthetic experiments. Columns correspond to distillation from Tf (a, d), Ts (b, e), and the joint Tfs (c, f). Top row: exact-match accuracy of sampled student rollouts on the distillation prompts; bottom row: reverse KL in nats. All runs train for three epochs over 20,000 prompts.
(a) Teacher: prefix, objective
Qwen3-4B
Gemma-2-2B
OLMo3-7B
Tf : student, reverse KL
2.01 / 2.4
1.33 / 3.9
3.08 / 2.3
Tf : student, forward KL
2.64 / 2.5
1.49 / 3.6
3.85 / 2.0
Tf : teacher, forward KL
2.86 / 2.5
1.52 / 3.6
3.89 / 2.1
Tf : teacher, reverse KL
2.27 / 2.5
1.34 / 3.8
3.20 / 2.1
Ts : student, reverse KL
1.27 / 3.3
1.13 / 3.9
2.50 / 2.6
Ts : student, forward KL
1.55 / 2.7
1.60 / 3.6
3.21 / 2.5
Appendix
Table 6 : Complete weight updates relative to S0 . (a) Global update norm / mean effective rank. (b) Cosine similarity between updates. (c) Mean embedding drift for held-out / initial fact identifiers. (d) Mean layer norms by depth quartile. (e) Norms by projection. Ranges cover the four prefix–KL combinations; control ranges cover two prompt sets.
On-policy distillation transfers reasoning ability through dense token-level supervision, yet the nature of the transferable signal remains unclear. We discover that reasoning chains contain two types of knowledge that require different discovery mechanisms: decisions (where to branch), which surface through student uncertainty, and evidence (intermediate steps that justify decisions), which hides in positions where the student is confident yet wrong. Current methods capture only decisions; the substantive knowledge in evidence tokens remains untransferred. We propose DEAR(Decision-Evidence Aware Reasoning Distillation), which first identifies decisions via student entropy, then discovers their supporting evidence through hidden-state cosine similarity to decision anchors, boosted by teacher-student divergence to prioritize the largest knowledge gaps. Across three student-teacher configurations on math and code benchmarks, DEAR consistently outperforms standard OPD, with up to +2.5pp on competition math and +5.7pp on code generation.
Jinwei Xiao, Zhuowen Han, Yueqing Sun +6
Meituan Longcat Team · TJUNLP Lab, College of Intelligence and Computing, Tianjin University · Nanjing University
Logit-based knowledge distillation (KD) is used to train smaller language models (LMs) via supervision from stronger teachers, but whether its benefits are consistent across training stages remains unclear. Through controlled experiments, we find that forward Kullback-Leibler (KL) distillation--the standard KD formulation--with post-trained teachers behaves fundamentally differently during mid-training, an intermediate phase of self-supervised learning on curated corpora. Surprisingly, while forward KD simultaneously improves reasoning and factual recall during pre-training relative to standard next-token prediction (NTP), it instead slows factual recall acquisition during mid-training despite continued reasoning gains. We trace this stage dependence to an asymmetry in teacher confidence across data domains and the student's evolving knowledge state: teachers are more confident on procedural than knowledge-intensive data, while students acquire low-entropy factual knowledge earlier in training. To mitigate this imbalance, we propose Switch Distillation, a simple mid-training objective that distills on tokens where the teacher is confident, using teacher predictive entropy as a lightweight routing signal, and otherwise falls back to cross-entropy. Switch Distillation consistently outperforms existing distillation objectives across teacher sizes. Relative to standard NTP, it achieves 1.61-1.71x the reasoning performance and 1.13-1.19x the knowledge and commonsense performance while preserving 96.7-96.8% of factual recall. Crucially, these benefits persist after post-training: Switch Distillation closes the factual recall gap while maintaining 1.25-1.32x and 1.13-1.20x gains in reasoning and knowledge and commonsense, respectively.
Jacqueline He, Howard Yen, Shuyue Stella Li +9
Meta AI · University of Washington · Princeton University
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu