On-policy distillation (OPD) strengthens language-model reasoning, yet whether students acquire new factual knowledge or compositional skill for multi-step reasoning remains unknown. We separate these capabilities using a controlled synthetic framework that measures the student's initial capabilities and independently controls the teacher's additional facts, compositional skill, or both. Across four models from three families, reverse-KL OPD reliably transfers compositional skill across unseen reasoning structures, but transfers minimal factual knowledge. Decoupling the distillation recipe reveals the source of this asymmetry: replacing reverse KL with forward KL restores factual transfer, whereas student rollouts specifically improve the execution of multi-step reasoning. Experiments on recent factual QA and competition mathematics show a similar asymmetry under reverse-KL OPD, yielding notable reasoning gains without factual memory expansion. Together, these results demonstrate that on-policy distillation does not expand a model's parametric knowledge, but instead teaches it to organize and compose the knowledge it already possesses.
Figures & tables
Figure 1 : Overview of the setup. (a) Random lookup tables represent facts; chains and trees specify how their query results are combined. (b) Starting from S0 , teachers add exposure to held-out facts ( ΔF ), branching trees ( ΔS ), or both, with frozen S0 as the baseline. (c) A fixed teacher supervises student rollouts with reverse KL. (d) Evaluation separates factual recall on held-out facts from compositional skill over initial facts.
(a) Student and teacher capabilities
Model
Factual knowledge
Compositional skill
S0
Initial facts
Chain execution
T∅=S0
Initial facts
Chain execution
Tf
Initial + held-out facts
Chain execution
Ts
Initial facts
Chain and branching-tree execution
Tfs
Initial + held-out facts
Chain and branching-tree execution
Table 1 : Student and teacher capabilities and distillation conditions.
Figure 2 : Student performance after on-policy distillation under each teacher. Top: greedy final-answer accuracy; bottom: final-answer Pass@32 (at least one of 32 sampled completions correct). Numbers above each group show the unweighted mean across four models. All values are percentages.
Figure 3 : Effects of prefix source and KL direction. Rows show Qwen3-4B and Gemma-2-2B. Factual knowledge is evaluated after distillation from Tf ; compositional skill after distillation from Ts . Accuracy measures greedy final answers; NLL is reported for reference completions in bits per completion. Reference prefixes come from teacher demonstrations.
Model
Role / Condition
Knowledge-2026
Math-2026
Qwen3-4B
Initial student S0
11.25
17.41
Knowledge teacher TK
84.55
13.93
Reasoning teacher TR
12.24
43.78
Distilled from TK
9.64
19.40
Distilled from TR
10.75
23.88
Table 2 : Greedy accuracy (%) before and after on-policy distillation on real-world benchmarks. Shaded cells indicate each teacher’s target capability.
Figure 4 : Greedy execution traces from Qwen3-4B students after distillation. (a) Seven-step chain over held-out facts: reverse KL maintains correct execution syntax but outputs incorrect lookup values ( red ), whereas forward KL restores correct lookups. (b) Unseen tree over initial facts: forward KL recalls all facts correctly but flattens the tree into a linear sweep ( red ), whereas reverse KL from Ts executes the valid postorder. Red highlights incorrect values or misrouted operands; green marks correct final answers.
(a) Parameter Metric
Qwen3-4B
Gemma-2-2B
OLMo3-7B
Update Norm: Tf (factual)
2.01–2.86
1.33–1.52
3.08–3.89
Update Norm: Ts (compositional)
1.27–1.55
1.13–1.60
2.50–3.21
Update Norm: T∅ (control)
0.30–0.32
0.20–0.24
0.48–0.55
Cosine Similarity: Tf vs. Ts
0.004–0.018
0.011–0.036
0.004–0.046
Cosine Similarity: Rev. vs. Fwd. KL ( Ts )
0.73–0.78
0.64–0.82
0.68–0.80
Table 3: LoRA weight updates relative to S0 . (a) Update magnitude and directional alignment across models. (b) Layer-wise norm distribution across depth quartiles in Qwen3-4B. Full breakdowns across all projections and embedding drift appear in Appendix E (Table 6 ).
Teacher Tf
Teacher Ts
Objective
Fact-Single
Comp-Unseen
Fact-Single
Comp-Unseen
S0 (no distillation)
0.1
1.8
0.1
1.8
Reverse KL
10.9
0.7
0.1
76.8
Forward KL
54.1
0.5
0.1
76.4
GKD ( β=0.5 )
9.1
0.7
0.1
70.6
ToDi
64.3
0.5
0.1
86.4
Table 4: Greedy final-answer accuracy (%) on Qwen3-4B with alternative distillation objectives. All distillation runs use student prefixes.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Factual Knowledge
Compositional Skill
Model
Profile
Fact-Single
Fact-Chain
Comp- Shared
Comp- TeacherOnly
Comp- Unseen
Qwen3-0.6B
S0 (student)
0.00
0.00
0.15
2.10
0.20
Tf
100
100
0.20
2.05
0.45
Ts
0.00
0.00
100
100
99.95
Tfs
100
100
100
99.95
100
Gemma-2-2B
S0 (student)
0.00
0.00
0.40
2.20
0.45
Appendix
Table 5 : Greedy final-answer accuracy (%) of S0 and the specialized teachers on factual knowledge and compositional skill tests. Test definitions appear in Appendix B.3 . All tests use exact match of the final four-digit answer. Green shading highlights high accuracy.
Figure 5 : Training curves for on-policy distillation in the synthetic experiments. Columns correspond to distillation from Tf (a, d), Ts (b, e), and the joint Tfs (c, f). Top row: exact-match accuracy of sampled student rollouts on the distillation prompts; bottom row: reverse KL in nats. All runs train for three epochs over 20,000 prompts.
(a) Teacher: prefix, objective
Qwen3-4B
Gemma-2-2B
OLMo3-7B
Tf : student, reverse KL
2.01 / 2.4
1.33 / 3.9
3.08 / 2.3
Tf : student, forward KL
2.64 / 2.5
1.49 / 3.6
3.85 / 2.0
Tf : teacher, forward KL
2.86 / 2.5
1.52 / 3.6
3.89 / 2.1
Tf : teacher, reverse KL
2.27 / 2.5
1.34 / 3.8
3.20 / 2.1
Ts : student, reverse KL
1.27 / 3.3
1.13 / 3.9
2.50 / 2.6
Ts : student, forward KL
1.55 / 2.7
1.60 / 3.6
3.21 / 2.5
Appendix
Table 6 : Complete weight updates relative to S0 . (a) Global update norm / mean effective rank. (b) Cosine similarity between updates. (c) Mean embedding drift for held-out / initial fact identifiers. (d) Mean layer norms by depth quartile. (e) Norms by projection. Ranges cover the four prefix–KL combinations; control ranges cover two prompt sets.
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu