Organizations: Shanghai Artificial Intelligence Laboratory, Shanghai, China · State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China
On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated by the student's policy. However, estimating the gradient of the reverse KL divergence in OPD remains a challenge. Using only the sampled token from the student-generated rollout is computationally cheap but provides limited distributional supervision, which will degrade accuracy. In addition, using the full vocabulary provides complete distributional supervision but is computationally expensive. Therefore, recent works propose Top-k OPD (TK-OPD) that use selected top-k tokens, which provides richer distributional supervision than sampled-token estimation at substantially lower computational cost than full-vocabulary estimation. Unfortunately, using only the selected top-k tokens induces bias, leading to accuracy degradation, as the probability mass outside the selected top-k tokens is discarded. To address the bias of TK-OPD, we propose Tail-Corrected Top-k On-Policy Distillation (TT-OPD). It preserves the advantages of TK-OPD, including rich distributional supervision and low computational cost, while providing an unbiased estimator of the gradient of the reverse KL divergence. The key insight of TT-OPD is to use not only the selected top-k tokens, but also the sampled token from the student-generated rollout, thereby recovering the discarded probability mass in expectation, avoiding the bias. Experimental results demonstrate that TT-OPD significantly outperforms other tested OPD variants.
Figures & tables
Figure 1: Overview of TT-OPD. The notation sg(⋅) denotes the stop-gradient operator. At each student-visited state Zt , the selected top- k tokens Stk provide rich distributional supervision, while the sampled token yt in the student-generated rollout provides an unbiased gradient tail correction. This only induces O(k+1) computational cost.
Unbiased Gradient?
Rich Distributional Supervision?
Computational Cost
ST-OPD
✓
×
O(1)
FV-OPD
✓
✓
O(∣V∣)
TK-OPD
×
✓
O(k)
TT-OPD
✓
✓
O(k+1)
Table 1: Comparison between our TT-OPD and other OPD variants. The notation V denotes the vocabulary.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
31.15
24.69
22.71
70.63
88.02
38.10
58.91
47.74
Qwen3-4B-Base Student
Student
8.02
5.31
6.15
30.80
54.07
22.79
25.22
21.77
ST-OPD
21.04
19.17
14.27
57.83
84.67
34.33
53.44
40.68
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
TT-OPD
26.77
23.33
22.92
64.87
86.08
36.81
54.70
45.07
Table 2: Main results on mathematical reasoning (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher for all tested algorithms. All tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected top- k tokens. Bold = best among the tested algorithms within each student block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
ST-OPD
21.04
19.17
14.27
57.83
84.67
34.33
53.44
40.68
k=8
TK-OPD
16.46
17.50
14.58
55.72
84.90
34.15
53.50
39.54
TT-OPD
27.60
22.60
21.88
62.88
85.70
36.35
55.96
44.71
k=16
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
Table 3: Sensitivity to the number of selected tokens (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected top- k tokens. ST-OPD does not depend on k . Bold = best among the tested algorithms for each value of k .
k
TK-OPD
TT-OPD
Qwen3-4B-Base Student
k=8
4 h 22 min
4 h 34 min
k=16
4 h 27 min
4 h 38 min
k=32
4 h 35 min
4 h 45 min
k=64
4 h 38 min
4 h 53 min
Qwen3-1.7B-Base Student
Table 4: Training time of TK-OPD and TT-OPD. Qwen3-8B-Base-GRPO-Math is the teacher. Except for k range in {8,16,32,64} , all other configurations are identical to those described in Section 5.1 .
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Student Top- k
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
TT-OPD
26.77
23.33
22.92
64.87
86.08
36.81
54.70
45.07
Teacher Top- k
TK-OPD
21.56
17.71
14.90
57.83
84.75
35.66
53.76
40.88
TT-OPD
27.08
25.83
20.21
64.83
86.58
35.66
54.70
44.98
Table 5: Results with different selected top- k tokens (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use DAPO-Math-17K as the training set. We set k=16 . Bold = best among the tested algorithms within each top- k block.
Figure 2: Code-generation results (average accuracy, %). Qwen3-8B-Base-GRPO-Code is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use the Code subset of Eurus-2-RL-Data as the training set. The student’s top- 16 tokens are used as the selected top- k tokens.
Figure 3: Tail-correction ablation (average accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base and Qwen3-1.7B-Base are the students, and all tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected top- k tokens.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
53.33
40.00
40.00
86.75
93.20
46.32
70.81
61.49
Qwen3-4B-Base Student
Student
26.67
26.67
20.00
63.86
76.20
36.03
42.37
41.69
ST-OPD
40.00
40.00
33.33
79.52
90.40
42.65
63.70
55.66
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
TT-OPD
56.67
43.33
43.33
80.72
92.20
37.87
64.89
59.86
Appendix
Table 7: Pass@ k results on mathematical reasoning (%). Unlike average accuracy, which averages correctness over all sampled responses, pass@ k is the percentage of problems for which at least one of the k sampled responses is correct. We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 2 . Bold = best among the tested algorithms within each student block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
ST-OPD
40.00
40.00
33.33
79.52
90.40
42.65
63.70
55.66
k=8
TK-OPD
40.00
23.33
26.67
73.49
90.60
41.18
62.22
51.07
TT-OPD
56.67
36.67
43.33
81.93
92.40
37.50
57.78
58.04
k=16
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
Appendix
Table 8: Pass@ k results for different numbers of selected tokens (%). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 3 . ST-OPD does not depend on k . Bold = best between TK-OPD and TT-OPD within each k block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Student Top- k
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
TT-OPD
56.67
43.33
43.33
80.72
92.20
37.87
64.89
59.86
Teacher Top- k
TK-OPD
46.67
30.00
30.00
72.29
88.20
37.87
62.37
52.49
TT-OPD
43.33
33.33
43.33
81.93
92.40
44.85
65.78
57.85
Appendix
Table 9: Pass@ k results with different selected top- k tokens (%). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 6 . Bold = best among the tested algorithms within each top- k block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
86.67
80.00
80.00
100.00
99.00
47.06
82.37
82.16
DeepSeek-R1-Distill-Qwen-7B Student
Student
83.33
73.33
76.67
96.39
98.00
44.85
82.52
79.30
ST-OPD
83.33
76.67
80.00
97.59
98.80
47.06
82.22
80.81
TK-OPD
80.00
70.00
76.67
98.80
98.60
47.06
83.26
79.20
TT-OPD
83.33
80.00
80.00
100.00
98.40
48.16
83.85
81.96
Appendix
Table 10: Pass@ k results on another model series (%). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 6 . Bold = best among the tested algorithms.
HumanEval+
MBPP+
LiveCodeBench
Avg.
Teacher
77.44
71.96
27.43
58.94
Qwen3-4B-Base Student
Student
18.14
13.16
16.14
15.81
ST-OPD
72.10
68.72
24.29
55.04
TK-OPD
74.70
68.72
24.00
55.81
TT-OPD
75.00
69.44
25.14
56.53
Appendix
Table 11: Per-dataset code-generation results (accuracy, %). Qwen3-8B-Base-GRPO-Code is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use the Code subset of Eurus-2-RL-Data as the training set. The student’s top- 16 tokens are used as the selected token set. Bold = best among the tested algorithms.
HumanEval+
MBPP+
LiveCodeBench
Avg.
Teacher
87.20
80.42
31.43
66.35
Qwen3-4B-Base Student
Student
48.78
41.27
28.00
39.35
ST-OPD
84.15
78.84
28.00
63.66
TK-OPD
81.71
77.51
28.00
62.41
TT-OPD
85.37
79.10
29.71
64.73
Appendix
Table 12: Per-dataset code-generation results (pass@4, %). All other settings are identical to those in Table 11 . Bold = best among the tested algorithms.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
31.15
24.69
22.71
70.63
88.02
38.10
58.91
47.74
Qwen3-4B-Base Student
Initial Student
8.02
5.31
6.15
30.80
54.07
22.79
25.22
21.77
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
TT-OPD w/o TC
21.56
20.63
19.58
57.27
72.70
33.92
49.30
39.28
TT-OPD
26.77
23.33
22.92
64.87
86.08
36.81
54.70
45.07
Appendix
Table 13: Per-dataset tail-correction ablation results (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base and Qwen3-1.7B-Base are the students, and all tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected token set. “w/o TC” removes only the tail-correction term, i.e., the second term in Eq. ( 15 ). Bold = best among the tested algorithms within each student block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
53.33
40.00
40.00
86.75
93.20
46.32
70.81
61.49
Qwen3-4B-Base Student
Initial Student
26.67
26.67
20.00
63.86
76.20
36.03
42.37
41.69
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
TT-OPD w/o TC
33.33
26.67
30.00
74.70
89.40
38.24
64.15
50.93
TT-OPD
56.67
43.33
43.33
80.72
92.20
37.87
64.89
59.86
Appendix
Table 14: Per-dataset tail-correction ablation results (pass@ k , %). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 13 . Bold = best among the tested algorithms within each student block.
On-Policy Distillation (OPD) has emerged as a dominant post-training paradigm for large language models, especially for reasoning domains. However, OPD remains unstable in practice due to the high gradient variance of its single-sample Monte Carlo estimator, and recipes for stable training are still immature. We propose vOPD (On-Policy Distillation with a control variate baseline), which casts OPD as policy-gradient RL and stabilizes it by introducing a control variate baseline-canonically a value function -- from the RL literature. We show that the OPD value function admits a closed form as the per-token negative reverse KL divergence between the student and the teacher, available directly from the already-computed forward pass with no additional critic or inference. Existing stabilization methods either compute the full token-level reverse KL over the entire vocabulary, adding significant overhead, or restrict it to a top-k support, biasing the objective. vOPD instead preserves the lightweight single-sample estimator, subtracting the value function as a detached baseline to keep the gradient unbiased while reducing variance. Furthermore, we show that a top-k approximation of the baseline further lowers cost without compromising performance. Across mathematical and scientific reasoning benchmarks, vOPD consistently outperforms vanilla OPD and matches the most expensive full-vocabulary baseline, offering an efficient stabilization of On-Policy Distillation through principled RL variance reduction.
Minjae Oh, Sangjun Song, Gyubin Choi +2
Graduate School of Data Science Seoul National University
Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.
Wenze Lin, Jiyuan Long, Jiale Zhao +11
LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University +3
On-Policy Distillation (OPD) is a fundamental technique for efficient post-training of large language models (LLMs), with broad applications in agent learning, multi-task enhancement, and model compression. However, OPD training becomes unstable when the teacher and student distributions differ substantially, as teacher supervision on student-generated tokens may yield unreliable policy gradients and even cause optimization failure. This work addresses reliable on-policy token-level supervision through credit assignment strategies, and proposes Trust Region On-Policy Distillation, TrOPD. It features the following characteristics: 1) Trust-Region On-Policy Learning: TrOPD performs OPD only in regions where the teacher provides reliable supervision, mitigating the optimization difficulty of the K1 reverse-KL estimator under distribution mismatch. 2) Outlier Estimation: For outlier regions, we explore gradient clipping, masking, and forward-KL estimation to reduce the adverse effects of unreliable supervision. 3) Off-Policy Guidance: The student continues generation from teacher prefixes and uses forward KL to imitate off-policy guidance, encouraging on-policy exploration toward reliable regions. Experiments show that TrOPD consistently outperforms SoTA OPD baselines, including OPD, EOPD, and REOPOLD, across mathematical reasoning, code generation, and general-domain benchmarks.
Xingrun Xing, Haoqing Wang, Boyan Gao +2
Samsung Research, Beijing, China · University of Oxford · Peking University