Organizations: Shanghai Artificial Intelligence Laboratory, Shanghai, China · State Key Laboratory of Novel Software Technology, Nanjing University, Nanjing, China
On-policy distillation (OPD) is becoming an important component of large language model (LLM) post-training for transferring the reasoning capability of a strong teacher LLM to a weaker student LLM. OPD trains the student by minimizing the reverse KL divergence between the teacher and the student via rollouts generated by the student's policy. However, estimating the gradient of the reverse KL divergence in OPD remains a challenge. Using only the sampled token from the student-generated rollout is computationally cheap but provides limited distributional supervision, which will degrade accuracy. In addition, using the full vocabulary provides complete distributional supervision but is computationally expensive. Therefore, recent works propose Top-k OPD (TK-OPD) that use selected top-k tokens, which provides richer distributional supervision than sampled-token estimation at substantially lower computational cost than full-vocabulary estimation. Unfortunately, using only the selected top-k tokens induces bias, leading to accuracy degradation, as the probability mass outside the selected top-k tokens is discarded. To address the bias of TK-OPD, we propose Tail-Corrected Top-k On-Policy Distillation (TT-OPD). It preserves the advantages of TK-OPD, including rich distributional supervision and low computational cost, while providing an unbiased estimator of the gradient of the reverse KL divergence. The key insight of TT-OPD is to use not only the selected top-k tokens, but also the sampled token from the student-generated rollout, thereby recovering the discarded probability mass in expectation, avoiding the bias. Experimental results demonstrate that TT-OPD significantly outperforms other tested OPD variants.
Figures & tables
Figure 1: Overview of TT-OPD. The notation sg(⋅) denotes the stop-gradient operator. At each student-visited state Zt , the selected top- k tokens Stk provide rich distributional supervision, while the sampled token yt in the student-generated rollout provides an unbiased gradient tail correction. This only induces O(k+1) computational cost.
Unbiased Gradient?
Rich Distributional Supervision?
Computational Cost
ST-OPD
✓
×
O(1)
FV-OPD
✓
✓
O(∣V∣)
TK-OPD
×
✓
O(k)
TT-OPD
✓
✓
O(k+1)
Table 1: Comparison between our TT-OPD and other OPD variants. The notation V denotes the vocabulary.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
31.15
24.69
22.71
70.63
88.02
38.10
58.91
47.74
Qwen3-4B-Base Student
Student
8.02
5.31
6.15
30.80
54.07
22.79
25.22
21.77
ST-OPD
21.04
19.17
14.27
57.83
84.67
34.33
53.44
40.68
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
TT-OPD
26.77
23.33
22.92
64.87
86.08
36.81
54.70
45.07
Table 2: Main results on mathematical reasoning (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher for all tested algorithms. All tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected top- k tokens. Bold = best among the tested algorithms within each student block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
ST-OPD
21.04
19.17
14.27
57.83
84.67
34.33
53.44
40.68
k=8
TK-OPD
16.46
17.50
14.58
55.72
84.90
34.15
53.50
39.54
TT-OPD
27.60
22.60
21.88
62.88
85.70
36.35
55.96
44.71
k=16
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
Table 3: Sensitivity to the number of selected tokens (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected top- k tokens. ST-OPD does not depend on k . Bold = best among the tested algorithms for each value of k .
k
TK-OPD
TT-OPD
Qwen3-4B-Base Student
k=8
4 h 22 min
4 h 34 min
k=16
4 h 27 min
4 h 38 min
k=32
4 h 35 min
4 h 45 min
k=64
4 h 38 min
4 h 53 min
Qwen3-1.7B-Base Student
Table 4: Training time of TK-OPD and TT-OPD. Qwen3-8B-Base-GRPO-Math is the teacher. Except for k range in {8,16,32,64} , all other configurations are identical to those described in Section 5.1 .
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Student Top- k
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
TT-OPD
26.77
23.33
22.92
64.87
86.08
36.81
54.70
45.07
Teacher Top- k
TK-OPD
21.56
17.71
14.90
57.83
84.75
35.66
53.76
40.88
TT-OPD
27.08
25.83
20.21
64.83
86.58
35.66
54.70
44.98
Table 5: Results with different selected top- k tokens (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use DAPO-Math-17K as the training set. We set k=16 . Bold = best among the tested algorithms within each top- k block.
Figure 2: Code-generation results (average accuracy, %). Qwen3-8B-Base-GRPO-Code is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use the Code subset of Eurus-2-RL-Data as the training set. The student’s top- 16 tokens are used as the selected top- k tokens.
Figure 3: Tail-correction ablation (average accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base and Qwen3-1.7B-Base are the students, and all tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected top- k tokens.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
53.33
40.00
40.00
86.75
93.20
46.32
70.81
61.49
Qwen3-4B-Base Student
Student
26.67
26.67
20.00
63.86
76.20
36.03
42.37
41.69
ST-OPD
40.00
40.00
33.33
79.52
90.40
42.65
63.70
55.66
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
TT-OPD
56.67
43.33
43.33
80.72
92.20
37.87
64.89
59.86
Appendix
Table 7: Pass@ k results on mathematical reasoning (%). Unlike average accuracy, which averages correctness over all sampled responses, pass@ k is the percentage of problems for which at least one of the k sampled responses is correct. We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 2 . Bold = best among the tested algorithms within each student block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
ST-OPD
40.00
40.00
33.33
79.52
90.40
42.65
63.70
55.66
k=8
TK-OPD
40.00
23.33
26.67
73.49
90.60
41.18
62.22
51.07
TT-OPD
56.67
36.67
43.33
81.93
92.40
37.50
57.78
58.04
k=16
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
Appendix
Table 8: Pass@ k results for different numbers of selected tokens (%). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 3 . ST-OPD does not depend on k . Bold = best between TK-OPD and TT-OPD within each k block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Student Top- k
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
TT-OPD
56.67
43.33
43.33
80.72
92.20
37.87
64.89
59.86
Teacher Top- k
TK-OPD
46.67
30.00
30.00
72.29
88.20
37.87
62.37
52.49
TT-OPD
43.33
33.33
43.33
81.93
92.40
44.85
65.78
57.85
Appendix
Table 9: Pass@ k results with different selected top- k tokens (%). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 6 . Bold = best among the tested algorithms within each top- k block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
86.67
80.00
80.00
100.00
99.00
47.06
82.37
82.16
DeepSeek-R1-Distill-Qwen-7B Student
Student
83.33
73.33
76.67
96.39
98.00
44.85
82.52
79.30
ST-OPD
83.33
76.67
80.00
97.59
98.80
47.06
82.22
80.81
TK-OPD
80.00
70.00
76.67
98.80
98.60
47.06
83.26
79.20
TT-OPD
83.33
80.00
80.00
100.00
98.40
48.16
83.85
81.96
Appendix
Table 10: Pass@ k results on another model series (%). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 6 . Bold = best among the tested algorithms.
HumanEval+
MBPP+
LiveCodeBench
Avg.
Teacher
77.44
71.96
27.43
58.94
Qwen3-4B-Base Student
Student
18.14
13.16
16.14
15.81
ST-OPD
72.10
68.72
24.29
55.04
TK-OPD
74.70
68.72
24.00
55.81
TT-OPD
75.00
69.44
25.14
56.53
Appendix
Table 11: Per-dataset code-generation results (accuracy, %). Qwen3-8B-Base-GRPO-Code is the teacher, Qwen3-4B-Base is the student, and all tested algorithms use the Code subset of Eurus-2-RL-Data as the training set. The student’s top- 16 tokens are used as the selected token set. Bold = best among the tested algorithms.
HumanEval+
MBPP+
LiveCodeBench
Avg.
Teacher
87.20
80.42
31.43
66.35
Qwen3-4B-Base Student
Student
48.78
41.27
28.00
39.35
ST-OPD
84.15
78.84
28.00
63.66
TK-OPD
81.71
77.51
28.00
62.41
TT-OPD
85.37
79.10
29.71
64.73
Appendix
Table 12: Per-dataset code-generation results (pass@4, %). All other settings are identical to those in Table 11 . Bold = best among the tested algorithms.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
31.15
24.69
22.71
70.63
88.02
38.10
58.91
47.74
Qwen3-4B-Base Student
Initial Student
8.02
5.31
6.15
30.80
54.07
22.79
25.22
21.77
TK-OPD
19.79
19.17
14.27
59.04
85.23
35.29
52.31
40.73
TT-OPD w/o TC
21.56
20.63
19.58
57.27
72.70
33.92
49.30
39.28
TT-OPD
26.77
23.33
22.92
64.87
86.08
36.81
54.70
45.07
Appendix
Table 13: Per-dataset tail-correction ablation results (accuracy, %). Qwen3-8B-Base-GRPO-Math is the teacher, Qwen3-4B-Base and Qwen3-1.7B-Base are the students, and all tested algorithms use DAPO-Math-17K as the training set. The student’s top- 16 tokens are used as the selected token set. “w/o TC” removes only the tail-correction term, i.e., the second term in Eq. ( 15 ). Bold = best among the tested algorithms within each student block.
AIME24
AIME25
AIME26
AMC
MATH
Minerva
Olympiad
Avg.
Teacher
53.33
40.00
40.00
86.75
93.20
46.32
70.81
61.49
Qwen3-4B-Base Student
Initial Student
26.67
26.67
20.00
63.86
76.20
36.03
42.37
41.69
TK-OPD
43.33
26.67
26.67
71.08
88.60
41.91
65.63
51.98
TT-OPD w/o TC
33.33
26.67
30.00
74.70
89.40
38.24
64.15
50.93
TT-OPD
56.67
43.33
43.33
80.72
92.20
37.87
64.89
59.86
Appendix
Table 14: Per-dataset tail-correction ablation results (pass@ k , %). We report pass@32 on AIME24, AIME25, AIME26, and AMC, and pass@8 on MATH, Minerva, and Olympiad. All other settings are identical to those in Table 13 . Bold = best among the tested algorithms within each student block.