Knowledge distillation transfers reasoning capabilities from large teachers to efficient students. However, token-level on-policy distillation (OPD) constrains student exploration and requires teacher token probabilities, precluding distillation from black-box teachers that provide only text outputs. We introduce On-policy Verbal Distillation (OVD), a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations. We analyze when ranking induced by verbal scores can guide distribution approximation: under a density-ratio calibration condition on acceptance probabilities and bounded teacher-replacement error, we bound the approximation error between the resulting mixed trajectory distribution and a teacher-preferred target. On Web Q&A, OVD achieves 41.09% average EM with teacher feedback at inference, exceeding the strongest evaluated baseline by 5.89 percentage points. On AMC23, OVD-FR improves accuracy over RLVR by 10.0 percentage points (52.5% to 62.5%) after 600 training steps on 128 problems. Further experiments suggest that retaining student-generated prefixes helps preserve exploration and mitigate trajectory-level entropy collapse. OVD also improves training efficiency: resampling selected suffixes rather than entire responses reduces mean per-step training time by 10.2% in the 128-problem setting. Project page: https://menik1126.github.io/ovd-project-page/.
Figures & tables
Figure 1: Two limitations of token-level distillation. (a,b) On Qwen2.5-7B, full-vocabulary FP32+BF16 logits scale linearly with context length: at L=8192 , they require 7.0 GB per trajectory versus 0.44 GB for the BF16 KV cache, and approximately 240 GB in aggregate for N=32 versus 13 GB for top- k ( k=8192 ). (c,d) On PopQA and HotpotQA, simulator validation EM generally increases over training under top-1/top-5 token-level distillation, whereas held-out live-search EM is much lower; top-5 underperforms top-1 on both datasets.
Figure 2: Initial-prefix selection and continuation EM on a fixed 128-problem subset of the DeepScaleR-Preview-Dataset, with 32 fixed student-generated prefixes per problem (4,096 total), reused across thresholds. The teacher scores only prefixes; those scoring at least τte receive student-generated continuations for EM evaluation. At τte=8 , 441 prefixes remain: EM is 61.22% for verbal-score selection, 49.06% for per-prompt random, and 17.36% for pooled random. Shading denotes 95% confidence intervals.
Figure 3: Student on-policy rollouts (left); teacher-guided rejection sampling and simulated search interaction for Web Q&A (right).
Figure 4: Verbal acceptance versus teacher-to-student log density ratio on 6,400 Math500 answers. Points show within-question-centered bin estimates; bars indicate 95% question-bootstrap intervals.
Table 5
Method
Steps
AIME24
AMC23
CARP
Middle
College
Gaokao-II
MATH
MAWPS
SAT-Math
Avg.
Initial Student
0
13.3
35.0
45.3
30.7
5.2
35.7
34.3
22.5
78.1
33.36
Training on One Fixed Problem
RLVR
300
6.7
50.0
51.3
42.6
33.0
50.0
66.0
91.3
81.3
52.46
RLVR
500
13.3
47.5
51.7
40.6
34.7
42.9
66.3
94.1
81.3
52.49
RLVR
600
10.0
50.0
51.7
39.6
34.5
50.0
66.9
94.1
81.3
53.13
RLVR
800
10.0
52.5
51.0
39.6
36.3
57.1
67.3
95.2
78.1
54.14
Table 3: Math accuracy after training on one fixed problem or 128 distinct problems. FR and HES use full-response and high-entropy-suffix resampling. Avg. is the macro average over the nine benchmarks; bold marks each method’s highest Avg. Benchmark selection and random-suffix results are given in Appendix A.5 .
Figure 5: Web Q&A performance under different training- and test-time rejection thresholds. Results are reported for each dataset and as a macro-average (Avg.).
Figure 8
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Verbal acceptance aτ(y∣q)=Pr[S(y,q)≥τ] versus teacher-to-student log density ratios for 6,400 base-student answers (200 Math500 questions × 32 samples). Left: raw; right: length-normalized. Points show acceptance rates in ten equal-count bins after within-question centering; bars indicate 95% question-bootstrap intervals. Teacher generation policy πE is a proxy for pT , not its definition.
Steps
SVAMP
ASDiv
MAWPS
TABMWP
Minerva
Olympiad
SAT-Math
MMLU
Gaokao
AIME24
Avg. 10
Avg. 23
Score scale 0–5, reject if score <4
300
89.90
92.28
95.98
81.80
28.68
30.52
75.00
52.02
30.77
10.00
58.69
56.51
500
90.30
91.96
96.22
81.50
28.31
28.89
71.88
51.16
32.97
16.67
58.99
56.95
600
89.90
91.74
96.13
81.70
30.15
29.63
78.12
52.05
30.77
16.67
59.69
57.78
800
89.50
92.01
96.66
82.80
27.57
28.59
81.25
51.86
36.26
13.33
59.98
57.55
Score scale 0–9 (baseline), reject if score <7
Appendix
Table 5: Effect of the teacher’s score range on one-problem training with full-response resampling. Scores are accuracy (%), evaluated on AMD GPUs. Avg. 10 and Avg. 23 average the ten displayed benchmarks and all 23 benchmarks, respectively.
Figure 9: Cumulative wall-clock training time under different resampling strategies. At 800 steps, OVD-HES and OVD-RS take 17.9% and 7.5% more time than OVD-FR with one fixed problem, but 7.4% and 30.5% less time with 128 problems. Panels use different vertical scales.
Figure 10: Average accuracy on Gaokao, SAT-Math, AIME24, Minerva, and MAWPS across training checkpoints. These five benchmarks were selected from the ten displayed in Table 5 by the largest gains of 0 – 5 over 0 – 9 at 800 steps; the same set is used for every curve and checkpoint. Rejection thresholds are 4, 7, and 16 for the 0 – 5 , 0 – 9 , and 0 – 20 scales, respectively. Complete results and scoring formats are reported in Table 5 .
Method
Steps
AIME24
AMC23
CARP
Middle
College
Gaokao-II
MATH
MAWPS
SAT-Math
Avg.
Training on 128 Distinct Problems
OVD -RS
300
16.7
47.5
52.8
44.6
35.8
42.9
68.2
93.0
81.3
53.62
OVD -RS
500
6.7
52.5
53.8
46.5
36.6
50.0
69.4
93.6
78.1
54.13
OVD -RS
600
16.7
60.0
54.4
44.6
38.0
50.0
70.7
95.8
71.9
55.78
OVD -RS
800
16.7
57.5
53.9
45.5
39.0
42.9
70.8
96.3
84.4
56.34
Appendix
Table 6: Random-suffix resampling on 128 training problems. Benchmarks and averaging follow Table 3 .
(a) Majority-vote accuracy
Steps
SVAMP
ASDiv
MAWPS
TABMWP
Minerva
Olympiad
SAT-Math
MMLU
Gaokao
AIME24
Avg.
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
OVD -FR
300
91.0
91.6
93.0
93.8
97.0
97.6
88.5
89.8
32.7
34.9
36.6
39.4
81.2
90.6
53.5
56.5
36.3
40.7
20.0
26.7
63.0
66.1
500
92.2
92.1
93.4
94.0
97.1
97.2
89.2
90.5
34.6
35.3
34.4
37.2
87.5
90.6
55.8
58.2
35.2
46.2
20.0
20.0
63.9
66.1
600
92.0
93.0
93.2
93.7
97.0
97.4
88.6
90.2
35.3
38.2
34.4
36.6
84.4
84.4
54.7
57.7
39.6
49.5
20.0
23.3
63.9
66.4
Appendix
Table 7: Sampling results for the three resampling strategies after training on the fixed 128-problem DeepScaleR subset, with (τtr,τte)=(7,0) . M@ k is majority-vote accuracy; tied votes count as incorrect. P@ k counts a question as correct if any of the k answers is correct. We sample 16 answers per question and use a fixed-seed random subset for k=8 . Temperature is 0.6 and top- p is 1.0.
Method
Steps
SVAMP
ASDiv
MAWPS
TABMWP
Minerva
Olympiad
SAT-Math
MMLU
Gaokao
AIME24
Avg.
Scaling RLVR (2 random examples)
Two-example
300
73.7
76.2
80.4
70.8
16.5
30.2
81.2
51.1
35.7
6.7
52.2
Two-example
500
77.1
80.0
86.0
75.17
16.5
27.7
81.2
51.3
42.9
13.3
55.1
Two-example
600
75.9
78.1
84.6
75.6
16.5
31.1
84.4
49.6
35.7
13.3
54.5
Two-example
800
81.1
80.0
87.6
73.7
15.8
29.6
81.2
49.0
35.7
13.3
54.7
OPD ∗ (reverse KL)
800
80.5
79.3
87.0
73.2
15.6
29.3
80.8
48.9
35.4
13.0
53.8
Appendix
Table 8: Math accuracy after training on two problems. All methods use Qwen2.5-Math-1.5B and are evaluated without a teacher. OPD ∗ denotes token-level distillation at 800 steps.
Figure 16
Figure 13: Valid search count across training steps. OVD (green curve) maintains higher valid-search rates throughout training compared to QueryLogitOPD (blue curve). The gap widens as training progresses, demonstrating that verbal step-level feedback helps the student policy learn more consistent search behaviors without requiring dense token-level logit matching at every query generation step.
Single-Hop QA
Multi-Hop QA
τte
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
Musique
Bamboogle
GAIA
Avg.
0
49.20
67.20
63.20
39.20
35.20
21.80
15.20
9.09
37.51
5
47.00
67.40
66.40
38.00
37.60
23.20
24.80
7.88
39.04
10
46.00
66.80
64.80
38.80
43.80
25.00
40.00
10.30
41.94
Appendix
Table 10: Effect of the test-time rejection threshold on Web Q&A exact match (%). The student is Qwen-2.5-3B-Base, the teacher is Qwen2.5-7B, and τtr=5 . Best results are in bold.
Score Decoding
τtr
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
Musique
Bamboogle
GAIA
Avg.
Greedy
5
47.00
67.40
66.40
38.00
37.60
23.20
24.80
7.88
39.04
Sampling
5
48.80
68.00
66.60
39.40
41.40
24.20
28.80
11.50
41.09
Sampling
10
25.60
39.40
43.20
22.20
15.80
10.80
12.80
12.72
22.82
Appendix
Table 11: Greedy versus sampled teacher scores on Web Q&A, with Qwen-2.5-3B-Base as student and Qwen2.5-7B as teacher. Greedy decoding takes the most likely score; sampling uses temperature 0.7 over the ten score tokens. We report exact match (%) at τte=5 . Best results are in bold.
Teacher
τte
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
Musique
Bamboogle
GAIA
Avg.
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
SFT teacher: Qwen2.5-7B-Base
7B
0
51.00
53.89
68.80
73.57
69.00
56.58
39.40
46.69
42.00
46.42
21.00
27.58
23.20
35.59
13.33
25.56
40.97
45.73
7B
5
50.00
52.89
67.80
72.57
69.80
57.38
39.00
46.29
43.70
48.12
27.90
34.48
32.00
44.39
12.12
24.69
42.79
47.60
7B
10
50.00
48.88
69.20
70.81
65.20
45.33
42.00
48.04
45.00
46.07
26.40
33.50
42.40
48.42
11.51
24.26
43.96
45.66
SFT teacher: Qwen2.5-14B-Base
Appendix
Table 12: Web Q&A results with 7B and 14B SFT teachers and a Qwen-2.5-3B-Base student. Training uses τtr=10 ; test thresholds are 0, 5, and 10. Scores are EM and F1 (%). Unavailable F1 scores are estimated from server logs; GAIA F1 is extrapolated from a fitted EM-to-F1 trend. Bold EM scores mark the best 7B-teacher setting for each benchmark and Avg.
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99% on 4B mathematics, while 8K full-parameter profiling shows 70.5% lower backward memory.
Yongliang Miao, Shuang Liu, Yanguang Liu +2
The Chinese University of Hong Kong, Shenzhen · Carnegie Mellon University · New Jersey Institute of Technology +1
On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates. Its {policy disagreement score} combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs. Across eight mathematical reasoning benchmarks, MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark. MAESTRO also reduces average training response length by 67.3% relative to standard OPD. The code is available at https://github.com/yhao-wang/MAESTRO.
Yuhao Wang, Ruiyang Ren, Yinan Zhang +3
Nanyang Technological University, Singapore · Baidu Inc.