Knowledge distillation transfers reasoning capabilities from large teachers to efficient students. However, token-level on-policy distillation (OPD) constrains student exploration and requires teacher token probabilities, precluding distillation from black-box teachers that provide only text outputs. We introduce On-policy Verbal Distillation (OVD), a framework that uses verbal scores from black-box teachers to rank student-generated sub-trajectories, retaining high-scoring ones and replacing low-scoring ones with teacher-generated continuations. We analyze when ranking induced by verbal scores can guide distribution approximation: under a density-ratio calibration condition on acceptance probabilities and bounded teacher-replacement error, we bound the approximation error between the resulting mixed trajectory distribution and a teacher-preferred target. On Web Q&A, OVD achieves 41.09% average EM with teacher feedback at inference, exceeding the strongest evaluated baseline by 5.89 percentage points. On AMC23, OVD-FR improves accuracy over RLVR by 10.0 percentage points (52.5% to 62.5%) after 600 training steps on 128 problems. Further experiments suggest that retaining student-generated prefixes helps preserve exploration and mitigate trajectory-level entropy collapse. OVD also improves training efficiency: resampling selected suffixes rather than entire responses reduces mean per-step training time by 10.2% in the 128-problem setting. Project page: https://menik1126.github.io/ovd-project-page/.
Figures & tables
Figure 1: Two limitations of token-level distillation. (a,b) On Qwen2.5-7B, full-vocabulary FP32+BF16 logits scale linearly with context length: at L=8192 , they require 7.0 GB per trajectory versus 0.44 GB for the BF16 KV cache, and approximately 240 GB in aggregate for N=32 versus 13 GB for top- k ( k=8192 ). (c,d) On PopQA and HotpotQA, simulator validation EM generally increases over training under top-1/top-5 token-level distillation, whereas held-out live-search EM is much lower; top-5 underperforms top-1 on both datasets.
Figure 2: Initial-prefix selection and continuation EM on a fixed 128-problem subset of the DeepScaleR-Preview-Dataset, with 32 fixed student-generated prefixes per problem (4,096 total), reused across thresholds. The teacher scores only prefixes; those scoring at least τte receive student-generated continuations for EM evaluation. At τte=8 , 441 prefixes remain: EM is 61.22% for verbal-score selection, 49.06% for per-prompt random, and 17.36% for pooled random. Shading denotes 95% confidence intervals.
Figure 3: Student on-policy rollouts (left); teacher-guided rejection sampling and simulated search interaction for Web Q&A (right).
Figure 4: Verbal acceptance versus teacher-to-student log density ratio on 6,400 Math500 answers. Points show within-question-centered bin estimates; bars indicate 95% question-bootstrap intervals.
Table 5
Method
Steps
AIME24
AMC23
CARP
Middle
College
Gaokao-II
MATH
MAWPS
SAT-Math
Avg.
Initial Student
0
13.3
35.0
45.3
30.7
5.2
35.7
34.3
22.5
78.1
33.36
Training on One Fixed Problem
RLVR
300
6.7
50.0
51.3
42.6
33.0
50.0
66.0
91.3
81.3
52.46
RLVR
500
13.3
47.5
51.7
40.6
34.7
42.9
66.3
94.1
81.3
52.49
RLVR
600
10.0
50.0
51.7
39.6
34.5
50.0
66.9
94.1
81.3
53.13
RLVR
800
10.0
52.5
51.0
39.6
36.3
57.1
67.3
95.2
78.1
54.14
Table 3: Math accuracy after training on one fixed problem or 128 distinct problems. FR and HES use full-response and high-entropy-suffix resampling. Avg. is the macro average over the nine benchmarks; bold marks each method’s highest Avg. Benchmark selection and random-suffix results are given in Appendix A.5 .
Figure 5: Web Q&A performance under different training- and test-time rejection thresholds. Results are reported for each dataset and as a macro-average (Avg.).
Figure 8
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Verbal acceptance aτ(y∣q)=Pr[S(y,q)≥τ] versus teacher-to-student log density ratios for 6,400 base-student answers (200 Math500 questions × 32 samples). Left: raw; right: length-normalized. Points show acceptance rates in ten equal-count bins after within-question centering; bars indicate 95% question-bootstrap intervals. Teacher generation policy πE is a proxy for pT , not its definition.
Steps
SVAMP
ASDiv
MAWPS
TABMWP
Minerva
Olympiad
SAT-Math
MMLU
Gaokao
AIME24
Avg. 10
Avg. 23
Score scale 0–5, reject if score <4
300
89.90
92.28
95.98
81.80
28.68
30.52
75.00
52.02
30.77
10.00
58.69
56.51
500
90.30
91.96
96.22
81.50
28.31
28.89
71.88
51.16
32.97
16.67
58.99
56.95
600
89.90
91.74
96.13
81.70
30.15
29.63
78.12
52.05
30.77
16.67
59.69
57.78
800
89.50
92.01
96.66
82.80
27.57
28.59
81.25
51.86
36.26
13.33
59.98
57.55
Score scale 0–9 (baseline), reject if score <7
Appendix
Table 5: Effect of the teacher’s score range on one-problem training with full-response resampling. Scores are accuracy (%), evaluated on AMD GPUs. Avg. 10 and Avg. 23 average the ten displayed benchmarks and all 23 benchmarks, respectively.
Figure 9: Cumulative wall-clock training time under different resampling strategies. At 800 steps, OVD-HES and OVD-RS take 17.9% and 7.5% more time than OVD-FR with one fixed problem, but 7.4% and 30.5% less time with 128 problems. Panels use different vertical scales.
Figure 10: Average accuracy on Gaokao, SAT-Math, AIME24, Minerva, and MAWPS across training checkpoints. These five benchmarks were selected from the ten displayed in Table 5 by the largest gains of 0 – 5 over 0 – 9 at 800 steps; the same set is used for every curve and checkpoint. Rejection thresholds are 4, 7, and 16 for the 0 – 5 , 0 – 9 , and 0 – 20 scales, respectively. Complete results and scoring formats are reported in Table 5 .
Method
Steps
AIME24
AMC23
CARP
Middle
College
Gaokao-II
MATH
MAWPS
SAT-Math
Avg.
Training on 128 Distinct Problems
OVD -RS
300
16.7
47.5
52.8
44.6
35.8
42.9
68.2
93.0
81.3
53.62
OVD -RS
500
6.7
52.5
53.8
46.5
36.6
50.0
69.4
93.6
78.1
54.13
OVD -RS
600
16.7
60.0
54.4
44.6
38.0
50.0
70.7
95.8
71.9
55.78
OVD -RS
800
16.7
57.5
53.9
45.5
39.0
42.9
70.8
96.3
84.4
56.34
Appendix
Table 6: Random-suffix resampling on 128 training problems. Benchmarks and averaging follow Table 3 .
(a) Majority-vote accuracy
Steps
SVAMP
ASDiv
MAWPS
TABMWP
Minerva
Olympiad
SAT-Math
MMLU
Gaokao
AIME24
Avg.
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
M@8
M@16
OVD -FR
300
91.0
91.6
93.0
93.8
97.0
97.6
88.5
89.8
32.7
34.9
36.6
39.4
81.2
90.6
53.5
56.5
36.3
40.7
20.0
26.7
63.0
66.1
500
92.2
92.1
93.4
94.0
97.1
97.2
89.2
90.5
34.6
35.3
34.4
37.2
87.5
90.6
55.8
58.2
35.2
46.2
20.0
20.0
63.9
66.1
600
92.0
93.0
93.2
93.7
97.0
97.4
88.6
90.2
35.3
38.2
34.4
36.6
84.4
84.4
54.7
57.7
39.6
49.5
20.0
23.3
63.9
66.4
Appendix
Table 7: Sampling results for the three resampling strategies after training on the fixed 128-problem DeepScaleR subset, with (τtr,τte)=(7,0) . M@ k is majority-vote accuracy; tied votes count as incorrect. P@ k counts a question as correct if any of the k answers is correct. We sample 16 answers per question and use a fixed-seed random subset for k=8 . Temperature is 0.6 and top- p is 1.0.
Method
Steps
SVAMP
ASDiv
MAWPS
TABMWP
Minerva
Olympiad
SAT-Math
MMLU
Gaokao
AIME24
Avg.
Scaling RLVR (2 random examples)
Two-example
300
73.7
76.2
80.4
70.8
16.5
30.2
81.2
51.1
35.7
6.7
52.2
Two-example
500
77.1
80.0
86.0
75.17
16.5
27.7
81.2
51.3
42.9
13.3
55.1
Two-example
600
75.9
78.1
84.6
75.6
16.5
31.1
84.4
49.6
35.7
13.3
54.5
Two-example
800
81.1
80.0
87.6
73.7
15.8
29.6
81.2
49.0
35.7
13.3
54.7
OPD ∗ (reverse KL)
800
80.5
79.3
87.0
73.2
15.6
29.3
80.8
48.9
35.4
13.0
53.8
Appendix
Table 8: Math accuracy after training on two problems. All methods use Qwen2.5-Math-1.5B and are evaluated without a teacher. OPD ∗ denotes token-level distillation at 800 steps.
Figure 16
Figure 13: Valid search count across training steps. OVD (green curve) maintains higher valid-search rates throughout training compared to QueryLogitOPD (blue curve). The gap widens as training progresses, demonstrating that verbal step-level feedback helps the student policy learn more consistent search behaviors without requiring dense token-level logit matching at every query generation step.
Single-Hop QA
Multi-Hop QA
τte
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
Musique
Bamboogle
GAIA
Avg.
0
49.20
67.20
63.20
39.20
35.20
21.80
15.20
9.09
37.51
5
47.00
67.40
66.40
38.00
37.60
23.20
24.80
7.88
39.04
10
46.00
66.80
64.80
38.80
43.80
25.00
40.00
10.30
41.94
Appendix
Table 10: Effect of the test-time rejection threshold on Web Q&A exact match (%). The student is Qwen-2.5-3B-Base, the teacher is Qwen2.5-7B, and τtr=5 . Best results are in bold.
Score Decoding
τtr
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
Musique
Bamboogle
GAIA
Avg.
Greedy
5
47.00
67.40
66.40
38.00
37.60
23.20
24.80
7.88
39.04
Sampling
5
48.80
68.00
66.60
39.40
41.40
24.20
28.80
11.50
41.09
Sampling
10
25.60
39.40
43.20
22.20
15.80
10.80
12.80
12.72
22.82
Appendix
Table 11: Greedy versus sampled teacher scores on Web Q&A, with Qwen-2.5-3B-Base as student and Qwen2.5-7B as teacher. Greedy decoding takes the most likely score; sampling uses temperature 0.7 over the ten score tokens. We report exact match (%) at τte=5 . Best results are in bold.
Teacher
τte
NQ
TriviaQA
PopQA
HotpotQA
2Wiki
Musique
Bamboogle
GAIA
Avg.
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
EM
F1
SFT teacher: Qwen2.5-7B-Base
7B
0
51.00
53.89
68.80
73.57
69.00
56.58
39.40
46.69
42.00
46.42
21.00
27.58
23.20
35.59
13.33
25.56
40.97
45.73
7B
5
50.00
52.89
67.80
72.57
69.80
57.38
39.00
46.29
43.70
48.12
27.90
34.48
32.00
44.39
12.12
24.69
42.79
47.60
7B
10
50.00
48.88
69.20
70.81
65.20
45.33
42.00
48.04
45.00
46.07
26.40
33.50
42.40
48.42
11.51
24.26
43.96
45.66
SFT teacher: Qwen2.5-14B-Base
Appendix
Table 12: Web Q&A results with 7B and 14B SFT teachers and a Qwen-2.5-3B-Base student. Training uses τtr=10 ; test thresholds are 0, 5, and 10. Scores are EM and F1 (%). Unavailable F1 scores are estimated from server logs; GAIA F1 is extrapolated from a fitted EM-to-F1 trend. Bold EM scores mark the best 7B-teacher setting for each benchmark and Avg.
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu