Beyond Prompt Count: How Data Shapes Transfer in On-Policy Distillation
Authors: Jiaxuan Wang, Jiafei Lyu, Yuchen Cai, Siye Wu, Pengyuan Wang, Jiashun Liu, Xiang Cheng, Kai Yang, +3 more
Organizations: State Key Laboratory of Novel Software Technology, Nanjing University · School of Intelligence Science and Technology, Nanjing University · LLM Department, Tencent · RUC
On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-student pairs remains poorly understood. We systematically study prompt quantity, source, and selection across RL- and SFT-continuation pairs and cross-model settings. We find that OPD can be highly prompt-efficient: a few prompts can approach large-pool performance, with four DAPO prompts matching the observed mathematics score of 3,840 DeepMath prompts. However, prompt utility is relational rather than intrinsic: changing only the teacher can reverse the relative effectiveness of mathematics and code prompts. To characterize these transfer differences, we analyze parameter and functional changes across prompt supports and model pairs. Functional alignment with the teacher varies across supports and target tasks; in continuation pairs, teacher-aligned prediction changes can coexist with weak parameter alignment. Continued OPD on effective supports can restore performance after unfavorable transfer. Finally, targeted selection does not consistently outperform uniform random sampling, and filtering out a source that performs poorly alone yields no consistent gain across three paired support draws. Overall, our results distinguish prompt efficiency from prompt interchangeability and show that effective data choice depends on the teacher-student pair and target capability, with random sampling providing a competitive baseline in the studied settings.
Figures & tables
Figure 1: Prompt efficiency does not imply prompt interchangeability. A single prompt can induce substantial transfer, but outcomes vary with prompt identity and composition. Moreover, a few prompts from one source can match the observed mathematics score of thousands from another. These examples motivate our study of prompt count, source, teacher–student pairing, and selection. In (b), P1 – P3 denote the individual code prompts described in Appendix C.4 , while P1+P2 denotes joint OPD using the first two prompts.
Table 1: Teacher–student pairs and OPD sources. The first three teachers are trained by us; the others are publicly released models. We list constituent data sources; support construction, including mixtures and source filtering, is detailed in Appendix B.3 .
Figure 2: Prompt-count scaling across six teacher–student pairs. Stars mark full-pool OPD; S/T denote the initial student and teacher. Circled points in (d) attain the same observed Math score with 4 DAPO versus 3,840 DeepMath prompts. Training budgets are comparison-specific; full results appear in Appendix C .
Figure 3: Data-source preferences across teacher–student pairs. (a) Source sensitivity with a 30B Instruct teacher. (b) Preference reversals between Math-RL and Code-RL teachers with the same initial student. (c) Cross-domain transfer from the Nemotron teacher. Avg is the unweighted mean of the four displayed benchmark groups; side annotations report differences from DeepMath. Darker and lighter shading mark the highest and second-highest OPD scores within each column.
Table 3: Data-source preferences with a multi-domain RL teacher. Qwen3-4B-Inst-Mix → Qwen3-4B-Instruct-2507; M=48 . Teal marks the best and second-best OPD scores; mixture construction is detailed in Appendix B.3 .
Figure 4: Parameter and functional alignment. (a) Parameter and all-domain functional cosines for 43 checkpoints; each line joins the same endpoint. (b) Teacher-direction projection α versus functional cosine on Math probes, grouped by model pair. (c) Functional cosine versus accuracy: 30B-to-4B Math (left) and Mix-RL-to-4B GPQA (right). Definitions: Appendix E.1 .
Figure 5: GPQA accuracy and output-length changes. All supports contain M=48 prompts; changes are relative to each pair’s initial student. Inset: mean output tokens.
Figure 6: Redirecting unfavorable transfer. (a) CKA relative to full-pool OPD. (b) Bidirectional hidden-state replacement. (c) Performance recovery through continued distillation; vertical ticks mark native references. (d) Functional versus parameter alignment to native endpoints. Reference checkpoints and unequal Code update budgets are specified in Appendix E.5 .
Figure 7: Continuation data shapes recovery at matched update counts. 30B Instruct → Qwen3-4B, with 15 updates per stage and a fresh Adam optimizer in every second-stage run; token budgets are not matched. Orange and blue indicate stage-2 GSM8K and DeepMath, respectively. (a) AIME24/25 validation mean@2; solid/dashed lines start from GSM8K/DeepMath checkpoints. Gray step-0 markers show the shared starting scores of 44.17 / 50.83 . (b) Formal Avg 4 (%), with the y-axis truncated at 30% , grouped by stage-1 source; annotations show the gains from using DeepMath rather than GSM8K in stage 2.
Method
Math
GPQA-D
HE+
LCB
Avg 4
GPU ⋅ h
M=8,C=384
Uniform random
31.09 ± 0.50
39.44 ± 1.24
62.40 ± 0.32
17.67 ± 0.86
37.65 ± 0.38
0
Stratified random
30.36 ± 0.39
38.19 ± 0.58
61.74 ± 0.46
17.93 ± 0.41
37.06 ± 0.25
1.1
Semantic diversity
30.21
39.46
62.88
17.36
37.48
<0.1
Hard selection
28.96
38.07
63.57
18.21
37.20
1.1
Shortest
28.75
39.08
61.28
16.86
36.49
2.2
Table 4: Prompt selection within DeepMath. JustRL-1.5B → DS-1.5B. Random selectors report mean ± sample SD over three sampled supports. Bold marks the column best within each M . Costs estimate selection-only GPU hours; 0 denotes no GPU screening.
Figure 12
Figure 9: Source composition and a four-prompt control. 30B Instruct → Qwen3-4B, 25 updates. DeepMath/GSM8K composition sweep (circles, M=48 ) and DeepMath4-only control (diamonds, M=4 ). Within each seed, the control reuses exactly the four DeepMath prompts in the corresponding 4/44 mixture. Error bars and bands show sample SD across selection seeds 42/43/44.
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Parameter
Math / OOD
Code (4B)
Direct-code
Samples per question ( k )
16 / 8
8
16
Temperature
1.0
1.0
0.6
Top- p
1.0
1.0
0.95
Top- k / min- p
−1 / 0.0
−1 / 0.0
−1 / 0.0
Max. generation length
16,384
16,384
4,096
Appendix
Table 11: Evaluation configurations. Lengths are in tokens; all settings report mean@ k and pass@ k , with k=16 for Math and k=8 for OOD.
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
22.29
19.79
11.04
8.75
15.47
42.74
77.06
17.71
Teacher
–
62.71
56.46
34.38
39.79
48.33
52.65
83.92
20.07
DeepMath
1
55.21
54.17
32.29
37.29
44.74
51.70
82.39
29.93
DeepMath
48
60.21
53.12
33.33
38.96
46.41
53.22
83.23
27.21
DeepMath
3,840
61.25
55.83
32.92
37.08
46.77
52.15
85.14
29.79
Appendix
Table 12: Math-RL continuation. Math GRPO-500 to Qwen3-4B ( Figure 2 a ), with 3,840 rollouts per OPD run. Math reports mean@16; other tasks report mean@8.
Source / support
M
HE+
MBPP+
LCB
Code macro
Initial
–
76.60
65.15
18.14
53.30
Teacher
–
86.36
71.03
28.00
61.80
Eurus Code (ID 5038)
1
79.04
68.39
26.43
57.95
Eurus Code
48
86.66
70.14
28.57
61.79
Eurus Code
3,840
86.05
70.90
28.14
61.70
Eurus Code (full)
22,618
86.74
71.06
27.86
61.89
Appendix
Table 13: Code-RL continuation. Code GRPO-300 to Qwen3-4B ( Figure 2 b ), at OPD step 100; all scores are mean@8. The principal M=1 support is ID 5038 ( B in Section 5 ); rows below the separator give additional single-prompt experiments.
Source / reference
M
HE+
MBPP+
LCB
Code macro
Initial: direct-code SFT-554
–
44.32
52.88
11.25
36.15
Teacher: Code GRPO-400
–
58.12
58.47
15.96
44.18
Open-R1 Code
1
58.82
58.74
15.96
44.51
Open-R1 Code
256
59.98
59.51
15.57
45.02
Open-R1 Code
3,840
59.68
59.62
16.61
45.30
Open-R1 Code (full)
6,519
58.73
58.80
15.50
44.34
Appendix
Table 14: Direct-code RL continuation. Code GRPO-400 to Qwen3-1.7B-Base after direct-code SFT at step 554 ( Figure 2 c ). OPD endpoints are at step 50; all scores are mean@16.
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
27.50
25.63
12.50
11.25
19.22
35.92
54.34
13.50
Teacher
–
53.96
39.38
21.67
23.75
34.69
40.47
64.56
19.36
DAPO
1
42.71
30.42
16.88
17.08
26.77
38.13
60.98
16.29
DAPO
2
49.58
36.25
20.83
20.83
31.87
38.64
62.27
16.71
DAPO
4
47.71
36.04
22.08
24.38
32.55
41.10
63.87
17.50
DAPO
8
51.04
35.00
21.25
22.71
32.50
40.91
64.10
18.64
Appendix
Table 15: JustRL continuation. JustRL-DeepSeek-1.5B to DeepSeek-R1-Distill-Qwen-1.5B ( Figure 2 d ), at OPD step 100. Math reports mean@16; other tasks report mean@8. Full denotes the available source pool; DeepMath M=8 is the random support in the prompt-count sweep.
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
22.29
19.79
11.04
8.75
15.47
42.74
77.13
17.71
Teacher
–
75.21
61.67
42.50
56.25
58.91
54.86
82.39
33.14
DeepMath
1
46.0
46.7
24.6
31.9
37.29
48.3
82.4
24.3
DeepMath
48
56.46
45.63
29.38
39.79
42.81
52.15
82.32
26.79
DeepMath
3,840
53.8
48.3
30.2
38.1
42.60
52.0
83.7
27.0
DAPO
1
49.4
40.4
26.2
31.5
36.88
49.7
83.5
26.3
Appendix
Table 16: Cross-model distillation. Qwen3-30B-A3B-Instruct-2507 to Qwen3-4B ( Figure 2 e ), at OPD step 15. Math reports mean@16; other tasks report mean@8. Full-pool endpoints at other training durations appear in Table C.2 .
Source
M
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Initial
–
13.33
10.00
5.00
5.63
8.49
29.10
59.83
12.36
Teacher
–
75.21
61.67
42.50
56.25
58.91
54.86
82.39
33.14
DeepMath
1
27.3
24.0
16.0
14.2
20.4
30.1
65.5
16.1
DeepMath
8
33.1
26.7
15.0
15.2
22.5
32.5
65.8
19.4
DeepMath
48
35.0
26.9
14.2
17.1
23.3
36.2
66.7
18.1
DeepMath
3,840
36.0
30.8
14.8
16.7
24.6
32.7
68.9
18.7
Appendix
Table 17: Cross-model distillation. Qwen3-30B-A3B-Instruct-2507 to the official Qwen3-1.7B checkpoint ( Figure 2 f ), not the SFT-554 initialization in Table C.1 . Math reports mean@16; other tasks report mean@8.
Source
Step
AIME24
AIME25
HMMT Feb
HMMT Nov
Math
GPQA-D
HE+
LCB
Math GRPO-500 → Qwen3-4B
DeepMath
25
62.92
54.58
32.29
38.75
47.14
52.78
84.91
29.86
DeepMath
50
64.79
56.04
34.38
41.46
49.17
54.29
84.45
28.07
Qwen3-30B-A3B-Instruct-2507 → Qwen3-4B
DeepMath
25
55.0
48.8
27.3
37.9
42.2
48.4
82.8
25.8
DeepMath
50
56.9
49.0
31.7
38.8
44.1
48.4
83.5
26.6
Appendix
Table 18: Additional full-pool OPD endpoints. Math reports mean@16; other tasks report mean@8. Math-RL runs use batch size 1,024, with 25,600 and 51,200 rollouts at OPD steps 25 and 50, respectively.
Figure 10: Support size and learning dynamics. All runs use the Qwen3-4B Math-RL continuation pair. (a–b) Training-time AIME24 and AIME25 validation accuracy (mean@2). (c) Mean rollout response length, in thousands of tokens. (d) Teacher–student discrepancy, reported as logged mean K1 on a logarithmic scale. These validation traces are distinct from the final mean@16/mean@8 evaluations in the complete result tables.
Figure 11: Teacher and prompt-source effects during Qwen3-4B OPD. Math-T and Code-T denote Math GRPO-500 and Code GRPO-300; Math48 and Code48 denote the fixed DeepMath and Eurus Code supports. (a) AIME25 validation with Math48. (b) TACO validation with Code48. Both report training-log mean@2. (c) Mean rollout response length. (d) Teacher–student discrepancy, reported as logged mean K1 (log scale). Panels (a–b) use different validation tasks; the common-benchmark source-preference reversal is shown separately in Figure 3 b .
Math-RL teacher
Code-RL teacher
Support / training seed
DeepMath48
Code48
DeepMath48
Code48
42 / 42
54.30
46.84
44.94
50.75
43 / 42
54.20
46.62
44.79
51.27
44 / 42
54.23
46.17
45.22
50.51
42 / 43
53.87
46.37
44.69
50.31
42 / 44
54.26
46.66
44.67
50.42
Appendix
Table 19: Seed robustness of teacher-dependent source preferences. Avg 4 (%) across support-selection and training seeds.
Data
Tokens (M)
Math
GPQA
HE+
LCB
GSM8K
1.19
31.15
46.84
81.25
16.79
+ token match
16.14
31.56
48.04
82.32
17.43
DeepMath
16.63
42.81
52.15
82.32
26.79
Appendix
Table 20: Benchmark-level supervision-token control. M=48 ; tokens are cumulative supervised tokens in millions, and GSM8K token matching to DeepMath is approximate.
Figure 12: Additional source-dependent functional diagnostics. (b) 30B-to-4B. (c) Nemotron. (d) Mix-RL. Upper plots report Math functional cosine and teacher-direction coefficient α ; lower matrices report task-specific KL closure (%). Closure uses teacher-first conditional KL on the initial student’s frozen top-128 support, not a task-score recovery rate. Blue and red indicate negative and positive closure, respectively; colors saturate at −100% and 100% .
Model / OPD support
GPQA-D
HumanEval+
LCB v6
(1) Math-RL: Qwen3-4B → Qwen3-4B
Initial (Compact)
42.74 / 870
77.06 / 325
17.71 / 557
Teacher (GRPO-500)
52.65 / 2,619
83.92 / 700
20.07 / 2,670
DeepMath M48 (Compact)
53.22 / 5,590
83.23 / 2,328
27.21 / 9,567
(2) Qwen3-30B-A3B-Instruct-2507 → Qwen3-4B
Initial
42.74 / 870
77.13 / 325
17.71 / 557
Appendix
Table 21: Accuracy and output length for seven teacher–student pairs. Each cell gives mean@8 accuracy (%) / mean output tokens; s denotes the OPD update step.
Figure 13: Code recovery in parameter and function space. PCA views of the initial student ( I ), degraded single-prompt endpoint ( A′ ), recovered endpoint ( R ), and native Code-M48 endpoint ( N ). Native Code-M48 is trained directly from I ; recovery continues from A′ on an effective 48-prompt support. From A′ to R , the reported distance to N increases from 1.46 to 1.69 in parameter space but decreases from 1.03 to 0.29 in function space. Axis percentages report explained variance; arrows connect endpoints rather than tracing intermediate updates.
Figure 14: Training trajectories under single-prompt and two-prompt supports. All runs use the same Qwen3-4B initialization and Code GRPO-300 teacher. A′ uses prompt 15240; M2 jointly uses prompts 15240 and 187 from initialization. (a) Code accuracy: the unweighted mean of mean@8 over HumanEval+, MBPP+, and LCB-v6. Markers for M2 and prompt 187 alone show only their step-100 endpoints. (b) Functional cosine to the teacher’s change. (c) Normalized functional-update magnitude of A′ , ∥vOPD∥/∥vT∥ .
Figure 15: Parameter displacements after bidirectional support switching. Cosines compare FP32 endpoint-minus-original-initialization vectors for the four paths in Figure 7 . Every endpoint has 30 total updates; G denotes GSM8K and D denotes DeepMath. The G→D versus D→D comparison uses the step-30 reference, unlike the step-15 Math reference in Figure 6 (d) .
Selection seed
Mixed
Filtered
Shared prompts
DeepMath / DAPO / GSM8K
DeepMath / DAPO / GSM8K
42
17 / 15 / 16
21 / 27 / 0
32
43
13 / 15 / 20
22 / 26 / 0
28
44
20 / 13 / 15
31 / 17 / 0
33
Appendix
Table 22: Mixed-pool support composition and overlap. Shared counts prompt instances selected by both arms under the same random ordering, not shared sampled responses.
Benchmark
Selection seed
Mixed
Filtered
Δ
Math average
42 / 43 / 44
43.2 / 40.3 / 38.8
41.4 / 39.3 / 40.7
−1.8 / −0.9 / +1.9
Mean ± SD
40.7±2.3
40.5±1.0
−0.3±2.0
AIME24
42 / 43 / 44
55.8 / 54.2 / 49.2
56.5 / 51.3 / 49.8
+0.6 / −2.9 / +0.6
Mean ± SD
53.1±3.5
52.5±3.5
−0.6±2.0
AIME25
42 / 43 / 44
50.4 / 48.1 / 43.8
44.4 / 46.7 / 46.9
−6.0 / −1.5 / +3.1
Mean ± SD
47.4±3.4
46.0±1.4
−1.5±4.6
Appendix
Table 23: Math results for mixed-pool and source-filtered sampling. 30B Instruct → Qwen3-4B, M=48 , 25 optimization steps; selection seeds 42/43/44 share a fixed training seed. Scores are mean@16 (%); Math average is the unweighted mean of the four benchmarks. Δ denotes paired Filtered − Mixed differences; means, SDs, and differences use unrounded scores.
Benchmark
Selection seed
Mixed
Filtered
Δ
OOD average
42 / 43 / 44
55.2 / 54.9 / 53.9
55.2 / 53.7 / 54.6
0.0 / −1.2 / +0.7
Mean ± SD
54.7±0.7
54.5±0.8
−0.2±0.9
GPQA-D
42 / 43 / 44
51.5 / 50.7 / 50.0
51.0 / 50.4 / 51.2
−0.4 / −0.3 / +1.2
Mean ± SD
50.7±0.7
50.9±0.4
+0.2±0.9
HE+
42 / 43 / 44
87.0 / 87.0 / 86.1
88.1 / 85.3 / 86.1
+1.1 / −1.8 / +0.1
Mean ± SD
86.7±0.6
86.5±1.4
−0.2±1.4
Appendix
Table 24: OOD results for mixed-pool and source-filtered sampling. 30B Instruct → Qwen3-4B, M=48 , 25 optimization steps; selection seeds 42/43/44 share a fixed training seed. Scores are mean@8 (%); OOD average is the unweighted mean of GPQA-D, HE+, and LCB. Δ denotes paired Filtered − Mixed differences; means, SDs, and differences use unrounded scores.
Method
HE+
MBPP+
LCB
Code Macro
GPU ⋅ h
Uniform random
88.26 ± 0.32
71.33 ± 0.34
27.50 ± 0.43
62.36 ± 0.22
0
Semantic diversity
88.03
71.83
28.14
62.67
0.1
Hard selection
88.03
71.56
28.14
62.58
6.9
T–S disagreement
87.58
70.92
27.86
62.12
8.6
Cost-D-opt
88.42
71.17
27.64
62.41
48.6
Appendix
Table 25: Prompt selection within Eurus Code. Code-RL 4B → Qwen3-4B; M=48 , C=2,304 , 100 updates. Scores (%) use mean@8; Code Macro averages the three benchmarks. Uniform random reports mean ± sample SD across three support draws. Bold marks column-best scores. Costs cover selection only; 0 denotes no GPU screening.
Method
Math
GPQA-D
HE+
LCB
Avg 4
GPU ⋅ h
Uniform random
40.56 ± 1.07
49.87 ± 0.55
86.16 ± 0.52
25.74 ± 0.52
50.58 ± 0.61
0
Stratified random
41.79 ± 1.13
50.65 ± 1.03
86.81 ± 1.47
26.81 ± 0.36
51.52 ± 0.83
5.7
Semantic diversity
40.21
50.44
86.20
27.36
51.05
<0.1
Hard selection
40.00
49.87
87.65
26.64
51.04
5.7
Shortest
35.05
46.53
83.61
20.21
46.35
5.7
Appendix
Table 26: Prompt selection within DeepMath L6. 30B Instruct → Qwen3-4B; M=48 , C=2,304 , 25 updates. Scores (%) use mean@16 for Math and mean@8 otherwise; Avg 4 averages the four benchmarks. Random selectors report mean ± sample SD across three support draws. Bold marks column-best scores. Costs cover selection only; 0 denotes no GPU screening.
Method
Selection seed
E(S) (%)
Cosine
Avg4 (%)
Uniform random
42
29.7109
0.9670
37.5227
Uniform random
43
29.7557
0.9565
38.0760
Uniform random
44
23.4674
0.9722
37.3466
Stratified random
42
21.7985
0.9833
37.0561
Semantic diversity
42
23.5230
0.9721
37.4771
Hard selection
42
39.0314
0.9606
37.2021
Appendix
Table 27: Per-support gradient representativeness and transfer. JustRL–DeepMath, C=384 , M=8 . Each row is one support in Figure 8 , not a mean across selection seeds. E(S) and Avg4 are percentages; cosine is dimensionless.
Training cost
Performance (%)
Method
F/U
Generated
Loss
Time
Math
GPQA-D
HE+
LCB
Avg 4
DeepSeek 1.5B
Fresh-only
67/67
128.19
128.19
590.5
32.55
39.02
63.64
18.57
38.45
Fresh-only
20/20
39.13
39.13
178.8
31.09
38.83
63.80
18.36
38.02
Uniform token replay
17/68
32.78
57.64
205.1
32.92
39.65
62.73
19.21
38.63
Full-batch reuse
17/68
32.60
130.40
226.0
32.14
40.15
64.94
18.50
38.93
Appendix
Table 28: Performance and training cost with trajectory reuse. F/U denotes fresh rollout batches / optimizer updates; generated and loss tokens are in millions (including repeated exposures for loss), and time is in minutes. Math reports mean@16 and the other benchmarks mean@8; Avg 4 is their unweighted mean.
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local confidence or teacher-student agreement to weight, filter, or truncate the sampled trajectory. These signals do not directly determine whether the teacher can continue a student prefix to a correct answer, and trajectory-level interventions can conflate one rollout's unreliability with low expected training value of its prompt. We define prompt-level teacher continuation reliability R as the teacher's probability of reaching a correct answer from a student prefix, averaged over prefixes and trajectories induced by the current student. Oracle experiments show that high-R prompts yield larger OPD gains and that descending-R training outperforms random and ascending orders on a fixed prompt pool. Because estimating R requires many teacher continuations, we use the maximum ROUGE-5 F1 between one independent student rollout and verifier-correct same-prompt teacher trajectories. Across ten equal-frequency bins of this actual score, mean R rises monotonically, showing that the proxy separates coarse reliability levels. ReOrder-OPD sorts prompts by the proxy, then draws independent on-policy training trajectories for vanilla OPD. It improves every matched aggregate comparison across Qwen3 and Gemma4 mathematics settings and Qwen3 code settings. Gains in all six FiRe-OPD and ExOPD settings show that prompt ordering complements within-trajectory supervision.
Ximo Zhu, Ruiqi Liu, Rong Wang +8
Hello Group Inc. · Institute of Automation, CAS · School of Advanced Interdisciplinary Sciences, UCAS +1
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a confidently wrong teacher can induce a strong yet misleading update. Distributional proxies, such as entropy or teacher-student likelihood agreement, measure uncertainty or agreement but do not directly verify outcome correctness. We introduce Teacher-Gated On-Policy Distillation (TGOPD), built on the principle that teacher reliability should be verified at the prompt level before dense supervision is admitted. TGOPD estimates reliability from a small set of verifier-scored teacher probes and routes each prompt exclusively to dense OPD when the reliability check passes or to verifier-grounded GRPO otherwise. Across 4B and 35B students in mathematics, code, and instruction following, TGOPD outperforms Vanilla OPD in all six single-domain settings and achieves higher seven-benchmark averages at both scales under multi-domain training. By using otherwise-idle teacher capacity for reliability estimation, TGOPD also reduces teacher-side compute waste in asynchronous OPD, increasing teacher-node GPU utilization from 9.8% to 78.9% in the measured 4B single-domain run.
On-policy distillation (OPD) trains a student language model with dense feedback from a stronger teacher on student-generated trajectories. Yet standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R2-OPD), which uses outcome agreement and the magnitude of teacher--student disagreement to continuously reallocate teacher supervision. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R2-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of 3.5 and 2.4 percentage points for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of 1.6 percentage points over standard OPD. These results highlight outcome-guided supervision allocation as an effective way to translate dense teacher feedback into stronger student performance across model scales and task domains.
Haofeng Xu, Junwei Su, Lansong Diao +2
The University of Hong Kong · Alibaba Group · University of Science and Technology of China