On-policy distillation (OPD) trains a student on its own generated responses using token-level teacher supervision. However, uniform weighting overlooks differences in token learning value, while existing weighting methods rely on predefined mappings from prediction signals to token weights. These mappings are not learned from the effectiveness of the resulting student updates, limiting their ability to adapt to evolving learning needs. In this paper, we propose MetaOPD, a bilevel optimization framework that jointly learns the student model and a lightweight token-weighting network. The inner objective updates the student through weighted OPD, while the outer objective optimizes the weighting network using validation loss on reference solutions after a virtual student update. Differentiating through this update connects weighting decisions to their effects on post-update performance, allowing the mapping from prediction signals to token weights to evolve alongside the student. Experiments on six mathematical reasoning and three out-of-domain datasets, covering two student scales and seven baselines, demonstrate the effectiveness of MetaOPD, with Avg@8/Pass@8 gains over OPD of 1.99/5.97 percentage points for the 0.6B student and 2.25/6.41 points for the 1.7B student.
Figures & tables
Figure 1: Motivating experiment on token weighting strategies. (a) Static proxy methods assign token weights using predefined signals, while MetaOPD learns weights through validation-loss feedback. (b) Validation loss reductions relative to Uniform OPD for MetaOPD, EOPD ( Jin et al., 2026 ) , TIP ( Xu et al., 2026b ) , and random token weighting. MetaOPD achieves the largest reduction, supporting the benefit of learning token weights aligned with the validation objective.
Figure 2: Overview of MetaOPD. (1) Standard OPD uses uniform token weights, while (2) static proxy weighting OPD assigns token weights through predefined rules based on proxy signals. (3) MetaOPD uses two branches: (a) the weighting network learns from the student’s validation loss after a virtual weighted OPD update; (b) the updated network recomputes weights for the actual student update. The schematic sums omit valid-token normalization, and weights in the actual-update branch refer to those recomputed after the weighting-network update.
Method
Avg@8
Pass@8
MATH500
AMC23
Minerva
HMMT
AIME24
AIME25
Avg.
MATH500
AMC23
Minerva
HMMT
AIME24
AIME25
Avg.
Teacher Qwen3-8B ⇝ Student Qwen3-0.6B-Base
KD
47.90
23.75
14.25
0.42
2.50
0.83
14.94
69.20
50.00
31.25
3.33
6.67
6.67
27.85
GRPO
53.20
28.44
16.31
0.83
4.17
1.25
17.37
75.00
57.50
32.35
6.67
6.67
10.00
31.36
OPD
50.28
25.00
15.95
0.42
2.92
1.25
15.97
73.00
55.00
31.62
3.33
10.00
6.67
29.94
EOPD
50.40
27.50
16.08
1.67
3.75
1.25
16.78
75.60
52.50
34.19
6.67
13.33
6.67
31.49
Table 1: Mathematical-reasoning results (%) with a Qwen3-8B teacher. Avg. is the unweighted average over six benchmarks; benchmark and Avg. scores are rounded separately. Abs. ΔOPD reports MetaOPD–OPD in percentage points. Rel. ΔOPD reports the same difference as a percentage of the OPD score. Delta values are calculated from the displayed scores. Bold marks the best result in each student setting, including ties.
Benchmark
Metric
Base Model
KD
GRPO
OPD
EOPD
FiRe-OPD
REOPOLD
TIP
MetaOPD
GPQA-Diamond
Avg@8
15.09
21.09
19.26
26.20
25.44
25.82
26.52
27.15
27.27
Pass@8
69.19
75.25
77.27
77.78
76.26
76.77
76.26
74.75
78.28
MMLU-Pro
Pass@1
7.28
38.13
41.86
41.57
40.93
41.51
39.52
41.90
42.30
AlpacaEval 2.0
LC-WR
1.21
14.72
13.12
15.96
15.19
15.14
16.06
16.70
15.98
WR
2.03
16.09
14.63
17.97
17.17
17.13
17.37
19.14
18.08
All metrics
Avg.
18.96
33.06
33.23
35.90
35.00
35.27
35.15
35.93
36.38
Table 2: Out-of-domain results for Qwen3-1.7B-Base after MATH training, with Qwen3-8B as the teacher for distillation methods. Base Model is the student before training. Scores are percentages; higher is better. GPQA-Diamond reports Avg@8/Pass@8, MMLU-Pro reports Pass@1, and AlpacaEval 2.0 reports length-controlled win rate (LC-WR) and win rate (WR). Avg. is the unweighted average of these five metric scores. Bold and underline mark the best and second-best results, respectively; the MetaOPD column is shaded.
Variant
Avg@8
Pass@8
Uniform OPD
25.07
40.41
Position only
25.21
39.32
Teacher entropy only
26.75
43.25
Log-probability gap only
25.42
43.62
Top-32 profiles only
25.70
42.24
MetaOPD (all features)
27.32
46.82
Table 3: Feature-ablation Avg. scores (%).
Figure 3: Effects of adding SFT to OPD and MetaOPD. Six-benchmark mean Avg@8/Pass@8 changes (percentage points) for Qwen3-1.7B-Base. (a) Adding the SFT loss to OPD at half, equal, and double the gradient-matched reference coefficient γ⋆≈2.16 . (b) Adding the SFT loss at γ⋆ to OPD or MetaOPD. Changes are relative to each method without the added SFT loss.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Feature group
Features
Dim.
Detached
Sampled action
Current-student log-probability, teacher log-probability, and teacher-minus-current log-probability gap
3
Yes
Uncertainty
Vocabulary-normalized student and teacher entropy
2
Yes
Top-1 behavior
Student and teacher top-1 probability; student–teacher top-1 identity match
3
Yes
Position and overlap
Position normalized by the common completion-window width; student–teacher top-32 set overlap
2
Yes
Ranked probabilities
Student top-32 and teacher top-32 full-vocabulary-normalized log-probability values
64
Yes
Total
Ten scalar statistics plus two ranked top-32 vectors
74
Yes
Appendix
Table 4: Composition of the 74-dimensional weighting network descriptor, with kf=32 .
Setting
Value
Student / teacher models
Qwen3-0.6B-Base or Qwen3-1.7B-Base / frozen Qwen3-8B
Training data / MetaOPD reference-solution data
MATH train / MATH train
MetaOPD rollout / reference sampling
Separate shuffled streams; overlap allowed; no meta split
Shared epochs / training seed
3 / 42
Shared rollout and update schedule
174 rollout iterations; four student updates per rollout; 696 student updates
Shared batch sizes
128 rollout / 32 student optimizer; global reference batch 32 where used
Appendix
Table 5: Shared training and evaluation configuration, including MetaOPD-specific settings.
Method
Configuration
FiRe-OPD
Trajectory drop fraction 0.2; teacher-confidence and student-confusion coefficients both 1.0
Soft-OR; response-level keep fraction 0.5; entropy clipping at the 98th percentile; exact-KL computation in 64-token chunks
Appendix
Table 6: Method-specific allocation settings within the shared sampled-OPD framework. Selection scopes and loss normalization are specified in the text.
Variant
Token weighting
Reference use
Student objective
Uniform OPD
Uniform
None
LOPD
Position only
Position
Outer
LOPDw
Teacher entropy
Teacher entropy
Outer
LOPDw
Log-probability gap
δi,t
Outer
LOPDw
Top-32 profiles
Student + teacher profiles
Outer
LOPDw
MetaOPD
All 74 descriptors
Outer
LOPDw
Appendix
Table 7: Ablation-variant mapping. “Outer” means that reference solutions update only the weighting network through the meta objective; “direct” means that their supervised loss enters the student objective.
Variant
Avg@8
Pass@8
MATH500
AMC23
Minerva
HMMT
AIME24
AIME25
Avg.
MATH500
AMC23
Minerva
HMMT
AIME24
AIME25
Avg.
Uniform OPD
67.23
38.75
27.80
1.67
8.33
6.67
25.07
84.20
67.50
47.43
6.67
20.00
16.67
40.41
Position only
68.50
39.69
26.01
1.67
9.58
5.83
25.21
88.00
67.50
43.75
3.33
20.00
13.33
39.32
Teacher entropy only
68.88
43.75
27.85
2.50
11.25
6.25
26.75
86.60
72.50
47.06
10.00
26.67
16.67
43.25
Log-probability gap only
68.50
39.38
26.75
1.67
8.75
7.50
25.42
87.80
75.00
45.59
6.67
26.67
20.00
43.62
Top-32 profiles only
69.05
40.94
27.57
1.67
8.33
6.67
25.70
88.40
75.00
46.69
3.33
20.00
20.00
42.24
Appendix
Table 8: Feature-composition ablation on Qwen3-1.7B-Base (%).
Variant
Avg@8
Pass@8
MATH500
AMC23
Minerva
HMMT
AIME24
AIME25
Avg.
MATH500
AMC23
Minerva
HMMT
AIME24
AIME25
Avg.
Uniform OPD
67.23
38.75
27.80
1.67
8.33
6.67
25.07
84.20
67.50
47.43
6.67
20.00
16.67
40.41
SFT
32.85
14.06
11.76
1.25
0.83
0.00
10.13
65.40
52.50
34.93
6.67
6.67
0.00
27.69
OPD + SFT ( 0.5γ⋆ )
67.45
38.44
23.99
2.50
7.92
6.67
24.49
86.60
75.00
43.75
10.00
16.67
16.67
41.45
OPD + SFT ( γ⋆ )
63.00
33.75
23.53
2.08
7.50
6.25
22.69
86.20
77.50
45.59
6.67
23.33
13.33
42.10
OPD + SFT ( 2γ⋆ )
57.25
27.81
22.75
1.25
8.33
5.42
20.47
86.20
67.50
47.79
10.00
20.00
20.00
41.92
Appendix
Table 9: SFT and combined OPD–SFT controls on Qwen3-1.7B-Base (%).
Method
Training time (h)
GPU-hours
Relative cost
OPD
5.62
45.0
1.00×
MetaOPD
11.06
88.5
1.97×
Appendix
Table 10: Training-only cost for the 1.7B student on eight NVIDIA A100 80GB GPUs, including four student and four teacher GPUs. Both methods complete 696 student updates; MetaOPD performs one meta update per step. Evaluation and diagnostic probes are excluded.
Figure 4: Training dynamics and update quality (1.7B). (a) Token-weight standard deviation over training. (b) Median within-response Spearman correlation between pre-update token weights and ci,t , the scaled negative validation-loss derivative, in epoch 1 of a matched diagnostic run; shading shows the interquartile range. (c) Paired validation loss difference ΔLLearned−Uniform after one update from identical starting parameters at each of three checkpoints; negative values favor Learned. Error bars show 95% reference-bank bootstrap intervals with checkpoints and rollouts fixed. Appendix D gives the full protocol.
(a) Agreement across independent reference batches
Table 11: Loss-derivative agreement and learned-weight correlations for MetaOPD at three training stages (1.7B). Agreement compares ci,t∘ computed from 32-example reference batches; weight correlations use ci,t∘ obtained after averaging d over 256 reference examples, weighted by valid reference-token counts, and centering within each response. Values summarize within-response comparisons. Brackets report 95% intervals from 2,000 prompt-cluster bootstrap resamples.
Checkpoint (step)
(a) Validation loss changes
(b) Paired contrasts
Uniform
Learned
Reweighted
Learned − Uniform
95% CI
Reweighted − Learned
Early (232)
+0.74
+0.09
−3.32
−0.65
[−0.72,−0.59]
−3.41
Middle (464)
+2.78
+0.80
−3.03
−1.97
[−2.16,−1.80]
−3.83
Final (696)
+6.26
+4.31
+0.61
−1.95
[−2.13,−1.78]
−3.70
Appendix
Table 12: One-step validation loss changes ΔLhstep and paired contrasts ( 103 scale; lower is better) at reference learning rate 3×10−6 . Learned − Uniform intervals use 10,000 reference-bank bootstrap resamples with checkpoint and rollouts fixed. Reweighted uses the evaluation bank to construct weights.
Input restriction
Spearman
Sign agreement
Top-20% Jaccard
103ΔLReweighted−Learned
Position only
0.978
0.994
0.962
-3.56
Teacher entropy
0.978
0.994
0.961
-3.92
Log-prob. gap
0.979
0.993
0.959
-3.84
Top-32 profiles
0.975
0.992
0.960
-3.30
Appendix
Table 13: Loss-derivative agreement and update comparisons for restricted-input variants (1.7B), all at student step 696. Agreement uses ci,t∘ computed from 32-example reference batches; ΔLReweighted−Learned uses equally weighted eight-example reference banks at reference rate 3×10−6 . Point estimates; negative contrasts favor Reweighted.
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by 9.7 points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals. Its sampled-token variant avoids the cost of full-vocabulary probabilities. Yet we find that training on fewer tokens can outperform full-token OPD, challenging the intuition that more supervision improves learning. This motivates selecting tokens by learning value. Existing disagreement-based criteria ignore probability scale: tokens assigned negligible probability by both models, termed low-low tokens, can receive large log-ratio rewards and hinder learning. We propose DIAL-OPD, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. A parameter beta controls this weighting, and the highest-scoring tokens are retained. Across 4 teacher-student pairs and 7 mathematical reasoning benchmarks, we compare DIAL-OPD with 9 baselines. Retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants, with mean accuracy gains reaching 5.25 percentage points over Vanilla OPD, and doubles AIME25 Pass@16 from 13.33% to 26.67%. It also achieves up to an 18% relative improvement in mean accuracy over the strongest token-selection baseline at matched retention ratios. With a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, showing that effective supervision allocation can outweigh teacher scaling. Further analysis shows that moderate beta balances suppressing low-low tokens against preserving useful disagreements. Token-level evidence reveals that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision critical to reasoning correctness.
Anhao Zhao, Haoran Xin, Junlong Tong +5
EIT-NLP Lab, Eastern Institute of Technology, Ningbo · The Hong Kong Polytechnic University · HKUST (GZ) +3
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision from teachers. In the standard KL objective of OPD, token-level losses are uniformly averaged, implying equal weights for all tokens. However, we discover that not all tokens are created equal: as student rollouts grow longer, they deviate further from the teacher's distribution, leading to degraded supervision quality at later positions. As a result, OPD using only the first 30% of tokens can perform comparably to using all tokens, whereas OPD using only the last 30% of tokens barely learns anything. In this work, we provide a principled understanding of this issue through the lens of constrained optimization. Based on these insights, we derive Importance-Weighted On-Policy Distillation (IW-OPD), in which the weight assigned to each token depends on the accumulated discrepancy between the student's and teacher's distributions, naturally upweighting earlier tokens and downweighting later ones with larger deviations. We show that IW-OPD converges significantly faster than OPD, with better learning efficiency, and achieves better final performance than standard OPD in both same-size and cross-scale settings, improving performance up to 6.9 points on AIME-2025.
Yan Xie, Sijie Zhu, Tiansheng Wen +2
Xidian University · Georgia Institute of Technology · Amazon AGI SF Lab