Organizations: Xi’an Jiaotong University · Zhejiang University · Georgia Institute of Technology · Boston College · Yale University · University of California, San Diego · Tsinghua University · Shanghai Innovation Institute · Massachusetts Institute of Technology
On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-k probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-k mass toward one even after the student matches the teacher's relative probabilities within the top-k set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top-k tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top-k mass m, the residual target is (1−β)(1−m) for β∈[0,1]: β=0 preserves the teacher's mass, and larger β moves more mass onto the top-k tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top-k mass lies between m and m+β(1−m) and increases monotonically with β; at β=0, underestimated top-k tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.
Figures & tables
Figure 1 : Residual probability as an explicit forward target. The teacher’s top- k probabilities carry selected mass m , leaving residual mass 1−m . EOPD renormalizes the selected probabilities to sum to one, assigning zero target mass to the residual. ReTaCo instead sets the residual target to (1−β)(1−m) and transfers β(1−m) to the selected tokens. Both targets preserve the teacher’s relative probabilities within the selected set. Bar widths are schematic.
Setting
Objective
MATH 500
Olympiad Bench
AMC
AIME 24
AIME 25
Forward supervision and residual-target choice
Full ReTaCo
R+B0+mC
83.78
49.65
52.56
28.75
21.67
No forward term
R
82.40
48.67
52.11
25.42
22.08
Partial trimming ( β=0.5 )
R+B0.5+r0.5C
78.88
43.61
47.74
24.17
21.67
Complete trimming ( β=1 )
R+B1+C
78.60
43.78
48.95
23.75
21.67
Component variants under one-epoch training
Table 2 : Qwen3-1.7B target and component comparisons (avg@8, %). Let R denote sampled reverse KL, Bβ=dBer(rβ∥P) , and C=DKL(qS∥pS) . The shaded row is the default, R+B0+mC , and bold marks the best score in each column. All variants train for one epoch from the same initial student.
Figure 2 : Selected and residual mass during training. Qwen3-1.7B student, Qwen3-30B-A3B-Instruct-2507 teacher, and k=16 ; EOPD and ReTaCo share the displayed interval of steps 1–202. (a) Student mass P on the teacher’s top- k set for EOPD (teal, dashed) and ReTaCo (wine, solid). Gray dotted curves show teacher mass m on each method’s own rollouts; m is also the ReTaCo forward target at β=0 . The black dashed line marks EOPD’s forward target r=1 . (b) The same data expressed as tail mass 1−P and 1−m , with a logarithmic vertical axis. Curves use exponential smoothing with span 11; faint traces in (a) show raw values.
Figure 3 : Residual-target control in analysis and experiments. (a) Selected-mass equilibria from Equation 22 at α=1 ; the dashed line denotes teacher mass. (b) Changes in Qwen3-1.7B avg@8 relative to β=0 , computed from the residual-target rows in Table 2 . Positive values indicate higher scores with trimming. All variants train for one epoch.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Objective
Reverse information
Forward support
Forward mass
Sampled reverse KL
Full KL in expectation
–
–
Renormalized forward KL
–
k in full softmax
P→1
Residual-target forward KL
–
k+1
P→rβ
ReTaCo
Full KL in expectation
k+1
P→rβ
Appendix
Table 3 : Support and mass control across objectives. The final column gives the mass preferred by the forward term alone. In the combined ReTaCo objective, reverse KL also influences the equilibrium.
Quantity
β=0
β=0.5
β=1
Predicted P∗
0.6802847194
0.7642681841
0.8703486945
Optimized P
0.6802847109
0.7642681719
0.8703486945
Appendix
Table 4 : Predicted and optimized selected mass. The teacher selected mass is m=0.6802847194 .
Setting
Value
Setting
Value
Student
0.6B
Teacher
8B
Training examples
144,490
Optimizer steps
2,257
Learning rate
3×10−6
Rollout batch
64
Samples / prompt
8
Response cap
4,096
α
1
β
0
Appendix
Table 5 : Training configuration for the β=0 response-length experiment.
Metric
Full-run mean
Last 100
Final
Response length
477.878
470.832
482.027
Response truncation (%)
0.461
0.441
1.172
Reverse loss
1.032
0.948
0.971
Mass-preserving forward loss
1.046
0.937
1.048
Teacher top- k mass
0.999742
0.999713
0.999664
Student top- k mass
0.993973
0.9993896
0.993832
Appendix
Table 6 : Response-length behavior at β=0 . Values are training diagnostics; response truncation denotes the fraction of responses reaching the generation limit.
Benchmark
Items
Task / split
MATH500
500
Mathematics test subset
OlympiadBench
675
English, text-only math
AMC
83
AMC12 2022–2023, integer answers
AIME24
30
2024 AIME I and II
AIME25
30
2025 AIME I and II
HumanEval+
164
EvalPlus Python tasks
Appendix
Table 7 : Evaluation benchmarks. HumanEval+ and MBPP+ use the full extended test suites.
Figure 4 : Policy entropy during training. Qwen3-1.7B student, Qwen3-30B-A3B-Instruct-2507 teacher, and k=16 . The curves show actor/entropy in nats for Sampled OPD (blue, dashed) and ReTaCo with β=0 (wine, solid), over the common interval of steps 1–277. Solid and dashed curves use exponential smoothing with span 11; faint traces show raw values. The last 20 steps average 0.305 and 0.399 nats, respectively.
Figure 5 : Mass and conditional matching during training. A single Qwen3-1.7B mathematics training run with k=16 , α=1 , and β=0 . (a) Mean B0=dBer(m∥P) and mC=mDKL(qS∥pS) from Equation 16 . (b) Mean ∣P−m∣ in percentage points. Faint lines show raw values; dark lines show centered 11-step moving averages, not confidence intervals. Steps denote rollout batches. This diagnostic uses a 7,168-token response cap and a 300-step budget, distinct from the main benchmark protocol.
Setting
Value
Optimizer / learning rate
AdamW / 3×10−6
AdamW betas / epsilon
(0.9,0.999) / 10−8
Weight decay
0
Schedule / warmup
Cosine / 3%
Precision / gradient clipping
BF16 / 1.0
Prompt batch / mini-batch
128 / 32
Appendix
Table 8 : Mathematics training settings shared by Qwen3-1.7B, Qwen3.5-2B, and Gemma4-E2B. H(q) is full-vocabulary teacher entropy.
Variant
Objective
Truncated (%)
Mass-only forward term
R+B0
47.51
Conditional-only forward term
R+mC
48.33
No reverse term
B0+mC
49.98
Appendix
Table 9 : Response truncation in the component ablations. Each variant trains for one epoch. Truncation denotes the fraction of responses reaching the 8,192-token generation limit.
Setting
MATH500
OlympiadBench
AMC
AIME24
AIME25
Forward weight
α=0.5
78.45
43.78
46.84
21.25
19.17
α=1 (default)
83.78
49.65
52.56
28.75
21.67
α=2
70.45
44.15
47.59
25.00
22.50
Selected support
k=8
79.25
45.76
49.85
23.33
22.50
Appendix
Table 10 : Qwen3-1.7B hyperparameter sensitivity (avg@8, %). Every setting trains for one epoch. Defaults are k=16 , α=1 , and β=0 ; each group varies the indicated parameter.