ReTaCo: Residual-Target Control for On-Policy Distillation
Organizations: Xi’an Jiaotong University · Zhejiang University · Georgia Institute of Technology · Boston College · Yale University · University of California, San Diego · Tsinghua University · Shanghai Innovation Institute · Massachusetts Institute of Technology
Abstract
On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top- probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top- mass toward one even after the student matches the teacher's relative probabilities within the top- set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top- tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top- mass , the residual target is for : preserves the teacher's mass, and larger moves more mass onto the top- tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top- mass lies between and and increases monotonically with ; at , underestimated top- tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.
Figures & tables
| Setting | Objective | MATH 500 | Olympiad Bench | AMC | AIME 24 | AIME 25 |
| Forward supervision and residual-target choice | ||||||
| Full ReTaCo | 83.78 | 49.65 | 52.56 | 28.75 | 21.67 | |
| No forward term | 82.40 | 48.67 | 52.11 | 25.42 | 22.08 | |
| Partial trimming ( ) | 78.88 | 43.61 | 47.74 | 24.17 | 21.67 | |
| Complete trimming ( ) | 78.60 | 43.78 | 48.95 | 23.75 | 21.67 | |
| Component variants under one-epoch training | ||||||
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
| Objective | Reverse information | Forward support | Forward mass |
| Sampled reverse KL | Full KL in expectation | – | – |
| Renormalized forward KL | – | in full softmax | |
| Residual-target forward KL | – | ||
| ReTaCo | Full KL in expectation |
| Quantity | |||
| Predicted | 0.6802847194 | 0.7642681841 | 0.8703486945 |
| Optimized | 0.6802847109 | 0.7642681719 | 0.8703486945 |
| Setting | Value | Setting | Value |
| Student | 0.6B | Teacher | 8B |
| Training examples | 144,490 | Optimizer steps | 2,257 |
| Learning rate | Rollout batch | 64 | |
| Samples / prompt | 8 | Response cap | 4,096 |
| 1 | 0 |
| Metric | Full-run mean | Last 100 | Final |
| Response length | 477.878 | 470.832 | 482.027 |
| Response truncation (%) | 0.461 | 0.441 | 1.172 |
| Reverse loss | 1.032 | 0.948 | 0.971 |
| Mass-preserving forward loss | 1.046 | 0.937 | 1.048 |
| Teacher top- mass | 0.999742 | 0.999713 | 0.999664 |
| Student top- mass | 0.993973 | 0.9993896 | 0.993832 |
| Benchmark | Items | Task / split |
| MATH500 | 500 | Mathematics test subset |
| OlympiadBench | 675 | English, text-only math |
| AMC | 83 | AMC12 2022–2023, integer answers |
| AIME24 | 30 | 2024 AIME I and II |
| AIME25 | 30 | 2025 AIME I and II |
| HumanEval+ | 164 | EvalPlus Python tasks |
| Setting | Value |
| Optimizer / learning rate | AdamW / |
| AdamW betas / epsilon | / |
| Weight decay | 0 |
| Schedule / warmup | Cosine / 3% |
| Precision / gradient clipping | BF16 / 1.0 |
| Prompt batch / mini-batch | 128 / 32 |
| Variant | Objective | Truncated (%) |
| Mass-only forward term | 47.51 | |
| Conditional-only forward term | 48.33 | |
| No reverse term | 49.98 |
| Setting | MATH500 | OlympiadBench | AMC | AIME24 | AIME25 |
| Forward weight | |||||
| 78.45 | 43.78 | 46.84 | 21.25 | 19.17 | |
| (default) | 83.78 | 49.65 | 52.56 | 28.75 | 21.67 |
| 70.45 | 44.15 | 47.59 | 25.00 | 22.50 | |
| Selected support | |||||
| 79.25 | 45.76 | 49.85 | 23.33 | 22.50 | |