cs.LGSep 30, 2026

ReTaCo: Residual-Target Control for On-Policy Distillation

Authors: Zixiang Ni, Zhuo Hu, Renjie Cao, Weijie Ren, Binqin Shi, Weijia Zhang, Shuheng Cao, Zhicheng Shi, +3 more

Organizations: Xi’an Jiaotong University · Zhejiang University · Georgia Institute of Technology · Boston College · Yale University · University of California, San Diego · Tsinghua University · Shanghai Innovation Institute · Massachusetts Institute of Technology

Abstract

On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-kk probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-kk mass toward one even after the student matches the teacher's relative probabilities within the top-kk set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top-kk tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top-kk mass mm, the residual target is (1−β)(1−m)(1-β)(1-m) for β∈[0,1]β\in[0,1]: β=0β=0 preserves the teacher's mass, and larger ββ moves more mass onto the top-kk tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top-kk mass lies between mm and m+β(1−m)m+β(1-m) and increases monotonically with ββ; at β=0β=0, underestimated top-kk tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

    Sep 29, 2026Qiwei Di, Xuheng Li, Kaixuan Ji +3Kullback-Leibler DivergenceFeedback

  2. Unbiased Top-kk Estimation for On-Policy Distillation

    Sep 28, 2026Linjian Meng, Siyuan Gan, YuHan Li +5Efficient On-Policy DistillationKullback-Leibler Divergence

  3. Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

    Sep 29, 2026Zhenyu Wang, Tianze Wang, Linjun Zhang +1Efficient On-Policy DistillationMean