On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-
k probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-
k mass toward one even after the student matches the teacher's relative probabilities within the top-
k set, so the teacher itself is not a stationary point whenever the omitted tokens have positive teacher probability. We propose ReTaCo (Residual-Target Control), which keeps the top-
k tokens individually and groups the remaining tokens into one residual symbol, and pairs this forward target with a single-sample estimator whose expectation equals the full-vocabulary reverse KL. With teacher top-
k mass
m, the residual target is
(1−β)(1−m) for
β∈[0,1]:
β=0 preserves the teacher's mass, and larger
β moves more mass onto the top-
k tokens without changing their relative probabilities. At a fixed prefix, we prove that the population objective has a unique optimum whose top-
k mass lies between
m and
m+β(1−m) and increases monotonically with
β; at
β=0, underestimated top-
k tokens still receive non-vanishing recovery gradients. Numerical optimization confirms these predictions, and across three teacher-student pairs, ReTaCo outperforms EOPD on most mathematics and code benchmarks.