cs.CLSep 28, 2026

Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation

Authors: Yongliang Miao, Shuang Liu, Yanguang Liu, Yandong Bai, Mengnan Du

Organizations: The Chinese University of Hong Kong, Shenzhen · Carnegie Mellon University · New Jersey Institute of Technology · Kuaishou

Abstract

On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99% on 4B mathematics, while 8K full-parameter profiling shows 70.5% lower backward memory.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Dr. OPD: Learning What to Follow for Optimal On-Policy Distillation of Large Language Models

    Sep 29, 2026Zhenyu Wang, Tianze Wang, Linjun Zhang +1Efficient On-Policy DistillationMean

  2. Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement

    Aug 31, 2026Yi Ding, Ruqi ZhangOnline-Policy DistillationTeacher

  3. Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

    May 29, 2026Yanjiang Liu, Jie Lou, Xinyan Guan +7TeacherDataset Distillation