cs.LGSep 29, 2026

Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

Authors: Qiwei Di, Xuheng Li, Kaixuan Ji, Chenggong Zhang, Heyang Zhao, Quanquan Gu

Organizations: Department of Computer Science, University of California, Los Angeles, CA 90095, USA

Abstract

On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Escaping the KL Agreement Trap in On-Policy Distillation

    Jun 8, 2026Haoran Xin, Anhao Zhao, Ying Sun +3Kullback-Leibler Divergence

  2. On the Position Bias of On-Policy Distillation

    Jun 21, 2026Yan Xie, Sijie Zhu, Tiansheng Wen +2Efficient On-Policy DistillationToken-Level Entropy

  3. On the Off-Policy Teacher in On-Policy Distillation

    Sep 29, 2026Langlin Huang, Hao Liu, Mononito Goswami +5Efficient On-Policy DistillationTeacher