cs.LGSep 27, 2026

Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models?

Authors: Wenze Lin, Jiyuan Long, Jiale Zhao, Shenzhi Wang, Xitai Jiang, Ce Luo, Rui Lan, Qianli Ma, +6 more

Organizations: LeapLab, Tsinghua University · Qiuzhen College, Tsinghua University · Beihang University · National University of Singapore · The Chinese University of Hong Kong · E Fund Management Co., Ltd.

Abstract

Since the advent of knowledge distillation, KL divergence has been the standard loss in distillation. Recently, on-policy distillation (OPD) has emerged as an efficient post-training paradigm for LLMs. As a distillation method, OPD naturally inherits KL divergence as its standard loss. However, in this work, we find that KL divergence may not be necessary for OPD. We show that simply preserving the update direction is sufficient for effective OPD. As long as the update direction is toward the teacher, OPD works. More precisely, it is not the direction of every token, but the direction of a small subset of tokens where the teacher and student disagree strongly. We first show that simply assigning a reward of (+1) to tokens where the teacher probability is higher than the student probability and (-1) where it is lower, which merely encourages updates toward the teacher, reproduces almost the same training mode as OPD with reverse KL. We further show that only the direction of a small subset of tokens with large teacher-student disagreement is critical, and training works as long as their update direction is toward the teacher, even if other tokens are pulled away from the teacher. And as an application of these findings, we introduce Consensus Multi-Teacher On-Policy Distillation (C-MOPD) to improve Multi-Teacher On-Policy Distillation (MOPD). Unlike MOPD, which routes each sample to a single teacher and may cause capability conflicts across domains, C-MOPD lets every sample be supervised by all teachers. Experiments show that C-MOPD consistently outperforms MOPD on both math and code benchmarks. Our code is available at https://github.com/LeapLabTHU/KL-Free-OPD.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Unbiased Top-kk Estimation for On-Policy Distillation

    Sep 28, 2026Linjian Meng, Siyuan Gan, YuHan Li +5Efficient On-Policy DistillationKullback-Leibler Divergence

  2. KL for a KL: On-Policy Distillation with Control Variate Baseline

    May 8, 2026Minjae Oh, Sangjun Song, Gyubin Choi +2Kullback-Leibler DivergencePost-Training

  3. Understanding Off- vs On-Policy Distillation: A Tale of Distinct Training Objectives

    Sep 29, 2026Qiwei Di, Xuheng Li, Kaixuan Ji +3Kullback-Leibler DivergenceFeedback