On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
Figures & tables
Figure 1: Forward- and reverse-KL aggregation with expert weight α=0.9 . Panels (a,b) use N=10 responses, with each teacher distributing its remaining probability uniformly over incorrect responses. (a) The other teachers are uniform, and expert confidence q varies above 0.5 ; the targets cross at qc≈0.7617 . (b) The expert’s probability is fixed at q=0.99 , while the other teacher’s correct-response probability ε varies on a logarithmic axis. (c) For binary responses, the horizon H varies while all token probabilities remain fixed. The expert chooses the correct first token a with probability 0.99 ; subsequent tokens have probabilities (0.99,0.01) after a and (1/2,1/2) after b . The other teacher is uniform at every prefix.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Sensitivity to expert weight and response-space size with uninformative teachers. Panels (a–c) vary α∈{0.5,0.9,0.99} at N=10 . Panels (d–f), together with (b), compare N∈{2,10,100,1000} at α=0.9 . Each curve varies q from 1/N toward 1 . Dotted vertical lines mark the interior crossing when it exists.
Figure 3: Sensitivity to expert weight and confidence with misleading teachers. Panels (a–c) fix q=0.99 and vary α . Panels (d–f), together with (b), fix α=0.9 and compare q∈{0.6,0.9,0.99,0.999} . Horizontal ranges differ: log10ε∈[−10,−1] in (a), [−600,−1] in (c), and [−60,−1] otherwise. Moving left makes the other teacher more misleading.
Figure 4: Sensitivity of first-token probabilities to the horizon. Top row: vary α with r=0.99 and δ=0.01 . Middle row: vary r with α=0.9 and δ=0.01 . Bottom row: vary δ∈{0.01,0.1,0.3} with α=0.9 and r=0.99 . The baseline is repeated in (b), (f), and (g). The horizontal axis is logarithmic; curves connect values computed at integer horizons.
On-policy distillation (OPD) provides dense token-level supervision by asking a teacher to score student-generated rollouts. However, when the student drifts into an unrecoverable prefix, the teacher may locally agree with the degraded state, producing low reverse KL but little corrective training signal. We identify this persistent regime as a low-KL agreement trap. Further analyses show that tokens during and after such traps produce less useful supervision signals. We propose KAT (KL Agreement Trap Termination), an online OPD termination rule that detects persistent low-KL agreement with a dynamic training-adaptive threshold. By filtering weak supervision from degenerate agreement, KAT improves avg@k accuracy by 2.66% and pass@k by 3.43% across four mathematical benchmarks, while reducing average rollout length by 59.73%.
Haoran Xin, Anhao Zhao, Ying Sun +3
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou) · The Hong Kong Polytechnic University · Eastern Institute of Technology, Ningbo +1
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision from teachers. In the standard KL objective of OPD, token-level losses are uniformly averaged, implying equal weights for all tokens. However, we discover that not all tokens are created equal: as student rollouts grow longer, they deviate further from the teacher's distribution, leading to degraded supervision quality at later positions. As a result, OPD using only the first 30% of tokens can perform comparably to using all tokens, whereas OPD using only the last 30% of tokens barely learns anything. In this work, we provide a principled understanding of this issue through the lens of constrained optimization. Based on these insights, we derive Importance-Weighted On-Policy Distillation (IW-OPD), in which the weight assigned to each token depends on the accumulated discrepancy between the student's and teacher's distributions, naturally upweighting earlier tokens and downweighting later ones with larger deviations. We show that IW-OPD converges significantly faster than OPD, with better learning efficiency, and achieves better final performance than standard OPD in both same-size and cross-scale settings, improving performance up to 6.9 points on AIME-2025.
Yan Xie, Sijie Zhu, Tiansheng Wen +2
Xidian University · Georgia Institute of Technology · Amazon AGI SF Lab
On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.
Langlin Huang, Hao Liu, Mononito Goswami +5
Washington University in St. Louis · AWS AI Labs · Carnegie Mellon University +1