On-policy distillation (OPD) learns from teacher feedback on student-generated responses and has shown promise in reducing forgetting relative to supervised fine-tuning (SFT). However, its benefits and fragility remain incompletely understood. We study sequential distillation from multiple teachers, where the student minimizes its average divergence from the teachers. Forward Kullback--Leibler (KL) divergence yields a weighted arithmetic mixture, while reverse KL yields a normalized weighted geometric aggregate. We develop algorithms that learn these targets under off-policy and on-policy feedback, respectively, establishing logarithmic regret bounds in the tabular setting and extending the analysis to function approximation. By analyzing these aggregation targets, we identify mechanisms that help explain both the benefits and fragility of OPD. Relative to forward KL, reverse KL can better retain a confident expert's preferences under uninformative feedback, but is more sensitive to teachers that assign very low probabilities to correct responses. Its token-level conditionals also reveal a dependence on continuation distributions that can favor incorrect prefixes over long horizons.
Figures & tables
Figure 1: Forward- and reverse-KL aggregation with expert weight α=0.9 . Panels (a,b) use N=10 responses, with each teacher distributing its remaining probability uniformly over incorrect responses. (a) The other teachers are uniform, and expert confidence q varies above 0.5 ; the targets cross at qc≈0.7617 . (b) The expert’s probability is fixed at q=0.99 , while the other teacher’s correct-response probability ε varies on a logarithmic axis. (c) For binary responses, the horizon H varies while all token probabilities remain fixed. The expert chooses the correct first token a with probability 0.99 ; subsequent tokens have probabilities (0.99,0.01) after a and (1/2,1/2) after b . The other teacher is uniform at every prefix.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Sensitivity to expert weight and response-space size with uninformative teachers. Panels (a–c) vary α∈{0.5,0.9,0.99} at N=10 . Panels (d–f), together with (b), compare N∈{2,10,100,1000} at α=0.9 . Each curve varies q from 1/N toward 1 . Dotted vertical lines mark the interior crossing when it exists.
Figure 3: Sensitivity to expert weight and confidence with misleading teachers. Panels (a–c) fix q=0.99 and vary α . Panels (d–f), together with (b), fix α=0.9 and compare q∈{0.6,0.9,0.99,0.999} . Horizontal ranges differ: log10ε∈[−10,−1] in (a), [−600,−1] in (c), and [−60,−1] otherwise. Moving left makes the other teacher more misleading.
Figure 4: Sensitivity of first-token probabilities to the horizon. Top row: vary α with r=0.99 and δ=0.01 . Middle row: vary r with α=0.9 and δ=0.01 . Bottom row: vary δ∈{0.01,0.1,0.3} with α=0.9 and r=0.99 . The baseline is repeated in (b), (f), and (g). The horizontal axis is logarithmic; curves connect values computed at integer horizons.
Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou) · The Hong Kong Polytechnic University · Eastern Institute of Technology, Ningbo +1