cs.LGSep 29, 2026

On the Off-Policy Teacher in On-Policy Distillation

Authors: Langlin Huang, Hao Liu, Mononito Goswami, Xinyu Li, Prithwith Jana, Nikos Kanakaris, Patrick Blöbaum, Purak Jain

Organizations: Washington University in St. Louis · AWS AI Labs · Carnegie Mellon University · Georgia Institute of Technology

Abstract

On-policy distillation (OPD) has recently emerged as a promising post-training paradigm in which the student learns from trajectories generated by its own policy under dense teacher supervision. However, OPD introduces a fundamental asymmetry: although the sampled trajectories are on-policy for the student, they are off-policy for the teacher. The teacher is typically optimized to continue from prefixes generated by its own policy, but during OPD it must instead supervise prefixes generated by the student. Empirically, we find that its continuation performance degrades as these prefixes grow longer. To address this issue, we propose Student-COnditioned Updates of the Teacher (SCOUT), a co-training framework that adapts the teacher to student-generated prefixes. Alongside standard OPD updates, SCOUT periodically optimizes the teacher's conditional ability using reinforcement learning with verifiable rewards, where the teacher generates continuations from student prefixes and learns from outcome rewards. Controlled experiments show that SCOUT improves the teacher's ability to continue from student-generated prefixes, supporting the intended mechanism of student-conditioned teacher adaptation. Across multiple teacher--student configurations, model scales, and reasoning domains, SCOUT also consistently improves the effectiveness of on-policy distillation.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Your Teacher Can't Help You Here: Combating Supervision Fidelity Decay in On-Policy Distillation

    May 29, 2026Yanjiang Liu, Jie Lou, Xinyan Guan +7TeacherDataset Distillation

  2. Trust-Region Behavior Blending for On-Policy Distillation

    May 29, 2026Daniil Plyusov, Alexey Gorbatovski, Alexey Malakhov +4Trust RegionDistribution Matching Distillation

  3. Pass the Baton: Trajectory-Relayed On-Policy Distillation

    Jul 28, 2026Haolei Xu, Xiaowen Xu, Haiwen Hong +5TeacherMathematical Reasoning Benchmarks