cs.CLOct 8, 2026

When Do We Need On-Policy Distillation? Distilling on Offline Student Rollouts Is Often Better

Authors: Siyan Zhao, Yonggan Fu, Jindong Jiang, Shih-Yang Liu, Song Bian, Byung-Kwan Lee, Sharath Turuvekere Sreenivas, Wenliang Dai, +3 more

Organizations: NVIDIA

Abstract

On-policy distillation (OPD) has become increasingly popular for transferring teacher capabilities to student models. In this work, we ask a critical research question: Is on-policy sampling always beneficial for distilling arbitrary teacher-student pairs? We show that a simple alternative, Semi-OPD, which distills from offline rollouts generated by the initial student, can often outperform OPD in both accuracy and training efficiency. Across 17 teacher-student pairs ranging from 1.5B to 235B parameters, Semi-OPD outperforms OPD in 14 cases, with up to +13.6% accuracy and 11.4x training speedup. We further find that the choice between OPD and Semi-OPD depends on the alignment between the initial teacher and student, quantified by an output-token overlap ratio: OPD is beneficial only when the two are highly aligned with high overlap ratios. Our deeper investigation suggests that effective distillation requires on-policyness w.r.t. both the student and the teacher. For misaligned pairs, student rollouts can become increasingly off-policy w.r.t. the teacher as context length grows, weakening the distillation signal. In contrast, Semi-OPD is often more stable, as it distills on shorter contexts while covering full trajectories and exposing the student to more teacher-preferred tokens. Beyond proposing Semi-OPD as an efficient alternative, our work motivates the community to rethink when to use OPD and to study stronger OPD variants with meaningful teacher-student pairs.

Explore similar work

CardsList
  1. Less is More: Early Stopping Rollout for On-Policy Distillation

    May 26, 2026Zhou Ziheng, Jiaqi Li, Huacong Tang +2Language Model DistillationOn-Policy Distillation

  2. Rethinking On-Policy Distillation of Large Language Models II: One Training Example

    Sep 3, 2026Zixuan Fu, Bingxiang He, Yuxin Zuo +10Language Model DistillationOn-Policy Distillation

  3. Are Full Rollouts Necessary for On-Policy Distillation?

    May 29, 2026Yaocheng Zhang, Jiajun Chai, Yuqian Fu +7RL for Language Model ReasoningLanguage Model Distillation