cs.LGOct 5, 2026

Flash-OPD: Fast On-Policy Distillation

Authors: Wei Chen, Junle Chen, Yitong Yang, Zhaoyang Xu, Jiaxin Lin, Yuxuan Liang, Xiaofang Zhou, Kai Wang, +1 more

Organizations: Tencent Hy · HKUST(GZ) · HKUST

Abstract

On-policy distillation (OPD) provides dense teacher supervision on student-generated trajectories, but generating and evaluating long rollouts incurs substantial training cost. Existing acceleration methods reduce this cost through open-loop rollout schedules or closed-loop horizon adaptation. However, supervision compatibility can vary substantially across trajectories, making a single rollout horizon difficult to match their heterogeneous reliable lengths: an overly short horizon may truncate useful supervision, while an overly long one wastes computation beyond reliable regions. Our key insight is that the trajectory-specific reliability boundary need not be predicted before generation. By viewing reliability as the first-passage of accumulated low teacher--student compatibility events, the boundary is inherently unknown before sampling, yet whether it has been reached can be determined exactly from the observed prefix. Building on this insight, we propose Flash-OPD, which shifts from rollout-horizon control to adaptive trajectory-level boundary verification. Flash-OPD interleaves cached generation with teacher verification and independently stops each trajectory according to its observed compatibility events. To reduce verification overhead, the recent event rate is used only to schedule the next verification point, while the actual stopping decision always relies on the exact cumulative count. This separation prevents estimation errors from causing premature termination while enabling efficient verification during generation. Extensive experiments across diverse datasets and teacher--student settings show that Flash-OPD achieves 2.2×2.2\times--7.5×7.5\times speedups over standard OPD while maintaining or improving accuracy.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Adaptive FastOPD: Progress-Aware Rollout Horizon Expansion for Efficient On-Policy Distillation

    Jul 31, 2026Qian Tan, Huaifei Liang, Xuanyu Zhu +2Efficient On-Policy DistillationAutoregressive Rollout

  2. Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts

    Jun 20, 2026Qingfei Zhao, Huan Song, Shuyu Tian +2Unsupervised On-Policy Self-DistillationRollout

  3. Are Full Rollouts Necessary for On-Policy Distillation?

    May 29, 2026Yaocheng Zhang, Jiajun Chai, Yuqian Fu +7Efficient On-Policy DistillationParallel Rollouts