cs.LGSep 28, 2026

DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents

Authors: Hanyang Wang, Zeyuan Liu, Zhengyu Chen, Jingqing Ruan, Chaoxu Pang, Zhongda Su, Wulin Xie, Zhizhao Zeng, +2 more

Organizations: University of Chicago · Meituan LongCat Interaction Team · University of the Chinese Academy of Sciences · The Hong Kong University of Science and Technology (Guangzhou)

Abstract

On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, wasting already-generated experience. To address this problem, we introduce DivOPD, a simple learner-side batch-selection method that spreads a fixed turn budget across more rollouts and, within each rollout, prioritizes turns with larger cumulative teacher-student disagreement. Turns without usable teacher feedback are excluded. The per-turn loss and optimizer remain fixed; selection only changes which student-visited turns receive training weight. For no-progress rollouts, an optional extension briefly hands control to the teacher before returning it to the student. Across six teacher-student settings on the simulated ALFWorld, ScienceWorld, and WebShop benchmarks, with 1.5B-7B students, DivOPD raises cross-setting mean peak success rate from 77.4 to 84.4 and mean success over the last five evaluations from 71.5 to 78.6. It reaches all reported setting-specific targets with geometric-mean speedups of 1.84x in training tokens and 1.87x in learner GPU time relative to vanilla OPD. Teacher intervention further raises this last-five mean to 82.4 while retaining about 1.7x learner-GPU speedup over vanilla OPD. Code will be released at https://github.com/HanyangWang0418-oss/DivOPD.

Figures & tables

Appendix figures & tables18 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation

    Date pendingYuchen Xia, Qianguo Sun, Chao Song +3Efficient On-Policy DistillationUni-Opd

  2. TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

    Apr 27, 2026Jiaqi Wang, Wenhao Zhang, Weijie Shi +2Efficient On-Policy DistillationCurriculum

  3. UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents

    Sep 27, 2026Wenbo Zhang, Pengcheng Xu, Weizhi Du +2Efficient On-Policy DistillationRp-Opsd