cs.LGJun 14, 2026

On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

Authors: Gengsheng LiMao ZhengMingyang SongRuiqi LiuTianyu YangJie SunQiyong ZhongHaiyun Guo+3 more

Organizations: 1Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · 3Large Language Model Department, Tencent · University of Science and Technology of China · 5Zhejiang University · 6National University of Singapore · 7Wuhan AI Research

Abstract

Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large models whose inference cost is prohibitive in practice. On-Policy Distillation (OPD) is a natural recipe for transferring such capabilities to smaller students, but we find that it suffers a characteristic failure mode in this setting: small student errors compound across turns and push the trajectory out of the teacher's familiar state distribution, so the teacher's supervision becomes least reliable precisely where the student needs it most. We propose Guided On-Policy Distillation (Guided-OPD), a simple yet effective algorithm that mixes teacher- and student-generated turns within each rollout and schedules the teacher's intervention probability along a curriculum that decays to zero. Strong guidance keeps early trajectories close to the teacher distribution and is then gradually withdrawn to recover the purely on-policy regime used at inference. On ALFWorld, ScienceWorld, and WebShop, distilling Qwen3 students from a Qwen3-30B-A3B teacher, Guided-OPD yields average relative gains of 21.1% in Score and 25.5% in Success Rate over vanilla OPD, with larger gains on smaller students.

Explore similar work

CardsList