cs.LGSep 28, 2026

Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents

Authors: Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao, Xuchen Li, Honglin Lin, Tao Cheng, Zhihan Yu, +3 more

Organizations: Qwen Large Model Application Team, Alibaba · Peking University · University of Chinese Academy of Sciences · Shanghai Jiao Tong University

Abstract

On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Look Ahead Before You Distill: Future Trajectory Validation of Teacher Guidance for Agentic On-Policy Distillation

    Aug 3, 2026Chishui Chen, Yaoyou Fan, Te Sun +11Efficient On-Policy DistillationTeacher

  2. On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

    Jun 14, 2026Gengsheng Li, Mao Zheng, Mingyang Song +8TeacherCurriculum

  3. TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents

    Apr 27, 2026Jiaqi Wang, Wenhao Zhang, Weijie Shi +2Efficient On-Policy DistillationCurriculum