cs.LGSep 27, 2026

UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents

Authors: Wenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang, Hengrui Cai

Organizations: University of California, Irvine · University of Michigan, Ann Arbor

Abstract

On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to 15.8%15.8\% relative to standard OPD.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation

    Date pendingYuchen Xia, Qianguo Sun, Chao Song +3Efficient On-Policy DistillationUni-Opd

  2. On-Policy Distillation with Curriculum Turn-level Guidance for Multi-turn Agents

    Jun 14, 2026Gengsheng Li, Mao Zheng, Mingyang Song +8TeacherCurriculum