On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to 15.8% relative to standard OPD.
Figures & tables
Figure 1 : OPD versus UOPD. OPD executes student actions ( atS ) given observations ( ot ) and distills using reverse KL. UOPD uses teacher signals for uncertainty quantification (UQ) and an adaptive threshold to determine when the teacher intervenes. This decision jointly determines the executed action and the learning objective: reverse KL for student actions, or SFT on teacher actions ( atT ). The latter minimizes forward KL in expectation.
ALFWorld
WebShop
Seen
Unseen
Test
Method
SR ↑
Turns ↓
SR ↑
Turns ↓
Score ↑
SR ↑
Turns ↓
RL-Qwen2.5-7B teacher → Qwen2.5-3B student
Student (zero-shot)
21.0 ± 0.8
43.8 ± 0.3
15.9 ± 3.4
46.0 ± 1.1
12.2 ± 2.1
1.0 ± 0.5
13.6 ± 0.3
Teacher (zero-shot)
97.4 ± 0.4
9.1 ± 0.1
92.8 ± 0.4
12.3 ± 0.4
85.4 ± 0.6
77.9 ± 1.2
6.6 ± 0.3
OPD
88.1 ± 1.8
14.2 ± 0.3
85.3 ± 1.6
16.3 ± 0.6
77.5 ± 2.2
66.9 ± 3.0
7.4 ± 0.3
Table 1 : Main results on ALFWorld and WebShop. We report success rate (SR, %) on ALFWorld, and mean task score, SR (%) on WebShop. We also include average number of turns. Results are means ± standard deviations over three evaluation seeds. Bold indicates the best result among distillation methods.
ALFWorld
WebShop
Configuration
Seen SR ↑
Unseen SR ↑
Score ↑
SR ↑
Turns ↓
UOPD ( β=1 , 30%→5% )
90.0
88.8
82.4
69.3
6.5
Intervention signal
Confidence gap
89.8
87.8
77.3
69.0
7.2
Random intervention
88.3
83.8
77.3
64.6
7.4
SFT weight
Table 2 : Effects of intervention selection, SFT weight, teacher execution, and schedule with a 3B student. The default UOPD configuration uses teacher confidence, β=1 , teacher execution, and a 30%→5% intervention schedule. Random intervention samples each turn with the scheduled probability. Values are means over three evaluation seeds, and bold marks the best value in each column.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Condition
Task success (%)
Trajectories with state returns (%)
Student
13.98 ( 13/93 )
93.55 ( 87/93 )
Random
16.13 ( 15/93 )
91.40 ( 85/93 )
Intervention
25.81 ( 24/93 )
80.65 ( 75/93 )
Appendix
Table 3: Outcomes on the same 93 ALFWorld tasks used in Fig. 2 (c,d). Random and Intervention each execute one teacher action before returning control to the student.
Jun 14, 2026·Gengsheng Li, Mao Zheng, Mingyang Song +8TeacherCurriculum
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Large Language Model Department, Tencent +4