Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents
Authors: Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao, Xuchen Li, Honglin Lin, Tao Cheng, Zhihan Yu, +3 more
Organizations: Qwen Large Model Application Team, Alibaba · Peking University · University of Chinese Academy of Sciences · Shanghai Jiao Tong University
On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.
Figures & tables
Figure 1: Illustration of the supervision–benefit mismatch. From the same state and interaction history, one branch takes the student’s original action and the other a teacher-proposed action. Both branches then continue interacting under the same frozen student policy. Moving toward the teacher’s preferred action does not necessarily improve the current student’s final task outcome.
Figure 2: Empirical analysis of final task outcomes. Bars show decision proportions by outcome category within each benchmark’s gap groups. Error bars indicate 95% decision-level bootstrap confidence intervals.
Figure 3: Overview of Outcome-Guided On-Policy Distillation. Trajectory-relative weighting assigns supervision weights, while gap magnitude and abrupt increases identify candidate turns. Outcome-based calibration compares paired student continuations and strengthens supervision on the student’s original trajectory when teacher guidance improves the final task outcome.
ALFWorld
ScienceWorld
WebShop
Method
Seen SR ↑
Unseen SR ↑
Hard SR ↑
Overall SR ↑
Rounds ↓
SR ↑
Score ↑
Rounds ↓
SR ↑
Score ↑
Rounds ↓
Qwen3-32B teacher → Qwen3-1.7B student
Student (zero-shot)
7.1
8.2
0.0
5.3
19.8
0.2
3.8
17.0
27.0
24.2
8.1
Teacher (zero-shot)
37.1
35.1
9.9
28.1
18.5
28.2
53.9
12.0
48.0
47.3
4.9
Vanilla OPD
26.7 ± 3.0
27.6 ± 3.3
4.1 ± 1.4
20.1 ± 2.1
17.7 ± 0.4
13.0 ± 1.4
48.1 ± 1.2
12.6 ± 0.3
28.0 ± 2.6
27.5 ± 2.6
6.8 ± 0.2
TCOD-F2B
23.8 ± 3.2
27.9 ± 3.0
7.4 ± 0.8
20.2 ± 2.0
19.2 ± 0.3
13.7 ± 0.7
50.0 ± 0.5
12.5 ± 0.2
30.7 ± 1.5
25.9 ± 1.8
7.7 ± 0.2
Table 1: Main results with the Qwen3-32B teacher and Qwen3-1.7B student. Trained methods report mean ± standard deviation across three training seeds. SR (%) denotes success rate, with ALFWorld Overall pooling all splits. Score uses a 0–100 scale, and Rounds averages interactions per task. Dark/light green mark the best/second-best means among methods, with the best in bold.
ALFWorld
WebShop
Method
Seen SR ↑
Unseen SR ↑
Hard SR ↑
Overall SR ↑
Rounds ↓
SR ↑
Score ↑
Rounds ↓
Qwen3-8B-RL teacher → Qwen3-4B student
Student (zero-shot)
32.1
26.9
9.9
23.5
17.5
27.0
26.5
8.0
Teacher (zero-shot)
81.4
79.1
33.9
66.1
10.3
54.0
56.4
4.4
Vanilla OPD
71.4 ± 3.7
67.4 ± 3.5
20.9 ± 2.4
54.6 ± 2.7
11.3 ± 0.3
51.7 ± 1.5
55.2 ± 2.2
4.2 ± 0.1
TCOD-F2B
77.9 ± 1.4
71.9 ± 1.7
28.1 ± 1.4
60.6 ± 1.2
11.5 ± 0.2
55.3 ± 2.3
57.8 ± 2.8
4.0 ± 0.1
Table 2: Main results with the Qwen3-8B-RL teacher and Qwen3-4B student. Each teacher is trained with GiGPO on its benchmark. Results for trained methods are mean ± standard deviation across three training seeds. SR (%), Score (0–100), and Rounds follow Table 1 . Dark/light green mark the best/second-best means among compared methods, with the best in bold.
Figure 4: Training dynamics on WebShop. The top row shows the Qwen3-32B teacher with the Qwen3-1.7B student. The bottom row shows the Qwen3-8B-RL teacher with the Qwen3-4B student.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Outcome category
Gap group
ScienceWorld
WebShop
Average
Benign gaps
Large
2/172 (1.2%)
7/73 (9.6%)
5.4%
Teacher rescue
Large
9/172 (5.2%)
9/73 (12.3%)
8.8%
Teacher harm
Large
4/172 (2.3%)
14/73 (19.2%)
10.8%
Teacher-only completion
Large
26/172 (15.1%)
19/73 (26.0%)
20.6%
Outcome-critical
Small
0/37 (0.0%)
5/33 (15.2%)
7.6%
Appendix
Table 3: Counts underlying the empirical analysis. Each benchmark entry gives the number of decisions in an outcome category divided by the corresponding gap group size, followed by the percentage in parentheses. Categories follow Equation 33 . The average gives equal weight to the two benchmarks.
Figure 6: Final task outcomes under gap magnitude ranking. Decisions are ranked by decreasing gap magnitude. Rescue precision (left) is the teacher rescue rate; net rescue (right) subtracts the teacher harm rate.
Parameter
Value
Pre-calibration weight cap βmax
1.2
Positive-outcome weight floor β↑
1.5
Gap-magnitude / positive-increase quantile
qν=0.5 / qυ=0.6
Candidate checks / teacher proposals per trajectory
At most 2 / 2
Paired-evaluation positions per trajectory
At most 1
Matched trials per selected position
1 pair
Appendix
Table 4: OG-OPD weighting and calibration settings. Candidate thresholds follow Equation 36 . Training uses one matched pair per selected position, whereas the empirical analysis uses five matched trials per decision.
Benchmark
Method
Rounds ↓
s/step ↓
Cost ↓
ScienceWorld
Vanilla OPD
12.6
156.97
1.00×
OG-OPD
11.6
223.58
1.42×
WebShop
Vanilla OPD
6.8
67.21
1.00×
OG-OPD
5.6
99.70
1.48×
Appendix
Table 5: Training cost of vanilla OPD and OG-OPD. Runs use 8× A800 GPUs with the Qwen3-32B teacher and Qwen3-1.7B student. OG-OPD timing includes outcome-based calibration, and Cost is normalized to vanilla OPD. Rounds reports mean evaluation interaction rounds. Boldface marks the best value per benchmark.
On-policy distillation (OPD) provides teacher supervision on states visited by the student, reducing the distribution gap between training and inference. However, in multi-turn agentic tasks, student deviations may accumulate over time, gradually moving the trajectory away from states where teacher guidance remains effective. Our quantitative analysis further shows that high-disagreement states offer promising opportunities for teacher guidance, but determining whether such guidance is beneficial requires examining its effect on subsequent student trajectories. We propose FutureBridge-OPD (FTB), which executes a short teacher bridge at a high disagreement state and uses the resulting student continuation to assess whether the bridge increases the density of positive distillation signals relative to the teacher. On ALFWorld, WebShop, and ScienceWorld, under the main Qwen3-32B teacher to Qwen3-1.7B student setting, FTB outperforms vanilla OPD and TCOD by an average of 16.6 and 7.6 points, respectively, and remains effective across student scales and teacher settings. Our code is publicly available at https://github.com/ChenChiShui/FutureBridge-OPD.
Chishui Chen, Yaoyou Fan, Te Sun +11
Meituan LongCat Interaction · Peking University · Shanghai Jiao Tong University +6
Multi-turn agents that plan, invoke tools, and interact with environments offer a promising paradigm for solving complex tasks, yet their capabilities typically rely on very large models whose inference cost is prohibitive in practice. On-Policy Distillation (OPD) is a natural recipe for transferring such capabilities to smaller students, but we find that it suffers a characteristic failure mode in this setting: small student errors compound across turns and push the trajectory out of the teacher's familiar state distribution, so the teacher's supervision becomes least reliable precisely where the student needs it most. We propose Guided On-Policy Distillation (Guided-OPD), a simple yet effective algorithm that mixes teacher- and student-generated turns within each rollout and schedules the teacher's intervention probability along a curriculum that decays to zero. Strong guidance keeps early trajectories close to the teacher distribution and is then gradually withdrawn to recover the purely on-policy regime used at inference. On ALFWorld, ScienceWorld, and WebShop, distilling Qwen3 students from a Qwen3-30B-A3B teacher, Guided-OPD yields average relative gains of 21.1% in Score and 25.5% in Success Rate over vanilla OPD, with larger gains on smaller students.
Gengsheng Li, Mao Zheng, Mingyang Song +8
Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Large Language Model Department, Tencent +4
On-policy distillation (OPD) has shown strong potential for transferring reasoning ability from frontier or domain-specific models to smaller students. While effective on static single-turn tasks, its behavior in multi-turn agent settings remains underexplored. In this work, we identify a key limitation of vanilla OPD in such settings, which we term Trajectory-Level KL Instability. Specifically, we observe that KL divergence increases together with a drop in success rate, and even after convergence, the KL remains high, leading to unstable training. This instability arises from inter-turn error compounding: as errors accumulate, the student is driven beyond the teacher's effective support, rendering the supervision signal unreliable. To address this, we propose TCOD (Temporal Curriculum On-Policy Distillation), a simple yet effective framework that controls the trajectory depth exposed to the student and progressively expands it from short to long with a curriculum schedule. Experimental results across four student-teacher pairs on three multi-turn agent benchmarks (ALFWorld, WebShop, ScienceWorld) show that TCOD mitigates KL escalation and enhances KL stability throughout training, improving agent performance by up to 18 points over vanilla OPD. Further evaluations show that TCOD can even surpass the teacher's performance and generalize to tasks on which the teacher fails. Our code is available at https://github.com/kokolerk/TCOD.
Jiaqi Wang, Wenhao Zhang, Weijie Shi +2
Tongyi Lab , Alibaba Group · † The Chinese University of Hong Kong.