Organizations: Department of Data Science and AI, Faculty of Information Technology, Monash University, Australia · The Hong Kong Polytechnic University, Hong Kong SAR, China · Zhongguancun Laboratory, Beijing, P.R. China
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
Figures & tables
Figure 1: Illustration of trajectory-level distribution shift and turn-level teacher intervention in on-policy distillation.
Figure 2: Overview of STI-OPD . A non-negative k3 discrepancy estimate determines whether to retain the student proposal or replace it with a teacher-generated response before execution. The executed response is used for training, with importance ratios recovering the conditional reverse-KL expectation for teacher-generated tokens.
Method
AIME24
AIME25
AIME26
HMMT26
GPQA
LiveCodeBench v6
Avg@32
Avg@8
Avg@8
Qwen3-4B
Teacher
71.7±0.3
65.0±0.3
60.0±0.3
39.8±0.3
50.9±0.3
70.5±0.3
Qwen3-0.6B
Vanilla
8.3±0.4
15.8±0.6
9.6±0.5
10.2±0.4
26.0±0.7
20.7±0.5
OPD
11.2±0.5
18.5±0.5
12.4±0.4
11.8±0.6
27.1±0.6
23.5±0.7
Table 1: Results on tool-integrated reasoning benchmarks. Bold and underlined values denote the best and second-best results, respectively.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Qwen3-4B
Teacher
87.5±0.4
66.2±0.4
35.5±0.4
57.6±0.4
46.4±0.4
83.3±0.4
26.1±0.4
Qwen3-0.6B
Vanilla
4.2±0.3
0.0
1.1±0.2
0.0
0.0
3.5±0.3
0.8±0.2
OPD
65.2±0.4
18.5±0.5
14.8±0.4
4.5±0.3
27.4±0.5
6.8±0.4
7.3±0.5
Table 2: Success rates on long-horizon agentic benchmarks. We report results for six ALFWorld task categories and the average across ScienceWorld tasks. Bold and underlined values denote the best and second-best results, respectively.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
STI-OPD
93.7 ±0.5
69.1 ±0.4
41.2 ±0.3
62.0 ±0.5
56.0 ±0.4
73.7 ±0.5
27.5 ±0.5
w/ k1 estimator
89.5±0.5
65.2±0.4
39.4±0.5
59.7±0.5
51.3±0.4
67.4±0.5
23.8±0.5
w/o importance weighting
86.2±0.4
66.7±0.4
36.8±0.3
56.2±0.4
46.4±0.5
61.9±0.4
21.5±0.3
w/ forward KL
91.3±0.5
67.2±0.4
38.2±0.4
58.8±0.5
54.9±0.4
69.6±0.5
23.1±0.4
Table 3: Ablation results on the long-horizon agentic benchmarks with the Qwen3-1.7B student. Each variant modifies one component of STI-OPD : replacing k3 with k1 , removing importance weighting, or replacing reverse KL with forward KL.
Figure 3: Training dynamics of STI-OPD and standard OPD on ALFWorld with Qwen3-1.7B as the student and Qwen3-4B as the teacher.
Figure 7
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Parameter
Value
Model
Student
Qwen3-0.6B; Qwen3-1.7B
Teacher
Qwen3-4B; Qwen3-32B
Data
Tool-Integrated Reasoning
# Training prompts
34,501
Training epochs
1
Batch size
128
Appendix
Table 4: Training and rollout parameters. Unless stated otherwise, these settings are shared by STI-OPD and all distillation baselines.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Qwen3-32B
Teacher
91.7±0.4
72.1±0.3
62.1±0.3
80.4±0.5
65.5±0.4
93.1±0.3
47.8±0.4
Qwen3-0.6B
Vanilla
4.2±0.3
0.0±0.2
1.1±0.2
0.0±0.2
0.0±0.2
3.5±0.3
0.8±0.2
OPD
66.9±0.4
22.4±0.5
34.7±0.4
13.3±0.3
42.2±0.5
16.0±0.4
22.6±0.5
Appendix
Table 5: Success rates on long-horizon agentic benchmarks with Qwen3-32B as teacher. We report results for six ALFWorld task categories and the average across ScienceWorld tasks. Bold and underlined values denote the best and second-best results, respectively.
Method
IFEval
Arena-Hard
Qwen3-0.6B
Vanilla
60.1
8.6
STI-OPD w/ forward KL
55.8 ↓ 4.3
7.2 ↓ 1.4
STI-OPD
64.3 ↑ 4.2
9.1 ↑ 0.5
Qwen3-1.7B
Vanilla
71.3
41.3
Appendix
Table 6: Impact of preserving the reverse-KL objective on mitigating forgetting.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Thre. 0.5
88.5±0.5
62.7±0.4
36.5±0.4
55.1±0.5
49.7±0.4
65.3±0.5
20.8±0.4
Thre. 1.0
89.8±0.4
64.8±0.4
38.0±0.3
57.2±0.5
51.8±0.4
68.4±0.4
23.2±0.5
Thre. 2.0
89.0±0.4
63.3±0.5
35.8±0.5
54.5±0.6
49.2±0.5
64.7±0.6
20.3±0.5
STI-OPD
93.7 ±0.5
69.1 ±0.5
41.2 ±0.2
62.0 ±0.6
56.0 ±0.6
73.7 ±0.5
27.5 ±0.6
Appendix
Table 7: Comparison of stochastic intervention and deterministic threshold on ALFWorld and ScienceWorld benchmarks.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Relay-OPD
84.9±0.5
44.8±0.5
33.5±0.4
32.4±0.4
47.3±0.5
35.1±0.5
14.6±0.4
Guided-OPD
86.2±0.4
50.3±0.5
34.8±0.4
39.5±0.4
48.2±0.4
45.4±0.5
13.1±0.5
FutureBridge-OPD
88.3±0.4
61.0±0.5
36.7±0.3
53.8±0.4
49.7±0.4
62.9±0.5
18.7±0.4
STI-OPD
90.8 ±0.4
65.8 ±0.4
38.1 ±0.2
58.6 ±0.5
52.8 ±0.4
69.9 ±0.4
24.2 ±0.4
Appendix
Table 8: Results of STI-OPD and baselines under the same teacher computation budget.
On-policy distillation (OPD) improves student models by training them on trajectories induced by their own policy, making it a promising approach for mitigating exposure bias in agent training. However, most OPD studies focus on single-turn settings, while realistic LLM agents interact with environments over multiple turns. In this regime, early errors can alter future observations and compound across the trajectory, and standard dense token-level OPD becomes brittle, as it may over-penalize semantically valid alternatives, reinforce local degeneracies such as repeated actions, and propagate unreliable teacher supervision on off-distribution histories. We propose SAGE-OPD, a verifier-free selective intervention framework specifically designed for multi-turn OPD. Instead of applying teacher supervision uniformly across all turns, SAGE-OPD first observes environment feedback and uses teacher judgment to decide whether each student response should be skipped or intervened on. To further address compounding errors, SAGE-OPD weights token-level distillation by teacher confidence, reducing the influence of uncertain teacher distributions on corrupted or ambiguous histories. Finally, SAGE-OPD applies loss normalization to preserve the overall loss scale of standard OPD while retaining selective turn-level weighting. Experiments on agent tasks show that SAGE-OPD consistently improves over baselines, achieving up to a 13.3% relative improvement in ALFWorld unseen success rate over standard OPD. Ablation studies further demonstrate that turn-level intervention, teacher confidence weighting, and loss normalization provide complementary benefits. Our results suggest that effective multi-turn OPD should remain on-policy, but teacher supervision should be selectively allocated to turns where intervention is necessary and reliable.
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn interaction histories. Fully online OPD is costly because each update requires fresh student rollouts through the environment and teacher queries at visited histories. We propose Replayed-Prefix On-Policy Distillation (ReOPD), an off-environment alternative that reuses pre-collected teacher trajectories as replayed prefixes: the student acts at selected steps, while the teacher provides dense per-step supervision without executing new environment interactions. We show that multi-turn OPD introduces a prefix trap: making histories more student-on-policy improves relevance to the student, but can query the teacher on histories where its target is unreliable. This creates a two-sided distribution shift between student occupancy and teacher reliability. ReOPD addresses this by treating multi-turn OPD as a reliability-aware prefix distribution design and implements it with a simple step-decaying sampling schedule that emphasizes early, lower-shift prefixes. Across mathematical reasoning with Python and search environments over multiple teacher and student model scales, ReOPD preserves or improves OPD-level accuracy, uses zero tool calls during student training, and is at least 4× faster per rollout than OPD. ReOPD therefore turns expensive agent-environment interaction into a reusable offline resource, enabling scalable distillation across tools, tasks, and environments.
On-policy distillation (OPD) trains a student on its own trajectories under token-level teacher supervision, but existing methods are capped by a single-teacher capability ceiling: when the teacher errs, the student inherits the error. OPD also remains largely unexplored in agentic tasks, where per-step errors compound across long trajectories and destabilize training. We propose MAD-OPD (Multi-Agent Debate-driven On-Policy Distillation), which breaks this ceiling by recasting the distillation teacher as a deliberative collective of teachers that debate over the student's on-policy state; the debate produces an emergent collective intelligence that supplies token-level supervision, with each teacher's contribution weighted by its post-debate confidence. To extend OPD to agentic tasks, we also introduce On-Policy Agentic Distillation (OPAD), which adds step-level sampling to stabilize training under multi-step error compounding. We additionally derive a task-adaptive divergence principle, selecting JSD (Jensen-Shannon divergence) for agentic stability and reverse KL (Kullback-Leibler) divergence for code generation, and verify it both theoretically and empirically. Across six teacher-student configurations (Qwen3 and Qwen3.5; 1.7B-14B students, 8B-32B teachers) and five agentic and code benchmarks, MAD-OPD ranks first across all six configurations; on the 14B+8B→4B setting it lifts the agentic average by +2.4% and the code average by +3.7% over the stronger single-teacher OPD.
Jianze Wang, Ying Liu, Jinlong Chen +7
School of Artificial Intelligence and Automation, Huazhong University of Science and Technology · Alibaba Group