Organizations: Department of Data Science and AI, Faculty of Information Technology, Monash University, Australia · The Hong Kong Polytechnic University, Hong Kong SAR, China · Zhongguancun Laboratory, Beijing, P.R. China
On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.
Figures & tables
Figure 1: Illustration of trajectory-level distribution shift and turn-level teacher intervention in on-policy distillation.
Figure 2: Overview of STI-OPD . A non-negative k3 discrepancy estimate determines whether to retain the student proposal or replace it with a teacher-generated response before execution. The executed response is used for training, with importance ratios recovering the conditional reverse-KL expectation for teacher-generated tokens.
Method
AIME24
AIME25
AIME26
HMMT26
GPQA
LiveCodeBench v6
Avg@32
Avg@8
Avg@8
Qwen3-4B
Teacher
71.7±0.3
65.0±0.3
60.0±0.3
39.8±0.3
50.9±0.3
70.5±0.3
Qwen3-0.6B
Vanilla
8.3±0.4
15.8±0.6
9.6±0.5
10.2±0.4
26.0±0.7
20.7±0.5
OPD
11.2±0.5
18.5±0.5
12.4±0.4
11.8±0.6
27.1±0.6
23.5±0.7
Table 1: Results on tool-integrated reasoning benchmarks. Bold and underlined values denote the best and second-best results, respectively.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Qwen3-4B
Teacher
87.5±0.4
66.2±0.4
35.5±0.4
57.6±0.4
46.4±0.4
83.3±0.4
26.1±0.4
Qwen3-0.6B
Vanilla
4.2±0.3
0.0
1.1±0.2
0.0
0.0
3.5±0.3
0.8±0.2
OPD
65.2±0.4
18.5±0.5
14.8±0.4
4.5±0.3
27.4±0.5
6.8±0.4
7.3±0.5
Table 2: Success rates on long-horizon agentic benchmarks. We report results for six ALFWorld task categories and the average across ScienceWorld tasks. Bold and underlined values denote the best and second-best results, respectively.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
STI-OPD
93.7 ±0.5
69.1 ±0.4
41.2 ±0.3
62.0 ±0.5
56.0 ±0.4
73.7 ±0.5
27.5 ±0.5
w/ k1 estimator
89.5±0.5
65.2±0.4
39.4±0.5
59.7±0.5
51.3±0.4
67.4±0.5
23.8±0.5
w/o importance weighting
86.2±0.4
66.7±0.4
36.8±0.3
56.2±0.4
46.4±0.5
61.9±0.4
21.5±0.3
w/ forward KL
91.3±0.5
67.2±0.4
38.2±0.4
58.8±0.5
54.9±0.4
69.6±0.5
23.1±0.4
Table 3: Ablation results on the long-horizon agentic benchmarks with the Qwen3-1.7B student. Each variant modifies one component of STI-OPD : replacing k3 with k1 , removing importance weighting, or replacing reverse KL with forward KL.
Figure 3: Training dynamics of STI-OPD and standard OPD on ALFWorld with Qwen3-1.7B as the student and Qwen3-4B as the teacher.
Figure 7
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Parameter
Value
Model
Student
Qwen3-0.6B; Qwen3-1.7B
Teacher
Qwen3-4B; Qwen3-32B
Data
Tool-Integrated Reasoning
# Training prompts
34,501
Training epochs
1
Batch size
128
Appendix
Table 4: Training and rollout parameters. Unless stated otherwise, these settings are shared by STI-OPD and all distillation baselines.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Qwen3-32B
Teacher
91.7±0.4
72.1±0.3
62.1±0.3
80.4±0.5
65.5±0.4
93.1±0.3
47.8±0.4
Qwen3-0.6B
Vanilla
4.2±0.3
0.0±0.2
1.1±0.2
0.0±0.2
0.0±0.2
3.5±0.3
0.8±0.2
OPD
66.9±0.4
22.4±0.5
34.7±0.4
13.3±0.3
42.2±0.5
16.0±0.4
22.6±0.5
Appendix
Table 5: Success rates on long-horizon agentic benchmarks with Qwen3-32B as teacher. We report results for six ALFWorld task categories and the average across ScienceWorld tasks. Bold and underlined values denote the best and second-best results, respectively.
Method
IFEval
Arena-Hard
Qwen3-0.6B
Vanilla
60.1
8.6
STI-OPD w/ forward KL
55.8 ↓ 4.3
7.2 ↓ 1.4
STI-OPD
64.3 ↑ 4.2
9.1 ↑ 0.5
Qwen3-1.7B
Vanilla
71.3
41.3
Appendix
Table 6: Impact of preserving the reverse-KL objective on mitigating forgetting.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Thre. 0.5
88.5±0.5
62.7±0.4
36.5±0.4
55.1±0.5
49.7±0.4
65.3±0.5
20.8±0.4
Thre. 1.0
89.8±0.4
64.8±0.4
38.0±0.3
57.2±0.5
51.8±0.4
68.4±0.4
23.2±0.5
Thre. 2.0
89.0±0.4
63.3±0.5
35.8±0.5
54.5±0.6
49.2±0.5
64.7±0.6
20.3±0.5
STI-OPD
93.7 ±0.5
69.1 ±0.5
41.2 ±0.2
62.0 ±0.6
56.0 ±0.6
73.7 ±0.5
27.5 ±0.6
Appendix
Table 7: Comparison of stochastic intervention and deterministic threshold on ALFWorld and ScienceWorld benchmarks.
Method
ALFWorld
Pick
Pick Two
Pick & Clean
Pick & Heat
Pick & Cool
Look
ScienceWorld
Relay-OPD
84.9±0.5
44.8±0.5
33.5±0.4
32.4±0.4
47.3±0.5
35.1±0.5
14.6±0.4
Guided-OPD
86.2±0.4
50.3±0.5
34.8±0.4
39.5±0.4
48.2±0.4
45.4±0.5
13.1±0.5
FutureBridge-OPD
88.3±0.4
61.0±0.5
36.7±0.3
53.8±0.4
49.7±0.4
62.9±0.5
18.7±0.4
STI-OPD
90.8 ±0.4
65.8 ±0.4
38.1 ±0.2
58.6 ±0.5
52.8 ±0.4
69.9 ±0.4
24.2 ±0.4
Appendix
Table 8: Results of STI-OPD and baselines under the same teacher computation budget.