cs.CLOct 7, 2026

Stochastic Teacher Intervention for Agentic On-Policy Distillation

Authors: Junnan Liu, Linhao Luo, Zhijun Chen, Qianren Mao, Thuy-Trang Vu, Gholamreza Haffari

Organizations: Department of Data Science and AI, Faculty of Information Technology, Monash University, Australia · The Hong Kong Polytechnic University, Hong Kong SAR, China · Zhongguancun Laboratory, Beijing, P.R. China

Abstract

On-policy distillation (OPD) efficiently transfers capabilities from a stronger teacher to a student language model through dense token-level supervision on student-generated rollouts and has shown promise on complex tasks such as mathematical reasoning. However, in multi-turn agentic tasks, student decisions shape subsequent observations, causing early errors to accumulate across turns. The resulting trajectories can drift away from the teacher's rollout distribution, making the teacher's token-level supervision less reliable or even counterproductive for OPD training. To address this issue, we introduce STI-OPD, a stochastic teacher intervention framework for multi-turn agentic OPD. During multi-turn interaction, STI-OPD uses teacher intervention guided by teacher-student policy discrepancy to replace the student's proposed action with a teacher-generated one to maximize the acquisition of reliable supervision. We further develop a stochastic intervention strategy, addressing the limitations of previous threshold-based or fixed-schedule approaches, that estimates policy discrepancy using KL divergence and maps it to an intervention probability. By sampling whether to intervene from this probability, STI-OPD adaptively balances teacher control with student exploration. To learn from the resulting mixed-policy trajectories, we introduce an Importance-Weighted Reverse KL objective that corrects the token sampling mismatch between teacher-generated responses and the student policy to preserve the original OPD objective. Across tool-integrated reasoning and long-horizon interaction, STI-OPD outperforms the strongest prior OPD baseline on every evaluated benchmark and student size. Ablations further show that both discrepancy-guided intervention and importance weighting contribute to these gains.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Multi-Turn On-Policy Distillation with Prefix Replay

    Jul 6, 2026Baohao Liao, Hanze Dong, Christof Monz +3LLM Agent TrainingOn-Policy Distillation

  2. MAD-OPD: Breaking the Ceiling in On-Policy Distillation via Multi-Agent Debate

    May 2, 2026Jianze Wang, Ying Liu, Jinlong Chen +7Multi-Agent DebateMulti-Teacher Knowledge Distillation