Learning from Teacher Continuations at Student States
Authors: Haojin Wang, Dylan Zhang, Huaibo Chen, Suhao Yu, Yihang Sun, Zhanyang Jin, Jiaying Ye, Dianqi Li, +3 more
Organizations: University of Illinois at Urbana-Champaign · Massachusetts Institute of Technology · University of Pennsylvania · University of Washington · International Business Machines
We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.
Figures & tables
Figure 1 : Overview of OLIVE . (a) Long chain-of-thought reasoning tasks: the student generates a reasoning prefix, which the teacher continues. (b) Long-horizon agentic tasks: the student interacts with the environment for part of an episode, then the teacher takes over. In both settings, prefixes are refreshed as the student policy changes.
Figure 2 : Synchronous and asynchronous implementations of OLIVE . (a) The student waits for the teacher continuation before updating on the same batch. (b) Prefix sampling for later batches overlaps with teacher generation, while student updates use completed continuations from earlier batches. Arrows connect teacher continuations to the updates that use them.
Table 3
ALFWorld
ScienceWorld
TextCraft
BabyAI
SearchQA
Method
SR ( % )
Turns
SR ( % )
Turns
SR ( % )
Turns
SR ( % )
Turns
SR ( % )
Turns
Student
19.38
26.97
0.12
28.63
23.00
23.89
38.33
14.52
30.50
11.91
Teacher
52.12
21.24
15.62
23.61
85.50
10.96
83.33
6.22
55.06
9.29
OPD
22.25
26.28
0.00
28.77
29.50
22.28
43.06
14.32
29.56
12.07
TCoD-B2F
37.10
23.30
0.75
29.28
39.50
19.98
66.30
9.03
37.69
11.12
TCoD-F2B
28.00
24.89
0.50
29.55
45.50
18.46
62.50
10.51
37.75
11.19
Table 2: Results on multi-turn agentic benchmarks (ALFWorld, ScienceWorld, TextCraft, BabyAI and SearchQA) with Qwen3-1.7B as the student and Qwen3-32B as the teacher. We report avg@4 success rate (SR, % ) and average trajectory score, along with the average turns across the five benchmarks.
Figure 4 : Training comparison between OPD, asynchronous OLIVE and synchronous OLIVE .
Figure 6
Figure 7 : Success rate on ScienceWorld over training 5 epochs and comparison with rollout number, with Qwen3-1.7B as the student and GPT-5.4-mini as the teacher.
Figure 8 : Success rate after 5 epochs of sequential training on each agentic environment.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Hyper-parameter
Value
Training temperature
1.0
Global batch size
64
Mini batch size
64
Student rollouts per prompt
4
LogProb top- K
16
Max response length
7168
Appendix
Table 3 : Default configuration of OPD used in our reasoning tasks experiments.
Hyper-parameter
Value
Supervision signal
Token-level CE on teacher continuation
Rollout mode
Online, asynchronous
Maximum staleness
3
Student rollouts per prompt
4
Prefix truncation point
4096
Teacher continuation length
1024
Appendix
Table 4 : Default configuration of OLIVE used in our reasoning tasks experiments.
Figure 9 : Avg@4 success rate of OLIVE on ALFWorld with different numbers of student turns before the handoff.
Figure 10 : Success-rate gain of OLIVE over offline distillation. OLIVE generalizes better.
Figure 11 : pass@4 on TBLite subsets with Qwen3.5-2B as the student.
On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts. In this work, we identify a common structural cause underlying OPD, which we call prefix failure. Under prefix failure, dense per-token supervision induces a bimodal teacher mixture and fragmented gradients that token-level loss truncation or reweighting fail to address. This observation motivates us to move beyond token-level loss interventions toward trajectory-level output corrections. We thus propose Trajectory-Refined Distillation (TRD), a trajectory-level correction method that revises the student's rollout under the teacher guidance while within on-policy support. By correcting problematic prefixes before distillation, TRD mitigates prefix failure at its source. Moreover, TRD improves the exploration by exposing the student to alternative valid derivations under teacher guidance, even when the original rolls are already correct. TRD can also be applied to on-policy self-distillation (OPSD), a parameter-sharing variant that uses the student model conditioned on privileged informations as the teacher. Across a wide range of benchmarks and base models at multiple scales, TRD consistently outperforms prior baselines, improving single-attempt accuracy and broadening reasoning coverage. Code is available at https://github.com/louieworth/trd
Li Jiang, Haoran Xu, Yichuan Ding +1
McGill University · Mila Quebec AI Institute · UT Austin
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
Distilling reasoning traces from strong large language models into smaller ones is a promising route to improve intelligence in resource-constrained settings. Existing approaches face a fundamental trade-off: offline distillation from teacher-generated traces provides high-quality, sample-efficient supervision but suffers from distributional drift: during training, the student model conditions on teacher-generated prefixes, whereas during inference the student autoregresses on self-generated prefixes, leading to compounding errors over long reasoning trajectories. Meanwhile, on-policy or self-distillation methods better match the student's inference-time distribution, but require costly online sampling and often produce low-quality traces in early training. We propose a principled offline reasoning distillation framework that preserves the efficiency and supervision quality of offline teacher-generated data while correcting teacher-student distribution drift. It adaptively emphasizes teacher supervision that is better aligned with the student's on-policy distribution. Evaluations on mathematical reasoning benchmarks of GSM8K, MATH, MATH500, and harder held-out competition-style tasks, including AMC, AIME, and OlympiadBench, show that our method improves reasoning accuracy over prior offline distillation algorithms and yields more stable reasoning traces while preserving instruction-following capabilities. Our work shows that lightweight, distribution-correction-aware training can substantially strengthen offline reasoning distillation without online rollouts.