cs.AISep 27, 2026

What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation

Authors: Zizhuo Lin, Quanling Liu, Yi Yang, Yawei Luo

Organizations: Zhejiang University

Abstract

On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Trajectory-Refined Distillation

    Jun 7, 2026Li Jiang, Haoran Xu, Yichuan Ding +1Symbolic DistillationTeacher

  2. Prefix-Guided On-Policy Distillation: Mining Golden Trajectories from Rollouts

    Jun 20, 2026Qingfei Zhao, Huan Song, Shuyu Tian +2Unsupervised On-Policy Self-DistillationRollout

  3. Pass the Baton: Trajectory-Relayed On-Policy Distillation

    Jul 28, 2026Haolei Xu, Xiaowen Xu, Haiwen Hong +5TeacherMathematical Reasoning Benchmarks