On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.
Figures & tables
Figure 1: TD-OPD eases confidence and agreement saturation by adding Trajectory Dropout during training, boosts the training signals.
Figure 2: Two routes to weak OPD signals. A: cumulative signal magnitude after sorting tokens from weakest to strongest. B: the own-logit correction vanishes as student confidence approaches one, for fixed teacher probability. C: the signal vanishes near zero advantage, for fixed score norm.
Figure 3: Weak signals in OPD.
Figure 4: TD-OPD training. The student first completes a standard full-context rollout. During training, it reprocesses the saved trajectory with restricted attention to selected spans. The teacher scores every target on its full causal prefix. Old and current student scores use the same saved mask. These are fixed and trainable parameter states of the same student. All response targets receive supervision; evaluation uses full context.
Figure 5: Pass@k performance of OPD and OPD+TD on AIME25
Mathematical reasoning
Out-of-domain generalization
Method
AIME24
AIME25
AIME26
Olymp.
HMMT Feb26
HMMT Nov25
Math Avg
MBPP+
GPQA -D
(a) Non-thinking inference
Student
12.50
9.58
7.50
38.56
6.44
4.58
13.19
55.62
27.90
SFT
23.33
19.58
16.25
46.63
12.50
6.67
20.83
56.08
28.91
KD
23.75
21.25
15.42
48.07
12.88
7.50
21.48
56.48
28.66
GRPO
24.58
22.08
15.83
48.74
14.39
9.58
22.54
57.21
29.55
Table 1: Performance (%) of Qwen3-1.7B under non-thinking and thinking inference. Panel (b) enables thinking only at inference. Math Avg is the unweighted mean over the six mathematical benchmarks, excluding MBPP+ and GPQA Diamond. TD variants are shaded, with absolute gains (percentage points) over their counterparts without TD shown below each score. Bold and underlined values mark the best and second-best reported scores within each panel and column.
Figure 6: Sensitivity of TD-OPD to protection-window length and dropout ratio. Average performance across six mathematical reasoning benchmarks with varying (a) protection-window length w and (b) dropout ratio ρ .
Teacher
AIME24
AIME25
AIME26
None (initial)
12.50
9.58
7.50
Qwen3-1.7B
10.83
10.83
7.50
Qwen3-4B- Instruct-2507
40.42
30.42
25.83
Table 2: Teacher ablation under OPD+TD. The student is Qwen3-1.7B throughout.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Training dynamics of Qwen3-1.7B with OPD and OPD+TD ( ρ = 0.2).
Method
AIME24
AIME25
AIME26
Olymp.
HMMT Feb26
HMMT Nov25
Avg
Student
1.67
2.50
0.83
16.37
0.76
3.75
4.31
SFT
5.00
7.50
4.17
26.74
1.89
2.92
8.04
KD
4.17
7.08
5.00
27.15
3.03
2.92
8.22
GRPO
8.33
15.42
10.00
35.41
7.95
6.25
13.89
OPD
13.33
16.67
12.08
41.33
7.58
5.42
16.07
OPD + TD
12.92
17.92
15.00
42.41
11.74
8.75
18.12
Appendix
Table 3: Mathematical reasoning performance of Qwen3-0.6B in non-thinking mode. All scores are percentages. Avg is the unweighted mean over the six benchmarks. TD variants are shaded. The best scores are bold, and the second-best scores are underlined.
Parameter
TD-OPD ( ρ=0.2 )
Student
Qwen3-1.7B
Teacher
Qwen3-4B-Instruct-2507
Training data
DAPO, 9,755 teacher-correct English examples
Objective
K1 reverse-KL, policy gradient
Optimizer
AdamW
Learning rate / schedule
10−6 / constant
Appendix
Table 4: Core experimental configurations for the 1.7B student (TD-OPD).
Benchmark
Version / subset
Questions
n
AIME24
AIME I and II, 2024
30
8
AIME25
AIME I and II, 2025
30
8
AIME26
MathArena, AIME 2026
30
8
Olymp
OlympiadBench, English text-only mathematics
675
4
HMMT Feb26
MathArena, February 2026
33
8
HMMT Nov25
MathArena, November 2025
30
8
Appendix
Table 5: Evaluation datasets and samples per problem ( n ). Question counts refer to the evaluated versions.
Setting
Excess KL improvement
Radius 0.25×∥ΔθOPD∥
+0.008[0.007,0.010]
Radius 0.5×∥ΔθOPD∥
+0.016[0.013,0.020]
Radius 1.0×∥ΔθOPD∥
+0.027[0.022,0.034]
Unrescaled (natural norm)
+0.027[0.021,0.033]
Unpreconditioned fixed-step-size update
+0.147[0.115,0.182]
Appendix
Table 6: Excess KL improvement of the OPD+TD update over the paired OPD update on held-out validation, at matched update norms (mean with 95% bootstrap CI). The top three rows use the three matched radii; the bottom two rows are the unrescaled natural-norm update and the unpreconditioned fixed-step-size gradient-update control.
Figure 8: Paired one-step updates under matched norms. (a) Absolute KL improvement of the OPD update and the norm-matched OPD+TD update across the three radii: equal step size, unequal effect—the OPD+TD update is higher at every radius. (b) The excess improvement of OPD+TD over the paired OPD update grows nearly linearly with the radius (dashed line: proportional reference), indicating a directional rather than step-size effect. Error bars are 95% bootstrap CIs; the excess is positive in 16/16 training batches at every radius.
Figure 9: Directional advantage: schematic and decomposition. (a) With equal update norms, the difference between the OPD+TD and OPD updates lies in their alignment with the validation-objective direction gval (a longer projection corresponds to a larger excess KL improvement); angles are illustrative, not measured. (b) At matched norm ( 1.0× ), changing only a single channel (attention masking only, −0.004 ; loss-view masking only, −0.002 ) performs worse than the full-context OPD update, and only the full OPD+TD (both channels changed jointly, +0.027 ) improves.
On-policy distillation (OPD) has become a central post-training tool for large language models (LLMs), providing dense per-token teacher supervision along the student's own rollouts. In this work, we identify a common structural cause underlying OPD, which we call prefix failure. Under prefix failure, dense per-token supervision induces a bimodal teacher mixture and fragmented gradients that token-level loss truncation or reweighting fail to address. This observation motivates us to move beyond token-level loss interventions toward trajectory-level output corrections. We thus propose Trajectory-Refined Distillation (TRD), a trajectory-level correction method that revises the student's rollout under the teacher guidance while within on-policy support. By correcting problematic prefixes before distillation, TRD mitigates prefix failure at its source. Moreover, TRD improves the exploration by exposing the student to alternative valid derivations under teacher guidance, even when the original rolls are already correct. TRD can also be applied to on-policy self-distillation (OPSD), a parameter-sharing variant that uses the student model conditioned on privileged informations as the teacher. Across a wide range of benchmarks and base models at multiple scales, TRD consistently outperforms prior baselines, improving single-attempt accuracy and broadening reasoning coverage. Code is available at https://github.com/louieworth/trd
Li Jiang, Haoran Xu, Yichuan Ding +1
McGill University · Mila Quebec AI Institute · UT Austin
On-policy distillation (OPD) improves reasoning models by applying dense teacher supervision on student-sampled trajectories. However, scaling OPD to long-horizon mathematical reasoning exposes a reliability and efficiency problem: standard OPD assigns every sampled candidate the same long rollout budget, even though some trajectories may quickly become weakly aligned with the teacher and provide less useful supervision. Prior analyses suggest that successful OPD depends on local teacher-student compatibility, which can be measured by top-k overlap on student-visited prefixes. When this overlap is low, continuing to generate or train on long suffixes may waste computation and introduce noisy learning signal. To address this, we introduce Prefix-Guided On-Policy Distillation (PG-OPD), a simple rollout-allocation framework that uses fixed-length prefixes to estimate trajectory value before expensive long-horizon generation. PG-OPD first decodes every sampled candidate to the same prefix length, computes teacher-student top-k overlap within an early probe window of that prefix, and selectively continues high-overlap candidates to a fixed long length. Low-overlap candidates stop at the fixed prefix, avoiding unnecessary suffix generation. Across diverse teacher-student combinations on AMC, AIME, and HMMT benchmarks, PG-OPD improves average accuracy by up to 4.80 points while reducing training time by up to 2.46x. These results suggest that prefix-level compatibility provides a practical signal for directing OPD computation toward trajectories that remain learnable from the teacher.
On-policy distillation (OPD) grounds token-level supervision in the student's own trajectory, yet suffers from prefix failure: once the student commits to a wrong reasoning direction, all subsequent generation builds on this deviation, producing misdirected continuations that elicit unreliable supervision and waste compute. We identify a teacher-student continuation asymmetry on failed prefixes, where the teacher tends to redirect while the student continues along the original direction, and convert it into a label-free handoff trigger in Relay On-Policy Distillation (Relay-OPD). During training, Relay-OPD constructs relay trajectories by letting the teacher briefly take over at detected trigger points to produce a teacher leg, after which the student resumes and is optimized on the resulting trajectory. A limited relay budget concentrates intervention on critical early positions while limiting departure from the student policy. With a Qwen3-4B-Instruct-2507 teacher and Qwen3-0.6B/1.7B-Non-Thinking students on eight mathematical reasoning benchmarks, Relay-OPD achieves the best or second-best results on every benchmark, outperforming standard OPD by +5.73% and the strongest baseline FastOPD by +1.49% on average for 1.7B, with consistent gains at 0.6B. Training trajectory length is reduced by over 50%.
Haolei Xu, Xiaowen Xu, Haiwen Hong +5
1Zhejiang University · 2Yuvion Team, Alibaba Group