On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Trajectory Dropout, a simple training-time intervention that exposes these weakened signals. The student first performs a standard full-context rollout to generate a complete trajectory. During training, we randomly drop a certain proportion of the student's reasoning trajectory, while the teacher continues to observe the complete trajectory for token-level supervision. This intervention strengthens corrections for overconfident predictions and introduces additional supervision at prefix-sensitive positions. Trajectory Dropout consistently improves average performance across teacher--student model pairs of different scales and six mathematical reasoning benchmarks, while also yielding gains on two out-of-domain benchmarks. It can also be flexibly integrated into existing OPD variants with negligible computational overhead, further improving their performance. These results demonstrate that Trajectory Dropout provides a simple mechanism for strengthening token-level supervision across model scales and OPD objectives.
Figures & tables
Figure 1: TD-OPD eases confidence and agreement saturation by adding Trajectory Dropout during training, boosts the training signals.
Figure 2: Two routes to weak OPD signals. A: cumulative signal magnitude after sorting tokens from weakest to strongest. B: the own-logit correction vanishes as student confidence approaches one, for fixed teacher probability. C: the signal vanishes near zero advantage, for fixed score norm.
Figure 3: Weak signals in OPD.
Figure 4: TD-OPD training. The student first completes a standard full-context rollout. During training, it reprocesses the saved trajectory with restricted attention to selected spans. The teacher scores every target on its full causal prefix. Old and current student scores use the same saved mask. These are fixed and trainable parameter states of the same student. All response targets receive supervision; evaluation uses full context.
Figure 5: Pass@k performance of OPD and OPD+TD on AIME25
Mathematical reasoning
Out-of-domain generalization
Method
AIME24
AIME25
AIME26
Olymp.
HMMT Feb26
HMMT Nov25
Math Avg
MBPP+
GPQA -D
(a) Non-thinking inference
Student
12.50
9.58
7.50
38.56
6.44
4.58
13.19
55.62
27.90
SFT
23.33
19.58
16.25
46.63
12.50
6.67
20.83
56.08
28.91
KD
23.75
21.25
15.42
48.07
12.88
7.50
21.48
56.48
28.66
GRPO
24.58
22.08
15.83
48.74
14.39
9.58
22.54
57.21
29.55
Table 1: Performance (%) of Qwen3-1.7B under non-thinking and thinking inference. Panel (b) enables thinking only at inference. Math Avg is the unweighted mean over the six mathematical benchmarks, excluding MBPP+ and GPQA Diamond. TD variants are shaded, with absolute gains (percentage points) over their counterparts without TD shown below each score. Bold and underlined values mark the best and second-best reported scores within each panel and column.
Figure 6: Sensitivity of TD-OPD to protection-window length and dropout ratio. Average performance across six mathematical reasoning benchmarks with varying (a) protection-window length w and (b) dropout ratio ρ .
Teacher
AIME24
AIME25
AIME26
None (initial)
12.50
9.58
7.50
Qwen3-1.7B
10.83
10.83
7.50
Qwen3-4B- Instruct-2507
40.42
30.42
25.83
Table 2: Teacher ablation under OPD+TD. The student is Qwen3-1.7B throughout.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Training dynamics of Qwen3-1.7B with OPD and OPD+TD ( ρ = 0.2).
Method
AIME24
AIME25
AIME26
Olymp.
HMMT Feb26
HMMT Nov25
Avg
Student
1.67
2.50
0.83
16.37
0.76
3.75
4.31
SFT
5.00
7.50
4.17
26.74
1.89
2.92
8.04
KD
4.17
7.08
5.00
27.15
3.03
2.92
8.22
GRPO
8.33
15.42
10.00
35.41
7.95
6.25
13.89
OPD
13.33
16.67
12.08
41.33
7.58
5.42
16.07
OPD + TD
12.92
17.92
15.00
42.41
11.74
8.75
18.12
Appendix
Table 3: Mathematical reasoning performance of Qwen3-0.6B in non-thinking mode. All scores are percentages. Avg is the unweighted mean over the six benchmarks. TD variants are shaded. The best scores are bold, and the second-best scores are underlined.
Parameter
TD-OPD ( ρ=0.2 )
Student
Qwen3-1.7B
Teacher
Qwen3-4B-Instruct-2507
Training data
DAPO, 9,755 teacher-correct English examples
Objective
K1 reverse-KL, policy gradient
Optimizer
AdamW
Learning rate / schedule
10−6 / constant
Appendix
Table 4: Core experimental configurations for the 1.7B student (TD-OPD).
Benchmark
Version / subset
Questions
n
AIME24
AIME I and II, 2024
30
8
AIME25
AIME I and II, 2025
30
8
AIME26
MathArena, AIME 2026
30
8
Olymp
OlympiadBench, English text-only mathematics
675
4
HMMT Feb26
MathArena, February 2026
33
8
HMMT Nov25
MathArena, November 2025
30
8
Appendix
Table 5: Evaluation datasets and samples per problem ( n ). Question counts refer to the evaluated versions.
Setting
Excess KL improvement
Radius 0.25×∥ΔθOPD∥
+0.008[0.007,0.010]
Radius 0.5×∥ΔθOPD∥
+0.016[0.013,0.020]
Radius 1.0×∥ΔθOPD∥
+0.027[0.022,0.034]
Unrescaled (natural norm)
+0.027[0.021,0.033]
Unpreconditioned fixed-step-size update
+0.147[0.115,0.182]
Appendix
Table 6: Excess KL improvement of the OPD+TD update over the paired OPD update on held-out validation, at matched update norms (mean with 95% bootstrap CI). The top three rows use the three matched radii; the bottom two rows are the unrescaled natural-norm update and the unpreconditioned fixed-step-size gradient-update control.
Figure 8: Paired one-step updates under matched norms. (a) Absolute KL improvement of the OPD update and the norm-matched OPD+TD update across the three radii: equal step size, unequal effect—the OPD+TD update is higher at every radius. (b) The excess improvement of OPD+TD over the paired OPD update grows nearly linearly with the radius (dashed line: proportional reference), indicating a directional rather than step-size effect. Error bars are 95% bootstrap CIs; the excess is positive in 16/16 training batches at every radius.
Figure 9: Directional advantage: schematic and decomposition. (a) With equal update norms, the difference between the OPD+TD and OPD updates lies in their alignment with the validation-objective direction gval (a longer projection corresponds to a larger excess KL improvement); angles are illustrative, not measured. (b) At matched norm ( 1.0× ), changing only a single channel (attention masking only, −0.004 ; loss-view masking only, −0.002 ) performs worse than the full-context OPD update, and only the full OPD+TD (both channels changed jointly, +0.027 ) improves.