On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99% on 4B mathematics, while 8K full-parameter profiling shows 70.5% lower backward memory.
Figures & tables
Figure 1: Full-vocabulary correction need not require full-vocabulary differentiation. Arrows show logit corrections in a constructed multimodal QA example. Full Vocabulary suppresses “union” and “sum” and promotes “overlap” and “intersection,” at high memory cost. Sampled Token incorrectly promotes “sum,” while TopK provides no direct correction to “overlap” or “intersection.” SparseOPD preserves the important full-vocabulary corrections with low memory use.
Figure 2: Overview of SparseOPD. (a) Full-vocabulary correction Ci,v=pi,v(ri,v−Di) . (b) Correction mass Mi=∥Ci∥1/2 allocates support. (c) Signed Top- C selects Si=Si+∪Si− . (d) Centered compensation yields Ci while preserving signed totals. (e) Sparse backpropagation passes only through selected logits zi,Si=WSihi .
Method
Mathematical Reasoning
Chemistry QA
MATH
Minerva
AMC23
AIME25
Avg.
Chemistry
GPQA
Avg.
Qwen3-1.7B → Qwen3-1.7B-Base
Teacher
72.3
28.8
44.4
8.3
38.4
42.0
26.5
34.2
Base
46.5
15.3
26.9
2.9
22.9
24.9
22.9
23.9
Sampled Token
54.1 +7.6
18.2 +2.9
32.8 +5.9
3.8 +0.9
27.2 +4.3
34.3 +9.4
28.6 +5.7
31.5 +7.6
TopK
63.6 +17.1
25.7 +10.4
34.1 +7.2
4.2 +1.3
31.9 +9.0
40.7 +15.8
29.2 +6.3
34.9 +11.0
Table 1: Mathematical reasoning and chemistry QA. Superscripts show percentage-point changes from Base; bold indicates the best result.
Method
MathVision
MathVista
WeMath
MathVerse
Avg.
MMFineReason-2B → Qwen3-VL-2B-Instruct
Teacher
24.4
65.1
66.1
46.0
50.4
Student
15.9
59.3
56.2
36.5
42.0
Sampled Token
22.2 +6.3
60.8 +1.5
62.9 +6.7
40.7 +4.2
46.7 +4.7
TopK
23.9 +8.1
62.2 +2.9
62.2 +6.0
43.3 +6.7
47.9 +5.9
SparseOPD(Ours)
24.2 +8.3
63.8 +4.5
64.8 +8.6
42.0 +5.5
48.7 +6.7
Table 2: Multimodal mathematical reasoning. Superscripts show percentage-point changes from Student; bold indicates the best result.
Figure 3: Task-average accuracy across student scales. SparseOPD outperforms both sparse baselines in all six settings and matches or exceeds Full Vocabulary in five.
Variant
MATH500
Minerva
AMC23
AIME25
Avg.
SparseOPD(Ours)
76.9
30.1
50.0
15.8
43.2
Student Probability Ranking
76.0 − 0.9
29.0 − 1.1
45.6 − 4.4
14.2 − 1.6
41.2 − 2.0
w/o Centered Compensation
76.3 − 0.6
29.2 − 0.9
47.5 − 2.5
14.2 − 1.6
41.8 − 1.4
w/o Batch-global Allocation
76.6 − 0.3
29.6 − 0.5
48.1 − 1.9
15.0 − 0.8
42.3 − 0.9
Table 3: Ablation of SparseOPD on mathematical reasoning. Results use Qwen3-4B → Qwen3-4B-Base under the same support budget. Red superscripts show drops from SparseOPD.
Method
Cos. ↑
Rel. ℓ2↓
Norm →1
Sign ↑
Full Vocabulary
1.00
0.00
1.00
1.00
Sampled Token
0.29
5.43
5.63
0.59
TopK
0.90
0.45
1.05
0.83
SparseOPD(Ours)
0.99
0.15
0.99
0.91
Table 4: Training-averaged gradient fidelity to Full Vocabulary OPD. Bold marks the best.
Length
Method
Memory (GiB)
Forward
Backward
4K
Full Vocabulary
12.82
10.09
SparseOPD(Ours)
1.94 ( ↓ 84.9%)
3.59 ( ↓ 64.5%)
8K
Full Vocabulary
25.50
20.05
SparseOPD(Ours)
3.88 ( ↓ 84.8%)
5.91 ( ↓ 70.5%)
16K
Full Vocabulary
50.85
39.96
Table 5: Memory usage. Peak extra GPU allocation above the pre-step baseline (Appendix B.4 ).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Optimizer
AdamW ( Loshchilov and Hutter, 2019 )
Initial learning rate / weight decay
10−5 / 0
Maximum gradient norm
1.0
Training hardware
8 NVIDIA H20 GPUs
Global batch size
64 trajectories
Training time per run
1.8 – 4 hours
Appendix
Table 6: Shared settings for the main accuracy experiments.
Setting
Mathematics
Chemistry
Multimodal
Student sizes
1.7B, 4B
1.7B, 4B
2B, 4B
Training questions
Full DAPO14k train split
1,890
5,782
Training epochs
1
3
1
Maximum response tokens
2,048
512
2,048
Distillation temperature
1.0
1.0
1.0
Appendix
Table 7: Task-specific training settings.
Setting
Mathematics
Chemistry
Multimodal
Generations per question
8
8
1
Temperature
0.6
0.6
1.0
Top- p
0.9
0.9
0.9
Maximum prompt tokens
1,024
1,024
1,024
Maximum output tokens
4,096
512
2,048
Appendix
Table 8: Evaluation decoding parameters. All tasks disable the model’s thinking mode and impose no top- k sampling restriction.
Setting
Value
Hardware
Two H200 GPUs
Teacher / student
Qwen3-4B / Qwen3-4B-Base
Updated parameters
All student parameters
Measured completion lengths
2,048 / 4,096 / 8,192 / 16,384 tokens
Global batch / microbatch per GPU
64 / 1 sequence
Gradient accumulation
32 microbatches per GPU
Appendix
Table 9: Protocol for the efficiency measurements.
Length
Method
Forward
Backward
Step (s)
Time (s)
Peak (GiB)
Time (s)
Peak (GiB)
2K
Full Vocabulary
6.25
84.82
20.49
83.46
33.31
Sampled Token
6.51
82.48
20.18
81.55
32.86
TopK
5.82
81.27
18.64
81.55
30.14
SparseOPD(Ours)
3.30
79.42
19.07
81.57
37.45
4K
Full Vocabulary
10.64
91.23
28.07
88.50
45.20
Appendix
Table 10: Efficiency results for all four methods. Memory columns are total peak Torch allocations during each stage, before baseline subtraction; the main table instead reports extra memory. Step time covers the complete update window, including vocabulary sparsification planning. Bold marks the lowest measured value at each length, including ties.
Setting
MATH500
Minerva
AMC23
AIME25
Avg.
SparseOPD(Ours)
64.5
26.1
35.9
5.8
33.1
Seed 43 ( K=20 )
65.10
25.28
34.69
5.42
32.62
Seed 44 ( K=20 )
65.70
25.60
37.19
3.75
33.06
Mean ± std
65.10±0.60
25.66±0.41
35.93±1.25
4.99±1.09
32.93±0.27
K=16
64.25
26.15
31.25
5.42
31.77
K=24
65.60
26.42
36.25
5.83
33.53
Appendix
Table 11: Parameter sensitivity on 1.7B mathematical reasoning. Scores are Mean@8 accuracy (%) under the same evaluation protocol; Avg. is the unweighted mean over benchmarks. SparseOPD(Ours) reproduces the main result ( K=20 , seed 42). The shaded summary reports mean ± sample standard deviation across seeds 42, 43, and 44, computed from the reported scores. Budget variants use seed 42.
On-policy distillation (OPD) trains a student on its own generated responses using dense, token-level supervision from a stronger teacher. Vanilla OPD treats all teacher signals equally, assuming that the teacher's supervision is equally important for every token. However, teacher signals at different tokens may have very different effects on the student's performance: some correct important reasoning errors, while others have little effect on the final answer. Motivated by this observation, we introduce Dr. OPD (OPD Done Right), which defines the optimal weighted OPD to maximize the student's performance. We formulate Dr. OPD as a bilevel optimization problem in which the student learns from weighted teacher supervision, while the weights are selected to maximize the expected reward of the resulting student. To solve Dr. OPD, we develop an efficient iterative solver that updates the token weights and student policy alternatively. At each round, it updates weights in closed form and then takes one gradient step on the resulting weighted OPD objective. Under regularity conditions, we show that this weighted update achieves a higher expected reward than a vanilla OPD update. Empirically, across strong-to-weak and same-size distillation on math and code, Dr. OPD consistently outperforms all evaluated baselines. In particular, in the strong-to-weak distillation setting, Dr. OPD improves average math performance by 9.7 points over vanilla OPD, and enables the smaller student to surpass its larger teacher.
On-policy distillation (OPD) offers dense token-level supervision as an alternative to the sparse outcome-level advantages of reinforcement learning with verifiable rewards (RLVR). However, the teacher scores student-generated trajectories that are inherently off-policy for it, so the reliability of its supervision, and hence the source of the student's improvement, remains unclear. We quantitatively analyze teacher supervision during OPD training and find substantial noise whose prevalence increases with teacher scale. Surprisingly, the student policy is insensitive to such noise, converging to comparable performance regardless of whether noisy supervision is retained or removed. Does OPD distill at all? By analyzing what drives its gains, we find that learning concentrates on low log-probability tokens, and using a single fixed negative advantage matches the performance of teacher-provided ones. This suggests that OPD works largely by suppressing low log-probability tokens, which requires no teacher. These findings motivate On-Policy Self-Adaptation (OPSA), a supervision-free method using entropy-adaptive negative advantages. It assigns stronger learning signals to high-entropy positions, suppressing tail tokens, and evenly redistributing probability mass among head tokens. Compared with the base \texttt{Qwen3-1.7B}, OPSA improves Avg@32 by 35.41 points on AIME24, corresponding to a 263% relative gain, and more than doubles Pass@32 across all three benchmarks. It also outperforms OPD by 16.77 points in Avg@32 on AIME24. Extensive experiments and analyses across model families and tasks further demonstrate its effectiveness and generalizability.
Yi Ding, Ruqi Zhang
Department of Computer Science, Purdue University, USA
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu