On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99% on 4B mathematics, while 8K full-parameter profiling shows 70.5% lower backward memory.
Figures & tables
Figure 1: Full-vocabulary correction need not require full-vocabulary differentiation. Arrows show logit corrections in a constructed multimodal QA example. Full Vocabulary suppresses “union” and “sum” and promotes “overlap” and “intersection,” at high memory cost. Sampled Token incorrectly promotes “sum,” while TopK provides no direct correction to “overlap” or “intersection.” SparseOPD preserves the important full-vocabulary corrections with low memory use.
Figure 2: Overview of SparseOPD. (a) Full-vocabulary correction Ci,v=pi,v(ri,v−Di) . (b) Correction mass Mi=∥Ci∥1/2 allocates support. (c) Signed Top- C selects Si=Si+∪Si− . (d) Centered compensation yields Ci while preserving signed totals. (e) Sparse backpropagation passes only through selected logits zi,Si=WSihi .
Method
Mathematical Reasoning
Chemistry QA
MATH
Minerva
AMC23
AIME25
Avg.
Chemistry
GPQA
Avg.
Qwen3-1.7B → Qwen3-1.7B-Base
Teacher
72.3
28.8
44.4
8.3
38.4
42.0
26.5
34.2
Base
46.5
15.3
26.9
2.9
22.9
24.9
22.9
23.9
Sampled Token
54.1 +7.6
18.2 +2.9
32.8 +5.9
3.8 +0.9
27.2 +4.3
34.3 +9.4
28.6 +5.7
31.5 +7.6
TopK
63.6 +17.1
25.7 +10.4
34.1 +7.2
4.2 +1.3
31.9 +9.0
40.7 +15.8
29.2 +6.3
34.9 +11.0
Table 1: Mathematical reasoning and chemistry QA. Superscripts show percentage-point changes from Base; bold indicates the best result.
Method
MathVision
MathVista
WeMath
MathVerse
Avg.
MMFineReason-2B → Qwen3-VL-2B-Instruct
Teacher
24.4
65.1
66.1
46.0
50.4
Student
15.9
59.3
56.2
36.5
42.0
Sampled Token
22.2 +6.3
60.8 +1.5
62.9 +6.7
40.7 +4.2
46.7 +4.7
TopK
23.9 +8.1
62.2 +2.9
62.2 +6.0
43.3 +6.7
47.9 +5.9
SparseOPD(Ours)
24.2 +8.3
63.8 +4.5
64.8 +8.6
42.0 +5.5
48.7 +6.7
Table 2: Multimodal mathematical reasoning. Superscripts show percentage-point changes from Student; bold indicates the best result.
Figure 3: Task-average accuracy across student scales. SparseOPD outperforms both sparse baselines in all six settings and matches or exceeds Full Vocabulary in five.
Variant
MATH500
Minerva
AMC23
AIME25
Avg.
SparseOPD(Ours)
76.9
30.1
50.0
15.8
43.2
Student Probability Ranking
76.0 − 0.9
29.0 − 1.1
45.6 − 4.4
14.2 − 1.6
41.2 − 2.0
w/o Centered Compensation
76.3 − 0.6
29.2 − 0.9
47.5 − 2.5
14.2 − 1.6
41.8 − 1.4
w/o Batch-global Allocation
76.6 − 0.3
29.6 − 0.5
48.1 − 1.9
15.0 − 0.8
42.3 − 0.9
Table 3: Ablation of SparseOPD on mathematical reasoning. Results use Qwen3-4B → Qwen3-4B-Base under the same support budget. Red superscripts show drops from SparseOPD.
Method
Cos. ↑
Rel. ℓ2↓
Norm →1
Sign ↑
Full Vocabulary
1.00
0.00
1.00
1.00
Sampled Token
0.29
5.43
5.63
0.59
TopK
0.90
0.45
1.05
0.83
SparseOPD(Ours)
0.99
0.15
0.99
0.91
Table 4: Training-averaged gradient fidelity to Full Vocabulary OPD. Bold marks the best.
Length
Method
Memory (GiB)
Forward
Backward
4K
Full Vocabulary
12.82
10.09
SparseOPD(Ours)
1.94 ( ↓ 84.9%)
3.59 ( ↓ 64.5%)
8K
Full Vocabulary
25.50
20.05
SparseOPD(Ours)
3.88 ( ↓ 84.8%)
5.91 ( ↓ 70.5%)
16K
Full Vocabulary
50.85
39.96
Table 5: Memory usage. Peak extra GPU allocation above the pre-step baseline (Appendix B.4 ).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Optimizer
AdamW ( Loshchilov and Hutter, 2019 )
Initial learning rate / weight decay
10−5 / 0
Maximum gradient norm
1.0
Training hardware
8 NVIDIA H20 GPUs
Global batch size
64 trajectories
Training time per run
1.8 – 4 hours
Appendix
Table 6: Shared settings for the main accuracy experiments.
Setting
Mathematics
Chemistry
Multimodal
Student sizes
1.7B, 4B
1.7B, 4B
2B, 4B
Training questions
Full DAPO14k train split
1,890
5,782
Training epochs
1
3
1
Maximum response tokens
2,048
512
2,048
Distillation temperature
1.0
1.0
1.0
Appendix
Table 7: Task-specific training settings.
Setting
Mathematics
Chemistry
Multimodal
Generations per question
8
8
1
Temperature
0.6
0.6
1.0
Top- p
0.9
0.9
0.9
Maximum prompt tokens
1,024
1,024
1,024
Maximum output tokens
4,096
512
2,048
Appendix
Table 8: Evaluation decoding parameters. All tasks disable the model’s thinking mode and impose no top- k sampling restriction.
Setting
Value
Hardware
Two H200 GPUs
Teacher / student
Qwen3-4B / Qwen3-4B-Base
Updated parameters
All student parameters
Measured completion lengths
2,048 / 4,096 / 8,192 / 16,384 tokens
Global batch / microbatch per GPU
64 / 1 sequence
Gradient accumulation
32 microbatches per GPU
Appendix
Table 9: Protocol for the efficiency measurements.
Length
Method
Forward
Backward
Step (s)
Time (s)
Peak (GiB)
Time (s)
Peak (GiB)
2K
Full Vocabulary
6.25
84.82
20.49
83.46
33.31
Sampled Token
6.51
82.48
20.18
81.55
32.86
TopK
5.82
81.27
18.64
81.55
30.14
SparseOPD(Ours)
3.30
79.42
19.07
81.57
37.45
4K
Full Vocabulary
10.64
91.23
28.07
88.50
45.20
Appendix
Table 10: Efficiency results for all four methods. Memory columns are total peak Torch allocations during each stage, before baseline subtraction; the main table instead reports extra memory. Step time covers the complete update window, including vocabulary sparsification planning. Bold marks the lowest measured value at each length, including ties.
Setting
MATH500
Minerva
AMC23
AIME25
Avg.
SparseOPD(Ours)
64.5
26.1
35.9
5.8
33.1
Seed 43 ( K=20 )
65.10
25.28
34.69
5.42
32.62
Seed 44 ( K=20 )
65.70
25.60
37.19
3.75
33.06
Mean ± std
65.10±0.60
25.66±0.41
35.93±1.25
4.99±1.09
32.93±0.27
K=16
64.25
26.15
31.25
5.42
31.77
K=24
65.60
26.42
36.25
5.83
33.53
Appendix
Table 11: Parameter sensitivity on 1.7B mathematical reasoning. Scores are Mean@8 accuracy (%) under the same evaluation protocol; Avg. is the unweighted mean over benchmarks. SparseOPD(Ours) reproduces the main result ( K=20 , seed 42). The shaded summary reports mean ± sample standard deviation across seeds 42, 43, and 44, computed from the reported scores. Budget variants use seed 42.
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu