On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
Figures & tables
Figure 1: Three diagnostics of Cross-Tokenizer OPD for Qwen2.5-7B-Instruct → Llama-3.2-3B-Instruct. Left: The logged student-side alignment ratio remains high during Strict full training despite substantially lower static vocabulary overlap. Middle: Adding supervision on mismatch groups lowers mathematics and code performance. Right: At strictly aligned positions in student-generated responses before distillation, the shared vocabulary retains nearly all teacher and student probability mass on average, and a student-selected top-16 subset retains most of it on both sides.
Teacher → Student
Jvocab
Coverage
Training steps
1–100
1–20
41–60
81–100
Qwen → Llama
64.32
C1:1θ
93.56
94.63
93.20
93.43
C1:1T
85.91
88.65
84.90
85.36
Granite → Phi
39.49
C1:1θ
96.98
96.65
96.95
97.12
C1:1T
97.26
95.83
97.50
97.96
Granite → Qwen
64.87
C1:1θ
85.57
85.23
85.06
85.88
Table 1: Static vocabulary Jaccard overlap Jvocab and strict token coverage on student trajectories from Strict full training. C1:1θ and C1:1T denote student and teacher strict token coverage over the indicated training steps.
Figure 2: Full-average accuracy (%) across mismatch weights. Dashed lines mark strict supervision ( λ=0 ), and annotations give the decrease at λ=1.5 in percentage points.
Figure 3: Probability mass retained by student-selected top- k subsets and the full shared vocabulary at strict positions before distillation.
Method
Qwen → Llama
Granite → Phi
Granite → Qwen
Math
Code
Full
Math
Code
Full
Math
Code
Full
Base
18.32
35.60
26.96
33.06
40.04
36.55
24.61
34.51
29.56
ULD
20.51
37.74
29.12
35.24
42.01
38.63
33.36
20.87
27.12
Extended ULD
23.01
39.26
31.13
35.28
41.42
38.35
33.14
32.00
32.57
GOLD
21.21
33.10
27.16
36.22
47.29
41.75
38.26
54.84
46.55
SimCT
25.31
37.86
31.59
36.67
46.70
41.68
38.64
53.56
46.10
Table 2: Cross-tokenizer distillation accuracy (%). Base denotes the student before distillation. Bold marks the best distilled score per column.
Figure 6
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Training and optimization
Training iterations
100
Batch size
512
Dynamic batching
Disabled
Optimizer
AdamW ( Loshchilov & Hutter, 2019 )
Learning rate
1×10−6
Appendix
Table 3: Training settings.
Teacher → Student
Vocabulary size
Overlap (%)
∣VT∣
∣Vθ∣
∣V∩∣
∣Vθ∣∣V∩∣
∣VT∣∣V∩∣
Jvocab
Qwen → Llama
151,665
128,256
109,567
85.43
72.24
64.32
Granite → Phi
100,352
200,029
85,034
42.51
84.74
39.49
Granite → Qwen
100,352
151,665
99,163
65.38
98.82
64.87
Appendix
Table 4: Static vocabulary overlap across the three teacher–student pairs. Vocabulary counts follow the shared-vocabulary convention in Section 2 . Overlap ratios are reported as percentages.
Subset
Qwen → Llama
Granite → Phi
Granite → Qwen
Student
Teacher
Student
Teacher
Student
Teacher
k=4
91.74
89.81
90.08
88.12
96.58
94.53
k=8
93.60
92.43
92.90
91.43
97.81
96.48
k=16
94.55
93.92
94.68
93.54
98.44
97.67
k=32
95.06
94.77
95.89
95.03
98.82
98.37
k=64
95.35
95.29
96.75
96.14
99.07
98.82
Appendix
Table 5: Mean retained probability mass (%) at strict positions before distillation. “Shared” denotes the full shared vocabulary, and bold values mark k=16 .
Step
Qwen → Llama
Granite → Phi
Granite → Qwen
cmis
cctrl
ρ
cmis
cctrl
ρ
cmis
cctrl
ρ
0
0.030
0.956
0.294
0.066
0.986
0.008
−0.360
0.977
0.015
20
0.058
0.671
1.231
0.008
0.955
0.021
−0.129
0.581
0.113
60
−0.011
0.296
1.877
0.044
0.733
0.100
0.002
0.049
0.222
100
−0.021
0.208
1.942
−0.004
0.246
0.233
−0.001
0.076
0.269
Appendix
Table 6: Gradient diagnostics at checkpoints from training with only the strict loss. cmis and cctrl denote the mismatch–strict and split-half strict gradient cosines, respectively. ρ is the unweighted span-to-strict norm ratio. Values are rounded to three decimals; figure growth factors use the unrounded medians.
Benchmark
Qwen → Llama
Granite → Phi
Granite → Qwen
Base
Teacher
Base
Teacher
Base
Teacher
MATH500
34.32
74.50
62.94
80.81
48.60
80.81
GSM8K
63.44
89.93
86.07
91.51
67.65
91.51
AIME-2024
2.29
11.88
6.35
22.71
5.10
22.71
AIME-2025
0.31
7.50
3.85
18.65
1.77
18.65
AIME-2026
0.52
7.50
4.69
12.81
2.81
12.81
Appendix
Table 7: Benchmark-level scores of the undistilled students and frozen teachers for the three main model pairs.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict
full
top-16
top-128
MATH500
40.73
47.33
41.70
50.55
52.74
52.92
52.08
GSM8K
62.11
63.96
70.00
78.20
80.21
80.26
80.08
AIME-2024
5.00
5.73
3.33
3.65
4.69
5.00
5.10
AIME-2025
0.52
1.15
0.62
1.56
1.67
1.88
1.98
AIME-2026
0.94
0.73
0.00
0.73
0.94
1.25
0.52
Appendix
Table 8: Benchmark-level comparison for Qwen2.5-7B-Instruct → Llama-3.2-3B-Instruct.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict
full
top-16
top-128
MATH500
68.20
68.11
69.06
68.86
70.55
70.10
70.26
GSM8K
86.00
86.25
87.06
86.95
87.53
87.56
87.43
AIME-2024
8.85
8.33
9.48
10.73
11.25
10.83
11.46
AIME-2025
3.44
5.21
6.56
5.83
6.46
7.19
6.98
AIME-2026
5.52
6.15
6.25
6.88
7.81
7.29
6.77
Appendix
Table 9: Benchmark-level comparison for Granite-4.1-8B → Phi-4-mini-instruct.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict
full
top-16
top-128
MATH500
65.84
66.16
73.31
73.73
74.91
75.25
75.31
GSM8K
85.63
83.97
88.03
87.75
88.64
88.81
88.87
AIME-2024
7.60
8.23
10.94
11.67
14.37
14.17
14.17
AIME-2025
3.85
4.38
8.33
9.06
9.27
8.75
9.69
AIME-2026
6.98
5.31
7.29
5.73
6.25
6.46
6.46
Appendix
Table 10: Benchmark-level comparison for Granite-4.1-8B → Qwen2.5-7B-Base.
Benchmark
Span-MSE weight λ
0
0.25
0.5
0.75
1
1.25
1.5
MATH500
52.74
52.38
51.21
51.19
50.50
50.18
49.95
GSM8K
80.21
79.96
79.35
79.65
79.11
78.94
78.75
AIME-2024
4.69
5.10
5.21
5.21
4.58
4.06
5.62
AIME-2025
1.67
1.77
1.56
1.56
1.35
1.25
1.25
AIME-2026
0.94
0.42
0.62
0.31
0.42
0.52
0.21
Appendix
Table 11: Benchmark-level span-MSE weight sweep for Qwen2.5-7B-Instruct → Llama-3.2-3B-Instruct.
Benchmark
Span-MSE weight λ
0
0.25
0.5
0.75
1
1.25
1.5
MATH500
70.55
71.04
70.01
70.44
69.59
70.47
70.08
GSM8K
87.53
87.64
87.31
87.57
87.09
87.31
87.46
AIME-2024
11.25
10.31
10.42
10.31
9.69
10.31
10.00
AIME-2025
6.46
6.67
6.46
6.77
7.08
6.25
6.46
AIME-2026
7.81
5.94
6.77
7.19
6.88
6.77
6.35
Appendix
Table 12: Benchmark-level span-MSE weight sweep for Granite-4.1-8B → Phi-4-mini-instruct.
Benchmark
Span-MSE weight λ
0
0.25
0.5
0.75
1
1.25
1.5
MATH500
74.91
75.45
74.97
74.84
74.54
74.72
74.81
GSM8K
88.64
88.79
88.75
88.77
88.83
88.95
88.77
AIME-2024
14.37
15.42
13.02
14.27
12.71
12.19
12.92
AIME-2025
9.27
8.02
9.27
9.69
8.23
8.02
8.02
AIME-2026
6.25
6.15
5.94
6.15
5.83
5.83
5.52
Appendix
Table 13: Benchmark-level span-MSE weight sweep for Granite-4.1-8B → Qwen2.5-7B-Base.
Benchmark
Base
Teacher
MATH500
80.81
95.88
GSM8K
91.51
94.24
AIME-2024
22.71
60.94
AIME-2025
18.65
51.98
AIME-2026
12.81
65.52
AMC23
64.53
90.00
Appendix
Table 14: Benchmark-level scores of the undistilled student and frozen teacher for Qwen3-235B-A22B-Instruct-2507 → Granite-4.1-8B.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict top-16
MATH500
81.31
87.29
85.91
86.08
89.75
GSM8K
91.59
93.06
93.12
93.15
93.69
AIME-2024
23.23
37.19
35.52
37.40
43.02
AIME-2025
21.15
26.98
27.19
29.17
35.10
AIME-2026
13.33
30.10
23.02
27.29
38.44
AMC23
67.34
72.66
69.30
74.22
76.80
Appendix
Table 15: Benchmark-level distillation results for Qwen3-235B-A22B-Instruct-2507 → Granite-4.1-8B. Bold marks the best Math, Code, and Full averages.
Method
Seen
Unseen
Overall
SR ↑
Turns ↓
SR ↑
Turns ↓
SR ↑
Turns ↓
Base
12.50
46.60
8.85
47.28
10.68
46.94
Teacher
29.69
41.07
29.43
41.22
29.56
41.15
ULD
1.56
49.61
1.30
49.74
1.43
49.68
Extended ULD
25.52
41.95
23.44
42.61
24.48
42.28
GOLD
3.65
49.30
1.56
49.58
2.61
49.44
Appendix
Table 16: ALFWorld results for Qwen3-4B-Instruct-2507 → Llama-3.2-3B-Instruct. Bold marks the highest SR and lowest Turns among distilled models.
On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and student predictions are comparable token by token, an assumption that fails whenever the two models tokenize the same text differently. Under heterogeneous tokenizers, exact shared-token matching silently discards a large fraction of the teacher signal at precisely the positions where vocabularies disagree. We propose \textbf{\underline{Sim}ple \underline{C}ross-\underline{T}okenizer OPD (SimCT)}, which restores this signal by enlarging the supervision space: alongside shared tokens, SimCT compares teacher and student over short multi-token continuations that both tokenizers can realize, leaving the OPD loss form itself unchanged. We show that these units are the finest jointly tokenizable supervision interface, and that coarser alternatives remove teacher-student distinctions that are useful for on-policy learning. Across three heterogeneous teacher-student pairs on mathematical reasoning and code-generation benchmarks, SimCT shows consistent gains over shared-vocabulary OPD and representative cross-tokenizer baselines, with ablations confirming that the improvements come from recovering supervision discarded by exact shared-token matching. Code is available at https://github.com/sunjie279/SimCT-.
Jie Sun, Mao Zheng, Mingyang Song +6
University of Science and Technology of China · Shanghai Innovation Institute · Large Language Model Department, Tencent +1
On-policy distillation transfers reasoning capabilities by training a student model on its own generated trajectories using token-level feedback from a teacher. However, we identify a critical bottleneck, \textbf{Supervision Fidelity Decay (SFD)}: as student-generated prefixes lengthen, the teacher's next-token distribution becomes less confident and less discriminative. Consequently, the teacher-dependent corrective signal in reverse-KL distillation weakens, causing student drift to compound across long reasoning chains. To mitigate SFD, we introduce \textbf{Lookahead Group Reward (\ours{})}. Building on the insight that next-step teacher confidence reflects the discriminative strength of future reverse-KL supervision, \ours{} evaluates the student's top-K candidate tokens by the teacher confidence they induce at the subsequent step and assigns a group-normalized reward. To maintain computational efficiency, we further design an entropy-triggered tree-attention mechanism. Across six math and code benchmarks, \ours{} improves mean@8 by \textbf{2.57} points over OPD for a 7B student, with gains increasing in longer-generation and reaching +\textbf{4.92} points on AIME-26 at 39k tokens.
Yanjiang Liu, Jie Lou, Xinyan Guan +7
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu
On-policy distillation (OPD) uses teacher correction on student-generated responses. Full-vocabulary correction can provide important corrections even for tokens that the student assigns low probability, but backpropagating through all token logits becomes memory-intensive for long sequences. Existing memory-saving approaches estimate corrections from sampled tokens or restrict supervision to the student's TopK tokens, introducing sampling noise or changing the full-vocabulary correction. We introduce \textbf{SparseOPD}, which uses full-vocabulary teacher correction to determine which corrections matter before selecting the token logits to differentiate. SparseOPD first constructs the full-vocabulary correction without retaining its backward graph, then selects tokens by correction magnitude rather than student probability. Signed residual compensation preserves the total promoting and suppressing correction mass, while correction-aware budget allocation distributes the sparse support across positions. Finally, the update backpropagates only through the selected token logits. Across six task--scale settings spanning mathematics, chemistry QA, and multimodal reasoning, SparseOPD outperforms Sampled Token and TopK in task-average accuracy and matches or exceeds Full Vocabulary. Gradient cosine similarity reaches 99% on 4B mathematics, while 8K full-parameter profiling shows 70.5% lower backward memory.
Yongliang Miao, Shuang Liu, Yanguang Liu +2
The Chinese University of Hong Kong, Shenzhen · Carnegie Mellon University · New Jersey Institute of Technology +1