On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predictions requires alignment at both sequence and vocabulary levels. In this paper, we examine whether expanding this alignment coverage improves learning. Across three heterogeneous teacher--student pairs on mathematical reasoning and code generation, strict 1:1 groups already cover most student-generated tokens despite substantial vocabulary mismatch. On responses sampled from the students before distillation, the shared vocabulary retains nearly all teacher and student probability mass at strictly aligned positions on average. Restricting reverse KL to a student-selected top-16 subset of the shared vocabulary at each strict position achieves accuracy comparable to full shared-vocabulary OPD, outperforming the evaluated cross-tokenizer baselines. Adding mean squared error supervision on span log-probabilities in mismatch groups gives complete supervision coverage, yet reduces accuracy. At checkpoints from training with only the strict loss, the span gradients show weak or negative directional agreement with the strict gradients and grow in magnitude relative to them. These diagnostics may help explain the accuracy drop from adding span supervision. Our findings motivate a shift from maximizing alignment coverage to prioritizing supervision reliability: compact supervision at strict positions can be more effective than broader coverage that introduces weakly aligned or conflicting training signals.
Figures & tables
Figure 1: Three diagnostics of Cross-Tokenizer OPD for Qwen2.5-7B-Instruct → Llama-3.2-3B-Instruct. Left: The logged student-side alignment ratio remains high during Strict full training despite substantially lower static vocabulary overlap. Middle: Adding supervision on mismatch groups lowers mathematics and code performance. Right: At strictly aligned positions in student-generated responses before distillation, the shared vocabulary retains nearly all teacher and student probability mass on average, and a student-selected top-16 subset retains most of it on both sides.
Teacher → Student
Jvocab
Coverage
Training steps
1–100
1–20
41–60
81–100
Qwen → Llama
64.32
C1:1θ
93.56
94.63
93.20
93.43
C1:1T
85.91
88.65
84.90
85.36
Granite → Phi
39.49
C1:1θ
96.98
96.65
96.95
97.12
C1:1T
97.26
95.83
97.50
97.96
Granite → Qwen
64.87
C1:1θ
85.57
85.23
85.06
85.88
Table 1: Static vocabulary Jaccard overlap Jvocab and strict token coverage on student trajectories from Strict full training. C1:1θ and C1:1T denote student and teacher strict token coverage over the indicated training steps.
Figure 2: Full-average accuracy (%) across mismatch weights. Dashed lines mark strict supervision ( λ=0 ), and annotations give the decrease at λ=1.5 in percentage points.
Figure 3: Probability mass retained by student-selected top- k subsets and the full shared vocabulary at strict positions before distillation.
Method
Qwen → Llama
Granite → Phi
Granite → Qwen
Math
Code
Full
Math
Code
Full
Math
Code
Full
Base
18.32
35.60
26.96
33.06
40.04
36.55
24.61
34.51
29.56
ULD
20.51
37.74
29.12
35.24
42.01
38.63
33.36
20.87
27.12
Extended ULD
23.01
39.26
31.13
35.28
41.42
38.35
33.14
32.00
32.57
GOLD
21.21
33.10
27.16
36.22
47.29
41.75
38.26
54.84
46.55
SimCT
25.31
37.86
31.59
36.67
46.70
41.68
38.64
53.56
46.10
Table 2: Cross-tokenizer distillation accuracy (%). Base denotes the student before distillation. Bold marks the best distilled score per column.
Figure 6
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Training and optimization
Training iterations
100
Batch size
512
Dynamic batching
Disabled
Optimizer
AdamW ( Loshchilov & Hutter, 2019 )
Learning rate
1×10−6
Appendix
Table 3: Training settings.
Teacher → Student
Vocabulary size
Overlap (%)
∣VT∣
∣Vθ∣
∣V∩∣
∣Vθ∣∣V∩∣
∣VT∣∣V∩∣
Jvocab
Qwen → Llama
151,665
128,256
109,567
85.43
72.24
64.32
Granite → Phi
100,352
200,029
85,034
42.51
84.74
39.49
Granite → Qwen
100,352
151,665
99,163
65.38
98.82
64.87
Appendix
Table 4: Static vocabulary overlap across the three teacher–student pairs. Vocabulary counts follow the shared-vocabulary convention in Section 2 . Overlap ratios are reported as percentages.
Subset
Qwen → Llama
Granite → Phi
Granite → Qwen
Student
Teacher
Student
Teacher
Student
Teacher
k=4
91.74
89.81
90.08
88.12
96.58
94.53
k=8
93.60
92.43
92.90
91.43
97.81
96.48
k=16
94.55
93.92
94.68
93.54
98.44
97.67
k=32
95.06
94.77
95.89
95.03
98.82
98.37
k=64
95.35
95.29
96.75
96.14
99.07
98.82
Appendix
Table 5: Mean retained probability mass (%) at strict positions before distillation. “Shared” denotes the full shared vocabulary, and bold values mark k=16 .
Step
Qwen → Llama
Granite → Phi
Granite → Qwen
cmis
cctrl
ρ
cmis
cctrl
ρ
cmis
cctrl
ρ
0
0.030
0.956
0.294
0.066
0.986
0.008
−0.360
0.977
0.015
20
0.058
0.671
1.231
0.008
0.955
0.021
−0.129
0.581
0.113
60
−0.011
0.296
1.877
0.044
0.733
0.100
0.002
0.049
0.222
100
−0.021
0.208
1.942
−0.004
0.246
0.233
−0.001
0.076
0.269
Appendix
Table 6: Gradient diagnostics at checkpoints from training with only the strict loss. cmis and cctrl denote the mismatch–strict and split-half strict gradient cosines, respectively. ρ is the unweighted span-to-strict norm ratio. Values are rounded to three decimals; figure growth factors use the unrounded medians.
Benchmark
Qwen → Llama
Granite → Phi
Granite → Qwen
Base
Teacher
Base
Teacher
Base
Teacher
MATH500
34.32
74.50
62.94
80.81
48.60
80.81
GSM8K
63.44
89.93
86.07
91.51
67.65
91.51
AIME-2024
2.29
11.88
6.35
22.71
5.10
22.71
AIME-2025
0.31
7.50
3.85
18.65
1.77
18.65
AIME-2026
0.52
7.50
4.69
12.81
2.81
12.81
Appendix
Table 7: Benchmark-level scores of the undistilled students and frozen teachers for the three main model pairs.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict
full
top-16
top-128
MATH500
40.73
47.33
41.70
50.55
52.74
52.92
52.08
GSM8K
62.11
63.96
70.00
78.20
80.21
80.26
80.08
AIME-2024
5.00
5.73
3.33
3.65
4.69
5.00
5.10
AIME-2025
0.52
1.15
0.62
1.56
1.67
1.88
1.98
AIME-2026
0.94
0.73
0.00
0.73
0.94
1.25
0.52
Appendix
Table 8: Benchmark-level comparison for Qwen2.5-7B-Instruct → Llama-3.2-3B-Instruct.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict
full
top-16
top-128
MATH500
68.20
68.11
69.06
68.86
70.55
70.10
70.26
GSM8K
86.00
86.25
87.06
86.95
87.53
87.56
87.43
AIME-2024
8.85
8.33
9.48
10.73
11.25
10.83
11.46
AIME-2025
3.44
5.21
6.56
5.83
6.46
7.19
6.98
AIME-2026
5.52
6.15
6.25
6.88
7.81
7.29
6.77
Appendix
Table 9: Benchmark-level comparison for Granite-4.1-8B → Phi-4-mini-instruct.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict
full
top-16
top-128
MATH500
65.84
66.16
73.31
73.73
74.91
75.25
75.31
GSM8K
85.63
83.97
88.03
87.75
88.64
88.81
88.87
AIME-2024
7.60
8.23
10.94
11.67
14.37
14.17
14.17
AIME-2025
3.85
4.38
8.33
9.06
9.27
8.75
9.69
AIME-2026
6.98
5.31
7.29
5.73
6.25
6.46
6.46
Appendix
Table 10: Benchmark-level comparison for Granite-4.1-8B → Qwen2.5-7B-Base.
Benchmark
Span-MSE weight λ
0
0.25
0.5
0.75
1
1.25
1.5
MATH500
52.74
52.38
51.21
51.19
50.50
50.18
49.95
GSM8K
80.21
79.96
79.35
79.65
79.11
78.94
78.75
AIME-2024
4.69
5.10
5.21
5.21
4.58
4.06
5.62
AIME-2025
1.67
1.77
1.56
1.56
1.35
1.25
1.25
AIME-2026
0.94
0.42
0.62
0.31
0.42
0.52
0.21
Appendix
Table 11: Benchmark-level span-MSE weight sweep for Qwen2.5-7B-Instruct → Llama-3.2-3B-Instruct.
Benchmark
Span-MSE weight λ
0
0.25
0.5
0.75
1
1.25
1.5
MATH500
70.55
71.04
70.01
70.44
69.59
70.47
70.08
GSM8K
87.53
87.64
87.31
87.57
87.09
87.31
87.46
AIME-2024
11.25
10.31
10.42
10.31
9.69
10.31
10.00
AIME-2025
6.46
6.67
6.46
6.77
7.08
6.25
6.46
AIME-2026
7.81
5.94
6.77
7.19
6.88
6.77
6.35
Appendix
Table 12: Benchmark-level span-MSE weight sweep for Granite-4.1-8B → Phi-4-mini-instruct.
Benchmark
Span-MSE weight λ
0
0.25
0.5
0.75
1
1.25
1.5
MATH500
74.91
75.45
74.97
74.84
74.54
74.72
74.81
GSM8K
88.64
88.79
88.75
88.77
88.83
88.95
88.77
AIME-2024
14.37
15.42
13.02
14.27
12.71
12.19
12.92
AIME-2025
9.27
8.02
9.27
9.69
8.23
8.02
8.02
AIME-2026
6.25
6.15
5.94
6.15
5.83
5.83
5.52
Appendix
Table 13: Benchmark-level span-MSE weight sweep for Granite-4.1-8B → Qwen2.5-7B-Base.
Benchmark
Base
Teacher
MATH500
80.81
95.88
GSM8K
91.51
94.24
AIME-2024
22.71
60.94
AIME-2025
18.65
51.98
AIME-2026
12.81
65.52
AMC23
64.53
90.00
Appendix
Table 14: Benchmark-level scores of the undistilled student and frozen teacher for Qwen3-235B-A22B-Instruct-2507 → Granite-4.1-8B.
Benchmark
ULD
Extended ULD
GOLD
SimCT
Strict top-16
MATH500
81.31
87.29
85.91
86.08
89.75
GSM8K
91.59
93.06
93.12
93.15
93.69
AIME-2024
23.23
37.19
35.52
37.40
43.02
AIME-2025
21.15
26.98
27.19
29.17
35.10
AIME-2026
13.33
30.10
23.02
27.29
38.44
AMC23
67.34
72.66
69.30
74.22
76.80
Appendix
Table 15: Benchmark-level distillation results for Qwen3-235B-A22B-Instruct-2507 → Granite-4.1-8B. Bold marks the best Math, Code, and Full averages.
Method
Seen
Unseen
Overall
SR ↑
Turns ↓
SR ↑
Turns ↓
SR ↑
Turns ↓
Base
12.50
46.60
8.85
47.28
10.68
46.94
Teacher
29.69
41.07
29.43
41.22
29.56
41.15
ULD
1.56
49.61
1.30
49.74
1.43
49.68
Extended ULD
25.52
41.95
23.44
42.61
24.48
42.28
GOLD
3.65
49.30
1.56
49.58
2.61
49.44
Appendix
Table 16: ALFWorld results for Qwen3-4B-Instruct-2507 → Llama-3.2-3B-Instruct. Bold marks the highest SR and lowest Turns among distilled models.
University of Chinese Academy of Sciences · 2Chinese Information Processing Laboratory Institute of Software, Chinese Academy of Sciences 2University of Chinese Academy of Sciences, Beijing, China · 3Xiaohongshu