On-policy distillation (OPD) transfers knowledge between language models through teacher supervision on student-generated trajectories. With different tokenizers, a single teacher token may require multiple student tokens to generate, creating intermediate states where the event is entered but not yet completed. Existing cross-tokenizer methods align tokens or text spans to construct comparable prediction targets. We study a complementary problem after partial generation: once the student produces a prefix of a teacher token, multiple next tokens may complete the same remaining bytes, but the teacher only specifies the required completion rather than how probability should be divided among these valid continuations. We introduce Event-Set Completion Distillation (ESCD), which complements cross-tokenizer probability alignment with completion-set supervision. ESCD aggregates prefix-related teacher events and supervises the total probability of byte-compatible one-step student completions, avoiding tokenizer-dependent probability splits among individual tokens. The method reuses student trajectories and predictions, requiring neither additional rollouts nor changes to the student vocabulary. Experiments demonstrate consistent gains in mathematics, code, and scientific reasoning across model families and tokenizers, extending to large-scale MoE distillation from a 1T teacher to a 35B student. Local analyses show that retaining completion sets better matches the reference supervision, while one-step completion covers over 99% of observed compatible teacher mass after partial event entry in the studied tokenizer pairs. These findings support event entry and event completion as complementary supervision targets for cross-tokenizer knowledge transfer. Code will be released on GitHub.
Figures & tables
Figure 1: Cross-tokenizer distillation results across model scales. Left: Teacher-normalized pass@ n scores on six mathematics and code benchmarks, with Qwen3.5-2B distilled from Qwen3-32B or GLM-Z1-9B. Recovery is the Student score divided by the corresponding teacher score, expressed as a percentage; 100% indicates matching teacher performance. We use n=8 for mathematics and n=2 for code. Right: Absolute scores on FrontierScience Olympiad and PHYRD-40 for two large-scale heterogeneous MoE pairs. Together, these results show consistent improvements across task domains, model families, and scales.
Method
FrontierScience Olympiad-100
PHYRD-40
Qwen3.5-397B-A17B †
70.0
79.8
Qwen3-30B-A3B-Thinking (base)
48.0
43.0
BPM
48.0
46.6
Ours
51.0
49.6
Δ vs. best baseline
+3.0
+3.0
Kimi-K2.7-Code †
75.0
83.6
Table 2: Cross-tokenizer OPD on heterogeneous MoE pairs with 397B and 1T teachers. In the Kimi setting, BPM and ESCD use the same SFT-initialized checkpoint due to direct OPD instability.
Teacher → Student
Root
Single
Event
Δ
Qwen3 → Qwen3.5
0.7383
0.4323
0.8495
+0.4172
GLM-Z1 → Qwen3.5
0.8848
0.7493
0.9341
+0.1848
Table 3: Local target comparison on contexts satisfying the COUF diagnostic criteria. We report reference-energy-weighted state-space COUF; Δ denotes Event minus Single.
Teacher → Student
Observed
1-step
Deeper
Qwen3 → Qwen3.5
81.83%
99.01%
0.99%
GLM-Z1 → Qwen3.5
83.29%
99.43%
0.57%
Table 4: Full-vocabulary child-supervision availability. Observed is relative to branchable teacher mass; 1-step and Deeper characterize representative residuals within visited mass.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Recognized termination
Preferred bridge stop
Qwen3 / Qwen3.5
model-declared EOS / stop tokens
<|im_end|>
GLM-Z1
<|endoftext|> and configured stops
<|endoftext|>
Kimi-K2.7-Code
[EOS] , <|im_end|>
<|im_end|>
Appendix
Table 5: Model-specific termination handling. The bridge stop is the token used to represent a normally completed response after cross-tokenizer retokenization.
Setting
Small-scale
Large-scale
Teacher
Qwen3-32B / GLM-Z1-9B
Qwen3.5-397B-A17B
Kimi-K2.7-Code
Student
Qwen3.5-2B
Qwen3-30B-A3B-Thinking
Qwen3.6-35B-A3B
Training framework
Slime
Slime
Slime
Student optimization
Megatron-LM
Megatron-LM
Megatron-LM
Rollout and Teacher serving
SGLang
SGLang
SGLang
Training data
DAPO-Math + TACO
SciDeriv
SciDeriv
Appendix
Table 6: Training configurations for OPD experiments. GPU totals include Teacher serving, Student optimization, and Student rollout, excluding separately scheduled evaluation jobs.
Task
Mathematics
Code
Source
DAPO-Math
TACO
Prompts
10,000
10,000
Appendix
Table 7: Training data for small-scale distillation.
Task
Mathematics
Code
Source
DAPO-Math
TACO
Prompts
10,000
10,000
Appendix
Table 7: Training data for small-scale distillation.
Task type
Prompts
Conclusion assessment
10,000
Conclusion reconstruction
7,577
Total
17,577
Appendix
Table 8: Composition of SciDeriv for large-scale distillation.
Evaluation setting
Context
Max output
Temp .
Top- p
Top- k
Samples/problem
AIME24/25/26
32,768
27,648
0.6
0.95
20
8
HMMT-26
32,768
27,648
0.6
0.95
20
8
LiveCodeBench
32,768
27,648
0.6
0.95
20
2
TACO
32,768
27,648
0.6
0.95
20
2
FrontierScience Olympiad
262,144
256,000
1.0
0.95
20
1
PHYRD-40
262,144
256,000
1.0
0.95
20
1
Appendix
Table 9: Generation settings for the controlled and large-scale experiments. Context and output limits are measured in tokens; the context limit includes both the input prompt and generated response.
Tokenizer pair
∣VT∣
∣VS∣
M
CT
CS
J
Qwen3 → Qwen3.5
151,669
248,077
131,612
86.78%
53.05%
49.08%
Qwen3.5 → Qwen3
248,077
151,669
131,612
53.05%
86.78%
49.08%
GLM-Z1 → Qwen3.5
151,343
248,077
142,628
94.24%
57.49%
55.54%
GLM-Z1 → Qwen3
151,343
151,669
123,118
81.35%
81.18%
68.44%
Kimi-K2.7 → Qwen3.6
163,840
248,077
121,757
74.31%
49.08%
41.96%
Appendix
Table 10: Static shared-vocabulary overlap under the deterministic one-to-one canonical-surface mapping. Vocabulary sizes exclude checkpoint padding.
Teacher token
qT(v∣x)
Student tokenization
Residual after Br
Bread
0.4
Br , ead
ead
Breads
0.2
Br , eads
eads
Break
0.3
Br , eak
eak
Cat
0.1
Cat
—
Appendix
Table 12: Illustrative teacher predictions and student realizations. Residuals are shown for individual teacher tokens after Br , before ESCD aggregation; Cat is incompatible with this action.
Teacher → Student
Contexts
Matched mass
Both 1-step
Qwen3 → Qwen3.5
32
0.8745%
0
GLM-Z1 → Qwen3.5
37
0.0662%
0
Appendix
Table 13: Available comparisons between visited and alternative child states. The last column counts contexts where both branches admit one-step completion of the same teacher event.
Teacher → Student
Root
Counterfactual
Visited
Qwen3 → Qwen3.5
0.999031
0.999982
0.999032
GLM-Z1 → Qwen3.5
0.998486
0.999997
0.998489
Appendix
Table 14: State-space COUF when the counterfactual branch may use deeper completion. The teacher event and root supervision are matched, but completion depth is not controlled.
Pair
mg
bg
Mass share
Child visited
One-step given visit
Qwen3 → Qwen3.5
1
1
0.0026%
34.5351%
96.0012%
1
>1
0.1089%
70.6500%
71.9885%
>1
1
88.2125%
80.9535%
99.9991%
>1
>1
11.6760%
88.6045%
92.4118%
Overall
100.0000%
81.8344%
99.0135%
GLM-Z1 → Qwen3.5
1
1
2.3143%
88.8103%
99.9993%
Appendix
Table 15: Root event structure and completion availability from the full teacher vocabulary after exact-shared and invalid-token filtering. Statistics use aligned, non-whitespace positions in 16 frozen student trajectories per pair, grouped by mg and bg .
teacher → Student
Events
Set size
Full prob.
Crossing prob.
Qwen3 → Qwen3.5
48
35.28
0.8846
2.10%
GLM-Z1 → Qwen3.5
287
14.05
0.9412
0.63%
Appendix
Table 16: Boundary-crossing ambiguity on materialized natural-child events. “Set size” is weighted by teacher source mass; “Crossing prob.” is the fraction of completion-event probability assigned to tokens that extend beyond the residual boundary.
Teacher → Student
Projection
Energy ratio
Cosine
Qwen3 → Qwen3.5
1.0566
1.1550
0.9832
GLM-Z1 → Qwen3.5
1.0302
1.0857
0.9887
Appendix
Table 17: Local target-gradient sensitivity to removing boundary-crossing completions.
Teacher → Student
Position upper bound
Observed
Top-1 observed
Qwen3 → Qwen3.5
93.47%
81.83%
71.55%
GLM-Z1 → Qwen3.5
95.49%
83.29%
76.23%
Appendix
Table 18: Compatible-child visitation under full-vocabulary analysis. All columns are relative to the same pooled branchable teacher mass. The position upper bound counts all branchable mass at positions with any compatible sampled action; Top-1 ranks branches by teacher source mass.
Teacher → Student
1-step
Deeper
Qwen3 → Qwen3.5
99.01%
0.99%
GLM-Z1 → Qwen3.5
99.43%
0.57%
Appendix
Table 19: One-step completion availability for representative residual constraints, weighted by teacher mass and conditional on a compatible child being visited.
Step
Mean tokens
Median tokens
Max tokens
1
7,698
7,550
14,470
16
10,587
7,605
50,590
20
3,022
2,518
8,099
26
247
24
1,818
30
56
8
566
32
8
6
20
Appendix
Table 20: Response-length collapse in the truncation-filtered Kimi-to-Qwen BPM run. Statistics describe the 16 responses retained at each listed step, rounded to the nearest token.
On-policy distillation (OPD) is a standard tool for transferring teacher behavior to a smaller student, but it implicitly assumes that teacher and student predictions are comparable token by token, an assumption that fails whenever the two models tokenize the same text differently. Under heterogeneous tokenizers, exact shared-token matching silently discards a large fraction of the teacher signal at precisely the positions where vocabularies disagree. We propose \textbf{\underline{Sim}ple \underline{C}ross-\underline{T}okenizer OPD (SimCT)}, which restores this signal by enlarging the supervision space: alongside shared tokens, SimCT compares teacher and student over short multi-token continuations that both tokenizers can realize, leaving the OPD loss form itself unchanged. We show that these units are the finest jointly tokenizable supervision interface, and that coarser alternatives remove teacher-student distinctions that are useful for on-policy learning. Across three heterogeneous teacher-student pairs on mathematical reasoning and code-generation benchmarks, SimCT shows consistent gains over shared-vocabulary OPD and representative cross-tokenizer baselines, with ablations confirming that the improvements come from recovering supervision discarded by exact shared-token matching. Code is available at \href{https://github.com/sunjie279/SimCT-}{https://github.com/sunjie279/SimCT-}.
Jie Sun, Mao Zheng, Mingyang Song +6
University of Science and Technology of China · Shanghai Innovation Institute · Large Language Model Department, Tencent +1
On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models. However, existing OPD distillation methods require teacher and student models to share the same tokenizer, restricting the applicability of OPD within the model series. Current mainstream practice typically employs Supervised Fine-Tuning (SFT) on teacher-generated responses for cross-tokenizer distillation, which fails to capture the rich knowledge embedded in the teacher's probability distribution. In this work, we enable the standard on-policy distillation method to operate across model families, ensuring that high-fidelity token-level signals can propagate across different tokenizers with a precise token-mapping algorithm. Extensive experiments show that cross-tokenizer OPD is significantly more compute-efficient than baselines on various benchmarks. Our results unlock a broader range of teacher-student pairs for OPD, opening up new avenues for adapting and enhancing interactions between LLMs.
Yifan Niu, Han Xiao, Dongyi Liu +4
The Hong Kong University of Science and Technology (Guangzhou) · Tencent · The Hong Kong University of Science and Technology
Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while existing cross-tokenizer methods may discard teacher probability mass or assign it to student tokens with unrelated content. We introduce Byte-Prefix Marginalization (BPM), which re-expresses the teacher's next-token distribution over the student vocabulary in a shared byte space. Specifically, BPM assigns each teacher token's probability to the longest student token whose byte representation is a prefix of the teacher token's bytes, aggregates mass mapped to the same student token, and places otherwise unmatched mass in an explicit residual category. This produces a vocabulary-complete, byte-aligned, and mass-preserving target for dense OPD. The target exactly recovers the teacher-induced byte-prefix marginal when the relevant prefix does not span multiple teacher tokens (a condition satisfied at more than 99% of training positions) and uses a mass-preserving, chain-factorized lower bound otherwise. Across Qwen3-32B, GLM-Z1-9B-0414, and MiniMax-M2.7 as teachers, BPM consistently outperforms current cross-tokenizer methods on six mathematics and programming benchmarks, improving six-benchmark avg@8 by 3.7-6.6 points over the strongest baselines.
Hao Wang, Kun Yuan, Wenlin Zhong +4
University of Chinese Academy of Sciences · KwaiKAT Team · Zhejiang University +1