On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
Figures & tables
Figure 1: (a) The key obstacle in cross-tokenizer distillation is that the teacher and student tokenizers segment the same text differently ( LS=LT ), while their distinct vocabularies yield logits of different dimensionalities ( ∣VS∣=∣VT∣ ). (b) Our method DVLM can make the teacher match the student’s tokenization schema and produce logits with the same dimensionality through the PTS input format and HPA mechanism, thereby achieving distributional alignment.
Figure 2: (a) The Dual-Vocabulary Language Model takes the Parallel-Tokenized Sequence as input, with its token visibility controlled by Hybrid-Prefix Attention, while only the student-vocabulary LM head is trained. (b) In the same way, the trained DVLM teacher can extract output distributions with exactly the same dimensionality as those of the student model for OPD.
Method
Mathematics
Coding
Average
GSM8K
MATH-500
AMC23
HumanEval+
CRUXEval
LCBench
ACC.
ACC.
pass@1
pass@1
pass@1
pass@1
Qwen3-4B
87.79
68.40
38.28
82.93
41.81
26.93
57.69
Llama-3.2-1B
44.05
27.00
4.68
31.10
14.25
1.47
20.43
SFT
45.06(0.4)
26.47(0.3)
4.84(0.3)
33.13(0.9)
8.38(0.4)
1.61(0.2)
19.92
GRPO
43.34(0.2)
27.00(0.4)
6.71(0.2)
31.30(0.9)
14.42(0.5)
1.93(0.1)
20.78
Table 1: Main results of our cross-tokenizer distillation method on six reasoning benchmarks with two student models. The best results are highlighted in dark yellow and bold, while the second-best cross-tokenizer distillation results are marked in light yellow. ‘ Δ w/ base’ denotes the improvement of our method over the student base model. Values in () indicate the variation across multiple runs.
Method
Mathematics
Coding
Average
GSM8K
MATH-500
AMC23
HumanEval+
CRUXEval
LCBench
ACC.
ACC.
pass@1
pass@1
pass@1
pass@1
Llama-3.2-1B
44.05
27.00
4.68
31.10
14.25
1.47
20.43
DVLM
45.89
29.93
9.06
36.99
16.44
2.86
23.53
w/o HPA
38.36
18.40
2.58
23.17
12.43
1.02
15.99
Table 2: Results of ablation studies. ‘w/o HPA’, ‘w/o T.Prefix’, ‘w/o Re-token.’, ‘w/o LM Head’, and ‘w/o L.Extract’ respectively mean standard causal attention, the DVLM teacher input with only re-tokenized tokens, the DVLM teacher input with only the teacher sequence, OPD with the original teacher model, and utilizing span-level distillation to replace logit extraction in our method.
Figure 3: Results of in-domain evaluation of two student models on two datasets.
Figure 4: The effectiveness of the DVLM teacher is evaluated by comparing both the training loss (the shaded areas represent the error ranges across multiple runs) and the performances on four understanding tasks against those of the original teacher model.
Figure 5: Sample difficulty scaling on four benchmarks with Llama-3.2-1B.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Stage I
Stage II
Training Objective
next-token prediction
forward KL
Learning Rate
1×10−5
2×10−6
Batch Size
48
64
Gradient Accumulation steps
1
1
Training Epochs
1
2
Learning-rate Schedule
Cosine
Cosine
Appendix
Table 3: Training configurations for the new LM head and the GOLD baseline.
Settings
Model Interface
new LM head
Teacher Model
Qwen3-4B
Prediction Format
single-token generation
Shot
0
Batch Size
32
Numerical Precision
BF16
Appendix
Table 4: Evaluation settings for the four understanding tasks.
Task
Batch Size
Shots
Max. Gen. Tokens
Precision
Temp.
Metrics
GSM8K
32
0
1024
BF16
0.0
ACC.
MATH-500
16
0
1024
BF16
0.0
ACC.
HumanEval+
32
0
1024
BF16
0.0
pass@ 1 (n= 1 )
CRUXEval
32
2
1024
BF16
0.2
pass@ 1 (n= 1 )
Appendix
Table 5: Evaluation settings for the LM-Evaluation-Harness framework.
Task
Batch Size
Shots
Max. Gen. Tokens
Precision
Temp.
Metrics
AMC23
16
0
1024
BF16
0.0
pass@ 1 (n= 16 )
LiveCodeBench v5
16
0
1024
BF16
0.0
pass@ 1 (n= 8 )
Appendix
Table 6: Evaluation settings for the Evalchemy framework.
Stage
Components
Runs
Trainable Parameters
Days
Stage I
Llama head DVLM
1×3
0.33 B
0.08 (single GPU)
OLMo head DVLM
1×3
0.26 B
0.09 (single GPU)
Stage II
Llama head DVLM → Llama
1×3
1.24 B
0.23 (four GPUs)
OLMo head DVLM → OLMo
1×3
1.48 B
0.27 (four GPUs)
Appendix
Table 7: Trainable parameters and training time of our method.
Teacher head
mean token loss
BPB ↓
Original Qwen head
1.94
≈0.37
DVLM (Llama vocabulary)
1.93
≈0.37
DVLM (OLMo vocabulary)
1.90
≈0.36
Appendix
Table 8: Held-out language-modeling evaluation on unseen AI-generated text.
Teacher head
Math-500
Humaneval+
Original Qwen head
68.40
82.93
DVLM (Llama vocabulary)
66.80
82.32
DVLM (OLMo vocabulary)
68.60
81.71
Appendix
Table 9: The evaluation of the DVLM teachers on Math-500 and Humaneval+.
Method
Mathematics
Coding
Average
GSM8K
MATH-500
AMC23
HumanEval+
CRUXEval
LCBench
ACC.
ACC.
pass@1
pass@1
pass@1
pass@1
OLMo-2-1B
69.37
21.00
5.00
23.17
12.31
0.27
21.85
Qwen3-4B
87.79
68.40
38.28
82.93
41.81
26.93
57.69
DVLM
70.15
22.87
8.39
27.44
17.94
1.50
24.72
OLMo-2-7B
80.36
35.40
7.57
25.60
16.06
6.90
28.65
Appendix
Table 10: Comparison of same-tokenizer and cross-tokenizer distillation for OLMo-2-1B-Instruct.
Method
Mathematics
Coding
Average
MATH-500
AMC23
HumanEval+
LCBench
ACC.
pass@1
pass@1
pass@1
Qwen3-4B
68.40
38.28
82.93
26.93
54.14
Llama-3.2-1B
27.00
4.68
31.10
1.47
16.06
TokAlign
24.40
2.82
18.29
0.57
11.52
DVLM
29.93
9.06
36.99
2.86
19.71
Appendix
Table 11: Comparison of the TokAlign transfer method and our DVLM method.
Open-weight language models from different families exhibit complementary capabilities, motivating their consolidation into a compact student through on-policy distillation (OPD). However, full-vocabulary OPD typically assumes a shared tokenizer, while existing cross-tokenizer methods may discard teacher probability mass or assign it to student tokens with unrelated content. We introduce Byte-Prefix Marginalization (BPM), which re-expresses the teacher's next-token distribution over the student vocabulary in a shared byte space. Specifically, BPM assigns each teacher token's probability to the longest student token whose byte representation is a prefix of the teacher token's bytes, aggregates mass mapped to the same student token, and places otherwise unmatched mass in an explicit residual category. This produces a vocabulary-complete, byte-aligned, and mass-preserving target for dense OPD. The target exactly recovers the teacher-induced byte-prefix marginal when the relevant prefix does not span multiple teacher tokens (a condition satisfied at more than 99% of training positions) and uses a mass-preserving, chain-factorized lower bound otherwise. Across Qwen3-32B, GLM-Z1-9B-0414, and MiniMax-M2.7 as teachers, BPM consistently outperforms current cross-tokenizer methods on six mathematics and programming benchmarks, improving six-benchmark avg@8 by 3.7-6.6 points over the strongest baselines.
Hao Wang, Kun Yuan, Wenlin Zhong +4
University of Chinese Academy of Sciences · KwaiKAT Team · Zhejiang University +1
On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models. However, existing OPD distillation methods require teacher and student models to share the same tokenizer, restricting the applicability of OPD within the model series. Current mainstream practice typically employs Supervised Fine-Tuning (SFT) on teacher-generated responses for cross-tokenizer distillation, which fails to capture the rich knowledge embedded in the teacher's probability distribution. In this work, we enable the standard on-policy distillation method to operate across model families, ensuring that high-fidelity token-level signals can propagate across different tokenizers with a precise token-mapping algorithm. Extensive experiments show that cross-tokenizer OPD is significantly more compute-efficient than baselines on various benchmarks. Our results unlock a broader range of teacher-student pairs for OPD, opening up new avenues for adapting and enhancing interactions between LLMs.
Yifan Niu, Han Xiao, Dongyi Liu +4
The Hong Kong University of Science and Technology (Guangzhou) · Tencent · The Hong Kong University of Science and Technology
On-policy distillation (OPD) offers superior capacity transfer by supervising student-sampled trajectories with dense token-level signals. To furnish high-quality supervision sources and thereby elevate the performance frontier of distillation, an intuitive direction is to infuse privileged information to either teacher or student itself. However, this additional input induces a potential failure mode we dub privilege illusion: a pattern that conflates the transferable capability gap that students are meant to close, and the information asymmetry gap that can only be mimicked but never replicated. This issue is further amplified by the inherent non-uniformity of token-level supervision, where only a small subset of tokens carries pivotal capability-bearing signals. To this end, we propose DOPD, an advantage-aware dual distillation paradigm that dynamically routes token-level supervision between privileged teacher and privileged student policies based on their advantage gap and relative probabilities. Each token receives supervision of different strength, objective, and strategy from either teacher or student itself, which transfers credible capability while simultaneously receiving auxiliary signals, to alleviate privilege illusion. Extensive experiments on both large language model (LLM) and vision-language model (VLM) settings demonstrate that DOPD consistently outperforms Vanilla OPD and other counterparts. Further results on stability, robustness, continual learning, and out-of-distribution tasks validate its superiority.