Dual-Vocabulary Language Model for Cross-Tokenizer Distillation
Organizations: East China Normal University · Shanghai Innovation Institute · Tsinghua University
Abstract
On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.
Figures & tables
| Method | Mathematics | Coding | Average | ||||
| GSM8K | MATH-500 | AMC23 | HumanEval+ | CRUXEval | LCBench | ||
| ACC. | ACC. | pass@1 | pass@1 | pass@1 | pass@1 | ||
| Qwen3-4B | |||||||
| Llama-3.2-1B | |||||||
| SFT | |||||||
| GRPO | |||||||
| Method | Mathematics | Coding | Average | ||||
| GSM8K | MATH-500 | AMC23 | HumanEval+ | CRUXEval | LCBench | ||
| ACC. | ACC. | pass@1 | pass@1 | pass@1 | pass@1 | ||
| Llama-3.2-1B | |||||||
| DVLM | |||||||
| w/o HPA | |||||||
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Stage I | Stage II | |
| Training Objective | next-token prediction | forward KL |
| Learning Rate | ||
| Batch Size | ||
| Gradient Accumulation steps | ||
| Training Epochs | ||
| Learning-rate Schedule | Cosine | Cosine |
| Settings | |
| Model Interface | new LM head |
| Teacher Model | Qwen3-4B |
| Prediction Format | single-token generation |
| Shot | |
| Batch Size | |
| Numerical Precision | BF16 |
| Task | Batch Size | Shots | Max. Gen. Tokens | Precision | Temp. | Metrics |
| GSM8K | BF16 | ACC. | ||||
| MATH-500 | BF16 | ACC. | ||||
| HumanEval+ | BF16 | pass@ (n= ) | ||||
| CRUXEval | BF16 | pass@ (n= ) |
| Task | Batch Size | Shots | Max. Gen. Tokens | Precision | Temp. | Metrics |
| AMC23 | BF16 | pass@ (n= ) | ||||
| LiveCodeBench v5 | BF16 | pass@ (n= ) |
| Stage | Components | Runs | Trainable Parameters | Days |
| Stage I | Llama head DVLM | B | (single GPU) | |
| OLMo head DVLM | B | (single GPU) | ||
| Stage II | Llama head DVLM Llama | B | (four GPUs) | |
| OLMo head DVLM OLMo | B | (four GPUs) |
| Teacher head | mean token loss | BPB |
| Original Qwen head | ||
| DVLM (Llama vocabulary) | ||
| DVLM (OLMo vocabulary) |
| Teacher head | Math-500 | Humaneval+ |
| Original Qwen head | ||
| DVLM (Llama vocabulary) | ||
| DVLM (OLMo vocabulary) |
| Method | Mathematics | Coding | Average | ||||
| GSM8K | MATH-500 | AMC23 | HumanEval+ | CRUXEval | LCBench | ||
| ACC. | ACC. | pass@1 | pass@1 | pass@1 | pass@1 | ||
| OLMo-2-1B | |||||||
| Qwen3-4B | |||||||
| DVLM | |||||||
| OLMo-2-7B | |||||||
| Method | Mathematics | Coding | Average | ||
| MATH-500 | AMC23 | HumanEval+ | LCBench | ||
| ACC. | pass@1 | pass@1 | pass@1 | ||
| Qwen3-4B | |||||
| Llama-3.2-1B | |||||
| TokAlign | |||||
| DVLM | |||||