cs.AISep 27, 2026

Dual-Vocabulary Language Model for Cross-Tokenizer Distillation

Authors: Kedi Chen, Chen Lin, Yutao Sun, Wei Zhang

Organizations: East China Normal University · Shanghai Innovation Institute · Tsinghua University

Abstract

On-policy distillation (OPD) bridges teacher supervision and student behavior, but different teacher-student tokenizers introduce misalignment in both input tokenization (#1) and output logits (#2). Existing approaches address the former by matching same-text spans or converting tokens to bytes, often losing fine-grained token information or disrupting the native-token paradigm, while for the latter, strategies such as ranking, padding, or key-token selection retain only shared logit dimensions, resulting in much distribution loss. In this paper, we propose Dual-Vocabulary Language Model (DVLM), which replaces the teacher's LM head with a new student-vocabulary projection head and obtains full-dimensional student logits (for #2). To support student tokens (for #1), it takes a Parallel-Tokenized Sequence (PTS) as input, which concatenates the original teacher-tokenized sequence and a re-tokenized sequence formed by independently converting each student token into a teacher-token group. To avoid inference inconsistency with the original teacher tokens, the Hybrid-Prefix Attention (HPA) further restricts re-tokenized groups to their corresponding teacher prefix and uses its last state as the aggregation of the original student-token representation for projection into the student vocabulary space. Similarly, via the combined use of PTS and HPA, the DVLM teacher can provide distribution-aligned supervision with the student's input-tokenization and output-logit during OPD. Experimental results demonstrate that our DVLM teacher has a similar converged loss as the original teacher model and enables student models to improve performance across six reasoning tasks.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization

    Jul 24, 2026Hao Wang, Kun Yuan, Wenlin Zhong +4

  2. Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families

    Jun 8, 2026Yifan Niu, Han Xiao, Dongyi Liu +4Teacher-Student DistillationTokenizer

  3. DOPD: Dual On-policy Distillation

    Jun 29, 2026Xinlei Yu, Gen Li, Qingyi Si +13Efficient On-Policy DistillationPrivileged Context