cs.CLOct 8, 2026

Latent Core Tokenizer: Compress, but Meaningfully

Authors: Felermino D. M. A. Ali, Millicent Ochieng, Ogbemi Ekwejunor-Etchie, Ade Famoti, Jacki O'Neill, Debjit Paul

Organizations: Microsoft Research Africa · Microsoft Research Accelerator · Microsoft Research India

Abstract

Tokenizers are commonly optimized for compression, but a compact vocabulary does not necessarily distribute its capacity evenly across languages. We introduce the Latent Core Tokenizer (LCT), a language-agnostic approach that separates structural discovery from vocabulary construction. LCT uses Minimum Description Length, entropy-based boundary signals, and morphotactic constraints to identify reusable linguistic units before constructing a shared vocabulary. Across 104 languages with a 200K-token vocabulary, LCT achieves lower fertility and higher MorphScore than BPE, Unigram, and parity-aware BPE, while maintaining comparable cross-lingual disparity in tokenization cost. Across four multilingual downstream benchmarks, LCT improves aggregate score by 1.48, 1.83, and 2.00 points over BPE, Unigram, and parity-aware BPE, respectively. Our findings show that compression alone does not predict representation quality and highlight the importance of morphology-driven structural discovery and how frequency is used to allocate the final vocabulary across languages.

Figures & tables

Appendix figures & tables10 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Tokenization with Split Trees

    May 21, 2026Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4Token EfficiencySubword Tokenization

  2. LangMAP: A Language-Adaptive Approach to Tokenization

    Jun 22, 2026Clara Meister, Suchir Salhan, Andrzej Szablewski +3Multilingual Language ModelsTokenizer Adaptation