cs.CLOct 1, 2026

Counting and Min-Cost Encoding for Tokenization in Large Language Models

Authors: Shuming Shi, Xiang Zhang, Hao Yu, Wenbo Fei, Changjian Wang, Zhan Wang, Guoqing Pang, Guangye Yu, +2 more

Organizations: Mashang Consumer Finance Co., Ltd., China · National-Mathematics Artificial Intelligence Institute in Chongqing (NMAII), China

Abstract

Mainstream large language models rely on a tokenizer to encode text into a token sequence. Different tokenizers may yield token sequences of substantially different lengths for the same text. With a fixed model architecture, shorter token sequences correspond to lower inference time. We propose a tokenizer training approach named Counting and Filtering (CNF) and a text encoding algorithm called Min-Cost Encoding (MCE). MCE defines a cost function over a text segment, and determines the best segmentation by globally minimizing the overall segmentation cost. CNF builds a raw vocabulary by directly counting valid substrings, and then constructs the final vocabulary through a filtering step based on actual token usage when segmenting the training corpus with MCE. The CNF-MCE conbination offers several advantages over BPE, including higher token efficiency, greater scalability, and lower dependency. Across six text categories and two vocabulary-size groups, CNF-MCE consistently achieves better compression than the evaluated BPE tokenizers. With a 250K vocabulary, CNF-MCE increases compression rate by 26% and 30% on English web text over the o200k_base and qwen250k tokenizers. Experiments scaling the vocabulary to 1M entries on English web text demonstrate sustained improvements over BPE, with a token efficiency improvement of over 60% and vocabulary utilization rising from 52.9% to 96.9%. The MCE algorithm does not depend on a merge list (as in BPE) or token probability (as in UnigramLM), making it applicable to a wide range of vocabularies, including those built from BPE, UnigramLM, CNF, and others. Language models trained from scratch at the 1.8B and 8B scales achieve comparable average performance to models using the BPE tokenizers across 11 benchmarks. These results demonstrate that CNF-MCE can improve token efficiency significantly while maintaining competitive downstream performance.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Tokenization with Split Trees

    May 21, 2026Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4Subword TokenizationUnigram Tokenizer

  2. More Than Words: Compositional Tokenization for Efficient Language Models

    Oct 4, 2026Yuval Reif, Guy Kaplan, Roy SchwartzTokenizerLanguage Modeling

  3. Incremental BPE Tokenization

    May 29, 2026Shenghu Jiang, Ruihao GongByte-Pair EncodingTokenizer