cs.CLOct 4, 2026

More Than Words: Compositional Tokenization for Efficient Language Models

Authors: Yuval Reif, Guy Kaplan, Roy Schwartz

Organizations: The Hebrew University of Jerusalem

Abstract

Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Less Is More: Reducing Token Counts Without Compromising Performance

    Jun 18, 2025Gyeongje Cho, Yeonkyoung So, Sangmin Lee +1Subword Tokenization

  2. Counting and Min-Cost Encoding for Tokenization in Large Language Models

    Oct 1, 2026Shuming Shi, Xiang Zhang, Hao Yu +7Counting

  3. Tokenization with Split Trees

    May 21, 2026Craig W. Schmidt, Michael Krumdick, Adam Wiemerslage +4Subword TokenizationUnigram Tokenizer