More Than Words: Compositional Tokenization for Efficient Language Models
Organizations: The Hebrew University of Jerusalem
Abstract
Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.
Figures & tables
| Text | BPE encoding # tokens | CoBPE encoding (1 compound token; base = ground ) |
|---|---|---|
| ground | ground 1 | — |
| the ground | the | _ground 2 | determiner=the |
| The ground | _The | _ground 2 | leading-space, determiner=the, capitalize-determiner |
| The Ground | _The | _Ground 2 | leading-space, capitalize, determiner=the, capitalize-determiner |
| In the ground | In | _the | _ground 3 | preposition=in, capitalize-preposition, determiner=the |
| in the ground, | in | _the | _ground | , 4 | preposition=in, determiner=the, punct-suffix=, |
| 24-layer (780M) | 28-layer (1.3B) | |||||
| Category | Task | BPE | SuperBPE ( ) | CoBPE | BPE | CoBPE |
| Knowledge | ARC-Easy (MC) | 68.3 | 69.7 | 74.1 | 73.8 | 76.1 |
| ARC-Challenge (MC) | 40.3 | 41.8 | 42.7 | 43.3 | 47.4 | |
| Jeopardy (EM) | 5.9 | 5.7 | 12.0 | 11.8 | 18.0 | |
| MMLU (MC) | 25.2 | 25.3 | 25.2 | 25.0 | 25.5 | |
| OpenBookQA (MC) | 39.8 | 42.6 | 42.2 | 42.8 | 39.2 | |
| Model | CORE | Tokens to target | Hours to target | Speedup |
|---|---|---|---|---|
| BPE | 25.91 | 5.84B | 7.92h | 1.00 |
| CoBPE | 26.00 | 4.74B | 6.45h | 1.23 |
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
| 24-layer (780M) | 28-layer (1.3B) | |||||
|---|---|---|---|---|---|---|
| Category | Task | BPE | SuperBPE ( ) | CoBPE | BPE | CoBPE |
| Reasoning / symbolic | Repeat-Copy-Logic (EM) | 0.0 | 3.1 | 0.0 | 0.0 | 3.1 |
| Reasoning / symbolic | GSM8K (EM) | 2.7 | 2.4 | 2.0 | 2.7 | 1.3 |
| Reading | HotpotQA (EM) | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Coding | HumanEval (P@10) | 0.0 | 0.0 | 0.0 | 0.3 | 0.0 |
| Coding | MBPP (P@10) | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| Group | Values used in the reported experiments |
|---|---|
| Leading space | Binary indicator: absent, present. |
| Preposition | No modifier, by , at , of , to , in , on , with , for , from . |
| Determiner | No modifier, the , a , an , my , your , his , her , our , their , its . |
| Preposition capitalization | Binary indicator: lowercase preposition, titlecased preposition. |
| Determiner capitalization | Binary indicator: lowercase determiner, titlecased determiner. |
| Base capitalization | Binary indicator: lowercase base, titlecased base. |
| Group | Values used in the reported experiments |
|---|---|
| Prefix punctuation | No modifier, " , ’ , ‘ , ( , [ , {, - , --- , ‘ , ’ , “ , ” |
| Suffix punctuation | No modifier, " , ’ , ‘ , ) , ] , }, . , ! , ? , , , ; , : , - , --- , ’ , possessive endings ’s , s’ , ’s , s’ |
| 24 layers | 28 layers | |
| Parameters | 780M | 1.3B |
| Layers | 24 | 28 |
| Hidden size | 1536 | 2048 |
| Intermediate size | 6144 | 6144 |
| Head dimension | 128 | 128 |
| Query heads | 12 | 16 |
| Setting | Value |
|---|---|
| Backbone scale | 24 layers, hidden size 1536, intermediate size 6144, head dimension 128. |
| Attention heads | 12 query heads and 12 KV heads. |
| Vocabulary / context | 32k vocabulary, maximum sequence length 2048. |
| Hardware | 8 L40S GPUs. |
| Batching | Device batch size 16; total batch size 1.05M tokens. |
| Target param:data ratio | 8.0 for BPE; 6.5 for the CoBPE variant |
| Language | Determiners | Adpositions | Attachment direction |
|---|---|---|---|
| Spanish | 17 | 14 | before |
| German | 24 | 16 | before |
| Hindi | 5 | 19 + 15 compound | before / after |
| Indonesian | 4 | 9 | before |