Language models process and generate text sequentially in token units, and the tokenizer determines how much text each inference step covers. Under standard tokenization, a short English phrase such as "On the table." is usually produced as four separate predictions for the preposition (On), article (the), noun (table), and punctuation (.), where each consumes a sequence position and adds inference cost. We introduce CoBPE, a compositional tokenization approach that represents such phrases as a lexical base token (table) attached with a small set of reusable surface modifiers, composed in embedding space at input and predicted jointly at output. In controlled pretraining from scratch at 780M and 1.3B scales, CoBPE shortens sequences by 30% and improves average downstream performance by 1.2 points relative to standard BPE under matched training compute. Our results suggest that part of what is now expressed through token sequences can instead be modeled through structured representations, opening a broad design space for more token-efficient and capable language models.
Figures & tables
Figure 1: Compositional tokenization shortens sequences compared to BPE by factoring recurring surface markers around a lexical base. Top: standard BPE represents the excerpt as a sequence of 18 tokens, while Compositional BPE ( CoBPE ) represents the same text using only 9 compound tokens built from a base token together with attached modifiers for leading space, capitalization, determiners, prepositions, and surrounding punctuation. Bottom: Encoding efficiency (bytes-per-token) of English text for BPE and CoBPE as a function of vocabulary size. Adding modifier families progressively improves compression over BPE, and the full CoBPE configuration covers roughly 1.4 × more bytes per token across vocabulary sizes.
Figure 2: A substantial share of token positions under a modern BPE tokenizer are grammatical markers. Breakdown of token positions occupied by determiners, prepositions, and punctuation in FineWeb text tokenized with GPT-4.
Table 1: Examples of CoBPE decompositions. A leading underscore in the BPE column denotes a space-prefixed token. Standard BPE assigns separate token identities to surface variants such as “ ground ” and “ Ground ”. It also represents the examples above using 1–5 sequential tokens. CoBPE instead keeps the lexical base token identity fixed and expresses variation through attached modifiers within a single structured compound token .
Figure 3: Adapting language models to CoBPE . (1) A span tokenized under CoBPE is represented as a structured compound token with a base token and attached modifiers. (2) On the input side, the model embeds this unit compositionally by summing the base embedding and the embeddings of its active modifiers. (3) The resulting sequence is processed by an otherwise standard transformer backbone. (4) On the output side, next-token prediction is factorized: the model first predicts the next base token, and then predicts the modifier values conditioned on that base.
24-layer (780M)
28-layer (1.3B)
Category
Task
BPE
SuperBPE ( t=27k )
CoBPE
BPE
CoBPE
Knowledge
ARC-Easy (MC)
68.3
69.7
74.1
73.8
76.1
ARC-Challenge (MC)
40.3
41.8
42.7
43.3
47.4
Jeopardy (EM)
5.9
5.7
12.0
11.8
18.0
MMLU (MC)
25.2
25.3
25.2
25.0
25.5
OpenBookQA (MC)
39.8
42.6
42.2
42.8
39.2
Table 2: Main downstream results under matched training-compute budgets. Models are trained for 15.6B tokens at the 24-layer, 780M scale and 26B tokens at the 28-layer, 1.3B scale. The primary comparison is BPE versus CoBPE at each scale. At the 24-layer scale, we also include SuperBPE as a secondary baseline, using the best transition point from our sweep for a 32k vocabulary. Best results per task and model size are in bold .
Model
CORE ↑
Tokens to target ↓
Hours to target ↓
Speedup ↑
BPE
25.91
5.84B
7.92h
1.00 ×
CoBPE
26.00
4.74B
6.45h
1.23 ×
Table 3: nanochat speedrun results on 8 L40S GPUs. We report CORE score at the stopping point, along with the training tokens and wall-clock hours used in each run. Times are comparable within this table, but not to the public nanochat H100 leaderboard.
Table 7
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Detailed breakdown of frequent grammatical markers. Top marker-level shares within determiners, prepositions, and punctuation under GPT-4 tokenization on FineWeb text.
24-layer (780M)
28-layer (1.3B)
Category
Task
BPE
SuperBPE ( t=27k )
CoBPE
BPE
CoBPE
Reasoning / symbolic
Repeat-Copy-Logic (EM)
0.0
3.1
0.0
0.0
3.1
Reasoning / symbolic
GSM8K (EM)
2.7
2.4
2.0
2.7
1.3
Reading
HotpotQA (EM)
0.0
0.0
0.0
0.0
0.0
Coding
HumanEval (P@10)
0.0
0.0
0.0
0.3
0.0
Coding
MBPP (P@10)
0.0
0.0
0.0
0.0
0.0
Appendix
Table 6: Results for the five tasks omitted from Table 2 . These tasks are included in the 30-task average reported in the main text.
Group
Values used in the reported experiments
Leading space
Binary indicator: absent, present.
Preposition
No modifier, by , at , of , to , in , on , with , for , from .
Determiner
No modifier, the , a , an , my , your , his , her , our , their , its .
Table 8: Exact punctuation strings supported by the CoBPE inventory. These are closed lists; punctuation outside them remains in the base-token stream.
24 layers
28 layers
Parameters
780M
1.3B
Layers
24
28
Hidden size
1536
2048
Intermediate size
6144
6144
Head dimension
128
128
Query heads
12
16
Appendix
Table 9: Exact model configurations for the main experiments.
Device batch size 16; total batch size 1.05M tokens.
Target param:data ratio
8.0 for BPE; 6.5 for the CoBPE variant
Appendix
Table 10: Exact configuration for the reported nanochat speedrun comparison. Both methods share the backbone and per-step optimization configuration. The target parameter-to-data ratio sets the training horizon for each method.
Language
Determiners
Adpositions
Attachment direction
Spanish
17
14
before
German
24
16
before
Hindi
5
19 + 15 compound
before / after
Indonesian
4
9
before
Appendix
Table 11: Language-specific modifier inventories. Counts give the retained surface forms after frequency, annotation-consistency, and direction filtering. Attachment indicates whether the markers normally occur before or after the lexical base. Hindi reports ordinary adpositions plus compound postpositions.