Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation
Organizations: School of Computer Science, East China Normal University, Shanghai, China · School of Computer Science, Fudan University, Shanghai, China
Abstract
Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields more accepted tokens than in subword Transformers.
Figures & tables
| CUTE (%; ) | OCRBench (%; ) | TextVQA (%; ) | ||||
|---|---|---|---|---|---|---|
| Size | Subword | Byte | Subword | Byte | Subword | Byte |
| 200M | 69.89 | 99.19 | 27.10 | 35.10 | 40.62 | 45.08 |
| 400M | 70.78 | 99.18 | 27.40 | 37.40 | 43.82 | 48.37 |
| 700M | 73.04 | 99.52 | 29.90 | 34.50 | 45.34 | 50.72 |
| 1B | 76.44 | 99.56 | 31.50 | 39.40 | 47.50 | 52.36 |
| 3B | 80.83 | 99.93 | 34.70 | 42.40 | 50.76 | 55.39 |
| Model | ||||
|---|---|---|---|---|
| Ours Byte -1B | 23.3 | 56.6 | 7.43 | 0.822 |
| BLT-1B | 25.7 | 59.2 | 6.62 | 0.730 |
| H-Net 1-stage XL | 22.5 | 55.8 | 7.53 | 0.878 |
| H-Net 2-stage XL | 23.1 | 56.2 | 7.48 | 0.867 |
| Ours Subword -1B | 0 1.5 | 16.3 | 2.35 | 0.218 |
| Llama-3.1-8B | 0 6.5 | 27.2 | 1.18 | 0.090 |
| 200M draft model | 50M draft model | |||
| Byte | Subword | Byte | Subword | |
| Acceptance rate (%) | 90.6 | 74.1 | 84.6 | 62.0 |
| Accepted tokens / forward | 9.527 | 2.831 | 5.471 | 1.618 |
| Byte / Subword ratio | ||||
| Bytes / accepted token | 1.000 | 3.364 | 1.000 | 3.159 |
| Accepted bytes / forward | 9.527 | 9.523 | 5.471 | 5.111 |
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
| Hyperparameter | Byte | Subword |
|---|---|---|
| Vocabulary size | 256 | 49,152 |
| Convolution kernel width | 16 | 4 |
| TST bag size / duration | 4 / first 30% of updates | |
| Hash -gram lengths | ||
| Hash tables slots per table | ||
| Size | Layers | Width | FFN | Head dim. GQA/GDN | GQA q/kv heads | GDN qk/v heads | Byte | Subword |
|---|---|---|---|---|---|---|---|---|
| 200M | 16 | 1,024 | 2,560 | 128/64 | 8/2 | 8/16 | 0.207 | 0.256 |
| 400M | 20 | 1,280 | 3,200 | 128/64 | 10/2 | 10/20 | 0.374 | 0.436 |
| 700M | 24 | 1,536 | 3,840 | 128/64 | 12/2 | 12/24 | 0.644 | 0.719 |
| 1B | 28 | 1,792 | 4,480 | 128/64 | 14/2 | 14/28 | 1.022 | 1.109 |
| 3B | 40 | 2,560 | 6,400 | 256/128 | 10/2 | 10/20 | 3.181 | 3.305 |
| Size | Layers | Expert width | Byte | Byte | Subword | Subword |
|---|---|---|---|---|---|---|
| 200M | 16 | 256 | 0.180 | 0 0.885 | 0.230 | 0 0.985 |
| 400M | 20 | 384 | 0.395 | 0 2.047 | 0.457 | 0 2.172 |
| 700M | 24 | 384 | 0.604 | 0 2.983 | 0.679 | 0 3.132 |
| 1B | 28 | 512 | 1.044 | 0 5.361 | 1.131 | 0 5.535 |
| 3B | 40 | 768 | 3.146 | 16.358 | 3.269 | 16.607 |
| Hyperparameter | Byte | Subword |
|---|---|---|
| Context length (tokens) | 8,192 | 2,048 |
| Global batch (tokens) | 1,048,576 | 262,144 |
| Base learning rate | 0.005 | |
| Optimizers | Muon (matrices), AdamW (other parameters) | |
| Scheduler | WSD, 1% warmup, 20% 1-sqrt decay | |
| AdamW | ||
| Recipe | |||||
|---|---|---|---|---|---|
| Byte raw | 1054.9 | 0.42293 | 13257 | 0.50811 | 0.87358 |
| Byte hash | 300.19 | 0.36138 | 53129 | 0.57329 | 0.86467 |
| Byte TST | 2547.6 | 0.46305 | 3691.2 | 0.42706 | 0.83219 |
| Byte TST+hash | 550.58 | 0.39097 | 9343.3 | 0.46963 | 0.81537 |
| Subword raw | 2410.6 | 0.46131 | 10460 | 0.48918 | 0.85900 |
| Subword hash | 3810 | 0.48746 | 10338 | 0.48703 | 0.86609 |
| 200M | 400M | 700M | 1B | 3B | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | Subword | Byte | Subword | Byte | Subword | Byte | Subword | Byte | Subword | Byte |
| ARC | 54.27 | 53.67 | 55.56 | 58.58 | 59.84 | 60.49 | 61.40 | 62.83 | 65.81 | 68.36 |
| MMLU | 29.68 | 29.57 | 31.09 | 31.50 | 32.83 | 33.63 | 34.02 | 35.64 | 37.08 | 38.63 |
| CSQA | 53.24 | 52.50 | 57.08 | 54.71 | 57.08 | 60.11 | 62.24 | 61.34 | 64.62 | 69.12 |
| HellaSwag | 47.24 | 51.32 | 52.48 | 56.91 | 57.14 | 62.17 | 60.94 | 65.56 | 67.84 | 72.89 |
| WinoGrande | 51.78 | 52.72 | 53.83 | 55.17 | 55.88 | 59.19 | 58.01 | 58.96 | 60.14 | 62.98 |
| 200M | 400M | 700M | 1B | 3B | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Task | Subword | Byte | Subword | Byte | Subword | Byte | Subword | Byte | Subword | Byte |
| Spelling | 51.70 | 99.80 | 56.10 | 99.80 | 60.00 | 99.90 | 68.60 | 99.80 | 81.80 | 100.00 |
| Spelling (random) | 99.10 | 99.90 | 99.00 | 99.90 | 99.40 | 99.00 | 99.60 | 100.00 | 99.80 | 100.00 |
| Inverse spelling | 59.90 | 100.00 | 63.20 | 100.00 | 73.90 | 100.00 | 78.30 | 100.00 | 82.00 | 100.00 |
| Inverse spelling (random) | 87.30 | 100.00 | 93.00 | 99.90 | 91.20 | 100.00 | 95.60 | 100.00 | 97.60 | 100.00 |
| Character containment | 67.20 | 99.20 | 70.60 | 99.60 | 72.60 | 100.00 | 74.00 | 99.80 | 78.60 | 100.00 |
| Task | Final question | Reference answer |
|---|---|---|
| Spelling (random) | Question: Spell out the word "bdj". | b d j |
| Inverse spelling | Question: Write the word "t h e". | the |
| Character insertion | Question: Add an "l" after every "t" in "little". | litltlle |
| Word swapping | Question: Swap "is" and "fun" in "It is fun.". | It fun is. |
| Benchmark / task | Question | Reference answer |
|---|---|---|
| OCRBench Irregular text | What is written in the image? | PARLIAMENT |
| OCRBench Handwriting | What is written in the image? | strictures |
| OCRBench Non-semantic text | What is written in the image? | ntishgcwi |
| TextVQA Scene text QA | What does the small white text spell? | copenhagen |
| TextVQA Scene text QA | What is the license plate number of this vehicle? | aj52uyv |
| TextVQA Scene text QA | What is the phone number listed to rent this billboard? | 648-3004 |
| OCR & document understanding | Knowledge-intensive visual reasoning | General visual understanding | ||||||||
| Size | Model | DocVQA | OCRBench | TextVQA | ChartQA | ScienceQA-IMG | MMBench | MME | GQA | POPE |
| 200M | Subword | 25.98 | 27.10 | 40.62 | 24.04 | 42.49 | 0 8.16 | 1164.20 | 48.26 | 85.37 |
| 200M | Byte | 28.56 | 35.10 | 45.08 | 26.56 | 37.88 | 0 2.06 | 1187.69 | 46.26 | 84.65 |
| 400M | Subword | 29.77 | 27.40 | 43.82 | 28.00 | 46.01 | 23.11 | 1245.94 | 49.67 | 82.28 |
| 400M | Byte | 32.64 | 37.40 | 48.37 | 28.72 | 48.19 | 13.48 | 1206.10 | 49.13 | 82.95 |
| 700M | Subword | 31.35 | 29.90 | 45.34 | 30.08 | 50.17 | 29.98 | 1332.91 | 51.87 | 85.39 |
| OCR & document understanding | Knowledge-intensive visual reasoning | General visual understanding | ||||||||
| Size | Model | DocVQA | OCRBench | TextVQA | ChartQA | ScienceQA-IMG | MMBench | MME | GQA | POPE |
| 200M | Subword | 25.87 | 26.00 | 39.79 | 23.04 | 50.37 | 0 2.66 | 1219.18 | 46.10 | 82.26 |
| 200M | Byte | 25.85 | 32.90 | 43.22 | 23.28 | 39.96 | 0 1.37 | 1154.95 | 43.81 | 81.36 |
| 400M | Subword | 28.68 | 27.10 | 44.01 | 23.48 | 44.97 | 0 6.35 | 1265.12 | 47.06 | 84.78 |
| 400M | Byte | 30.84 | 35.90 | 49.61 | 27.32 | 47.30 | 18.21 | 1292.77 | 47.61 | 84.75 |
| 700M | Subword | 34.26 | 33.40 | 49.28 | 34.48 | 52.85 | 35.57 | 1335.96 | 53.65 | 85.17 |
| Model | Near-zero BPB | Tail BPB | ||
|---|---|---|---|---|
| Ours Byte -1B | -0.946 | 1.898 | ||
| BLT-1B | -0.952 | 1.855 | ||
| H-Net 1-stage XL | -0.916 | 1.867 | ||
| H-Net 2-stage XL | -0.918 | 1.899 | ||
| Ours Subword -1B | -0.612 | 1.684 | ||
| Llama-3.1-8B | -0.696 | 1.407 |