Byte Language Models: Scaling, Emergent Abstractions, and Information Allocation
Authors: Jie Wang, Shiwei Luo, Qi Zhang, Yuanbin Wu
Organizations: School of Computer Science, East China Normal University, Shanghai, China · School of Computer Science, Fudan University, Shanghai, China
Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to 25% of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields 3.4× more accepted tokens than in subword Transformers.
Figures & tables
Figure 1: Pareto frontiers of compute-optimal models. (a) Byte has lower optimal BPB L⋆ at the same optimal parameter count N⋆ , while (b) Subword has lower L⋆ at the same training compute C . (c) Byte optima generally allocate more source bytes per parameter. Markers indicate fitted optima at the evaluated compute budgets (see details in Appendix C ).
CUTE (%; ↑ )
OCRBench (%; ↑ )
TextVQA (%; ↑ )
Size
Subword
Byte
Subword
Byte
Subword
Byte
200M
69.89
99.19
27.10
35.10
40.62
45.08
400M
70.78
99.18
27.40
37.40
43.82
48.37
700M
73.04
99.52
29.90
34.50
45.34
50.72
1B
76.44
99.56
31.50
39.40
47.50
52.36
3B
80.83
99.93
34.70
42.40
50.76
55.39
Table 1: Byte consistently outperforms Subword on fine-grained tasks across model scales.
Figure 2: Compute scaling with shared depth and sparse experts. Four-pass Subword closes most of the BPB gap to Byte with higher QA RC, while sparse MoE lowers Byte BPB with less activated computation. Markers are observations and curves are fits.
Figure 3: Byte -1B attention maps during copying. Rows denote copied target bytes and columns denote source bytes from the previous occurrence; each map shows target-to-source attention. All test strings have internal spaces removed. Blue boxes indicate original word chunks, and blue dashed lines mark their final bytes. In intermediate layers, attention concentrates near chunk boundaries, and multiple subsequent target bytes repeatedly attend to the same aggregated positions, indicating the formation of meaningful local representations.
Figure 4: Performance retention under frozen layer intervention. Each cell shows downstream performance relative to the uncompressed baseline for a compressed layer interval G=[a,b] . Byte results are shown in the top row and Subword results in the bottom row. 100% indicates performance at or above the uncompressed baseline.
Figure 5: Token-span averaging largely closes the BPB distribution gap between Byte and Subword . The BPB distribution of Byte has more near-zero and high-BPB mass. All panels are byte-weighted. The dotted curve averages fixed Byte losses over Subword token spans, preserving mean BPB on complete spans (Appendix F.3 ).
Model
FX(10−3)
FX(0.1)
SX(4)
SX(8)
Ours Byte -1B
23.3
56.6
7.43
0.822
BLT-1B
25.7
59.2
6.62
0.730
H-Net 1-stage XL
22.5
55.8
7.53
0.878
H-Net 2-stage XL
23.1
56.2
7.48
0.867
Ours Subword -1B
0 1.5
16.3
2.35
0.218
Llama-3.1-8B
0 6.5
27.2
1.18
0.090
Table 2: Byte models place more probability mass near zero and at high BPB in their native distributions (Section 4.1 ). Values are percentages of retained text bytes. Figure 11 shows the full BPB distributions.
200M draft model
50M draft model
Byte
Subword
Byte
Subword
Acceptance rate (%)
90.6
74.1
84.6
62.0
Accepted tokens / forward
9.527
2.831
5.471
1.618
Byte / Subword ratio
3.366×
3.381×
Bytes / accepted token
1.000
3.364
1.000
3.159
Accepted bytes / forward
9.527
9.523
5.471
5.111
Table 3: Higher byte acceptance rates offset shorter token spans at both draft sizes, giving comparable accepted text per target forward pass. Targets are same-family 1B models, evaluated by an oracle simulation on the given evaluation text (Appendix G.1 ). Forward denotes a simulated target verification pass; counts include only accepted draft tokens (Appendix G.2 ).
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Hyperparameter
Byte
Subword
Vocabulary size
256
49,152
Convolution kernel width
16
4
TST bag size / duration
4 / first 30% of updates
Hash n -gram lengths G
{5,6,7,8}
{2}
Hash tables × slots per table
4×32,768
1×131,072
Appendix
Table 4: Architectural differences between Byte and Subword . TST and hash settings apply when the corresponding components are enabled.
Size
Layers
Width d
FFN
Head dim. GQA/GDN
GQA q/kv heads
GDN qk/v heads
Byte N
Subword N
200M
16
1,024
2,560
128/64
8/2
8/16
0.207
0.256
400M
20
1,280
3,200
128/64
10/2
10/20
0.374
0.436
700M
24
1,536
3,840
128/64
12/2
12/24
0.644
0.719
1B
28
1,792
4,480
128/64
14/2
14/28
1.022
1.109
3B
40
2,560
6,400
256/128
10/2
10/20
3.181
3.305
Appendix
Table 5: Dense configurations and active parameter counts N (billions). In the head-count columns, kv indicates equal key and value counts; qk indicates equal query and key counts.
Size
Layers
Expert width
Byte N
Byte Ntotal
Subword N
Subword Ntotal
200M
16
256
0.180
0 0.885
0.230
0 0.985
400M
20
384
0.395
0 2.047
0.457
0 2.172
700M
24
384
0.604
0 2.983
0.679
0 3.132
1B
28
512
1.044
0 5.361
1.131
0 5.535
3B
40
768
3.146
16.358
3.269
16.607
Appendix
Table 6: MoE configurations and parameter counts (billions).
Hyperparameter
Byte
Subword
Context length (tokens)
8,192
2,048
Global batch (tokens)
1,048,576
262,144
Base learning rate η0
0.005
Optimizers
Muon (matrices), AdamW (other parameters)
Scheduler
WSD, 1% warmup, 20% 1-sqrt decay
AdamW (β1,β2)
(0.8,0.95)
Appendix
Table 7: Default pretraining hyperparameters for the main model suite.
Figure 6: TST+hash lowers fitted optimal BPB in both families at shared budgets, with larger gains for bytes. Dots are measurements; curves are fixed-budget slices of L(M,V) , and stars are fitted optima. Colors denote training FLOP budgets.
Figure 7: At matched N,V , Byte TST+hash improves over raw Subword as source-byte exposure grows. Here, N excludes the output embedding to align parameter counts between the two model families. Panels (a–b) show fitted BPB, and (c) shows their difference. Stars are fitted compute-optimal allocations; black dots are training configurations, and dashed contours mark 1019 FLOPs. Surfaces cover only the sampled parameter–data regions.
Recipe
A
α
B
β
E
Byte raw
1054.9
0.42293
13257
0.50811
0.87358
Byte hash
300.19
0.36138
53129
0.57329
0.86467
Byte TST
2547.6
0.46305
3691.2
0.42706
0.83219
Byte TST+hash
550.58
0.39097
9343.3
0.46963
0.81537
Subword raw
2410.6
0.46131
10460
0.48918
0.85900
Subword hash
3810
0.48746
10338
0.48703
0.86609
Appendix
Table 8: Fitted parameters of L(M,V)=E+AM−α+BV−β (Section 2.2 ).
200M
400M
700M
1B
3B
Task
Subword
Byte
Subword
Byte
Subword
Byte
Subword
Byte
Subword
Byte
ARC
54.27
53.67
55.56
58.58
59.84
60.49
61.40
62.83
65.81
68.36
MMLU
29.68
29.57
31.09
31.50
32.83
33.63
34.02
35.64
37.08
38.63
CSQA
53.24
52.50
57.08
54.71
57.08
60.11
62.24
61.34
64.62
69.12
HellaSwag
47.24
51.32
52.48
56.91
57.14
62.17
60.94
65.56
67.84
72.89
WinoGrande
51.78
52.72
53.83
55.17
55.88
59.19
58.01
58.96
60.14
62.98
Appendix
Table 9: QA RC scores (%; ↑ ) for our models.
200M
400M
700M
1B
3B
Task
Subword
Byte
Subword
Byte
Subword
Byte
Subword
Byte
Subword
Byte
Spelling
51.70
99.80
56.10
99.80
60.00
99.90
68.60
99.80
81.80
100.00
Spelling (random)
99.10
99.90
99.00
99.90
99.40
99.00
99.60
100.00
99.80
100.00
Inverse spelling
59.90
100.00
63.20
100.00
73.90
100.00
78.30
100.00
82.00
100.00
Inverse spelling (random)
87.30
100.00
93.00
99.90
91.20
100.00
95.60
100.00
97.60
100.00
Character containment
67.20
99.20
70.60
99.60
72.60
100.00
74.00
99.80
78.60
100.00
Appendix
Table 10: CUTE scores (%; ↑ ) for our main models.
Task
Final question
Reference answer
Spelling (random)
Question: Spell out the word "bdj".
b d j
Inverse spelling
Question: Write the word "t h e".
the
Character insertion
Question: Add an "l" after every "t" in "little".
litltlle
Word swapping
Question: Swap "is" and "fun" in "It is fun.".
It fun is.
Appendix
Table 11: CUTE evaluation prompt and task examples. The spelling prompt is shown in full; the additional examples show only the final question. Reference answers are displayed separately from the model input.
Benchmark / task
Question
Reference answer
OCRBench Irregular text
What is written in the image?
PARLIAMENT
OCRBench Handwriting
What is written in the image?
strictures
OCRBench Non-semantic text
What is written in the image?
ntishgcwi
TextVQA Scene text QA
What does the small white text spell?
copenhagen
TextVQA Scene text QA
What is the license plate number of this vehicle?
aj52uyv
TextVQA Scene text QA
What is the phone number listed to rent this billboard?
648-3004
Appendix
Table 12: OCRBench and TextVQA examples with questions and reference answers. Images are omitted.
OCR & document understanding
Knowledge-intensive visual reasoning
General visual understanding
Size
Model
DocVQA
OCRBench
TextVQA
ChartQA
ScienceQA-IMG
MMBench
MME
GQA
POPE
200M
Subword
25.98
27.10
40.62
24.04
42.49
0 8.16
1164.20
48.26
85.37
200M
Byte
28.56
35.10
45.08
26.56
37.88
0 2.06
1187.69
46.26
84.65
400M
Subword
29.77
27.40
43.82
28.00
46.01
23.11
1245.94
49.67
82.28
400M
Byte
32.64
37.40
48.37
28.72
48.19
13.48
1206.10
49.13
82.95
700M
Subword
31.35
29.90
45.34
30.08
50.17
29.98
1332.91
51.87
85.39
Appendix
Table 13: Vision–language evaluation for dense models.
OCR & document understanding
Knowledge-intensive visual reasoning
General visual understanding
Size
Model
DocVQA
OCRBench
TextVQA
ChartQA
ScienceQA-IMG
MMBench
MME
GQA
POPE
200M
Subword
25.87
26.00
39.79
23.04
50.37
0 2.66
1219.18
46.10
82.26
200M
Byte
25.85
32.90
43.22
23.28
39.96
0 1.37
1154.95
43.81
81.36
400M
Subword
28.68
27.10
44.01
23.48
44.97
0 6.35
1265.12
47.06
84.78
400M
Byte
30.84
35.90
49.61
27.32
47.30
18.21
1292.77
47.61
84.75
700M
Subword
34.26
33.40
49.28
34.48
52.85
35.57
1335.96
53.65
85.17
Appendix
Table 14: Vision–language evaluation for MoE models.
Figure 8: In the main model suite, dense Byte achieves lower BPB and higher QA RC than depth-matched dense Subword , with greater training compute. BPB and QA RC use base models, and CUTE uses SFT models. Points are observations; curves summarize trends within each family’s observed range (Appendix D ).
Figure 9: Ablations on separators: Local retrieval persists with alternative separators. Layer-24 Byte -1B maps use the copying probe in Figure 3 , with words separated by spaces, | , ; , or and . Dashed lines mark original word ends; colors share the same source-normalized attention scale.
Figure 10: Additional performance retention results under the frozen layer intervention. Each cell reports downstream performance relative to the uncompressed baseline for a compressed layer interval G=[a,b] . The axes denote the start layer a , end layer b , and interval length ∣G∣=b−a+1 . Byte results are shown in the top row and Subword results in the bottom row. 100% indicates performance at or above the uncompressed baseline.
Figure 11: Byte models show more near-zero and high-BPB mass than subword models. All panels are byte-weighted.
Model
Near-zero BPB
α
Tail BPB
β
Ours Byte -1B
[5.4×10−5,2.15]
-0.946
[3,15.5]
1.898
BLT-1B
[5.4×10−5,2.86]
-0.952
[3.5,15.5]
1.855
H-Net 1-stage XL
[4.1×10−5,5.07]
-0.916
[5.5,16]
1.867
H-Net 2-stage XL
[3.1×10−5,5.07]
-0.918
[5.5,16.5]
1.899
Ours Subword -1B
[4.0×10−4,0.164]
-0.612
[3.5,13]
1.684
Llama-3.1-8B
[1.3×10−5,0.164]
-0.696
[2.5,11.25]
1.407
Appendix
Table 15: Finite-range density fits fX∝xα near zero and fX∝e−x/β in the tail, with β in BPB. Byte exponents are nearer −1 ; tail scales are comparable across granularities.
Figure 12: Near-zero power laws and exponential tail cores describe finite BPB ranges across models. All panels use byte-weighted distributions. Dashed lines show fits, and shading marks fitting ranges. Local exponents αlocal and decay scales βlocal indicate where the approximations hold. Dotted tail segments indicate sparse data. Fitting details are in Appendix F.2 .
Figure 13: Token-span averaging reduces near-zero and high-BPB mass for both Byte -1B and BLT-1B. All panels are byte-weighted. Tokenizers yield similar distributions; uniform 4-byte grouping leaves more near-zero mass and less far-tail mass (Appendix F.3 ).
Figure 14: Tokenizers share most boundaries, consistent with their similar aggregation curves. Values are boundary Jaccard similarities (%).
Figure 15: Small drafts achieve high acceptance rates on low-BPB byte tokens. Both families use 16-layer drafts and same-family 1B targets with verification on the given evaluation text within native windows. Left: acceptance rate within each target-BPB interval. Right: target-BPB distributions split into accepted (dark) and rejected (light) tokens, with each native token weighted equally. Colors identify model families.