Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch. BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict. We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution. Under tight patch budgets, entropy-triggered layouts skip these positions, and accuracy on them collapses. In Meta's BLT-1B with patch starts on 10% of bytes, the entropy rule puts a patch start at 16% of the computed results in GSM8K solutions and gets 19.0% of them exactly right; a boundary after each "=" at the same patch count gets 51.8%, and entropy combined with a label-free boundary-dependence signal gets 67.1% (default layout at 26% of bytes: 76.8%). The gap survives adapting BLT-1B to the budget with low-rank fine-tuning (32.9% vs 72.7%, three runs per rule, paired p < 1e-200) and grows with model size in byte models trained from scratch at a 10% budget: at 1M, 12M and 50M parameters, boundary dependence beats entropy on final answers by -1.6, +10.1 and +19.8 points, and at 50M it gets 35.9% of computed results against 13.9% (3 seeds each). BLT's entropy-jump rule helps neither target at 50M. The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%. The effect is specific to computed values: copies and lookups gain little, and values the model cannot compute gain nothing. Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels: combined with entropy it beats the hand-written rule on computed results.
Figures & tables
Layout
Patch rate
Results covered
Computed results
Final answers
default (BLT-1B’s own)
0.255
79%
76.8%
73.5%
entropy@15
0.157
48%
47.0%
65.0%
results@15 (hand-written)
0.157
100%
66.1%
67.0%
dep@15
0.153
100%
77.3%
40.8%
entdep@15
0.157
100%
76.5%
67.3%
entropy@10
0.106
16%
19.0%
51.4%
Table 1
Figure 1: BLT-1B patch starts in one GSM8K solution at a 10% budget, with equal patch counts for the two layouts. Bars mark patch starts and shading marks in-line computed results. Entropy starts no patch at any of the six results; entropy plus dependence starts one before each.
Budget
Entropy
Entropy jump (BLT’s monotonic rule)
Dependence (label-free)
Hand-written results rule
10%
7.1% / 10.4% / 1.654
7.0% / 64.1% / 1.639
17.1% / 67.4% / 1.678
24.7% / 72.5% / 1.640
15%
8.3% / 32.5% / 1.552
not run
14.8% / 68.2% / 1.638
22.7% / 74.8% / 1.528
20%
8.4% / 44.0% / 1.475
not run
17.4% / 68.2% / 1.577
23.9% / 75.1% / 1.451
Table 3
Parameters
Rule
Seeds
Final answers (by seed)
Computed results (by seed)
Bits per byte
Eval patch rate
1.1M
entropy
3
7.0% (8.8, 4.8, 7.4)
4.4% (3.6, 4.9, 4.8)
2.093
11.2%
1.1M
dependence
2
5.4% (6.2, 4.5)
4.2% (4.3, 4.1)
2.196
10.9%
1.1M
hand-written
3
45.5% (54.8, 25.3, 56.2)
4.6% (3.9, 4.1, 5.9)
2.109
12.1%
11.7M
entropy
3
46.6% (42.0, 49.1, 48.8)
5.8% (4.8, 6.4, 6.3)
1.657
11.2%
11.7M
dependence
3
56.7% (58.5, 53.9, 57.6)
7.7% (7.2, 5.6, 10.3)
1.673
10.9%
11.7M
hand-written
3
76.3% (76.4, 73.9, 78.5)
12.9% (14.9, 10.3, 13.5)
1.645
12.1%
Table 4
Figure 2: Exact match against parameter count for models trained from scratch at a 10% budget (small markers: seeds; large markers: means). The advantage of the label-free dependence rule over BLT’s entropy rule grows with size, and computed results separate only once the models can compute. BLT’s jump rule and entropy at twice the budget were run at 50M only.
Trained and tested under
Untrained
Runs
Mean
entropy
19.0%
32.8, 32.9, 33.0
32.9%
results (hand-written)
51.8%
62.4, 62.9, 62.7
62.7%
entdep (label-free)
67.1%
72.9, 70.9, 74.4
72.7%
Table 6
Target
Default
Entropy
Entdep
Forced
Kind
GSM8K computed results
76.8%
19.0%
67.1%
51.8%
computed, skill present
Python identifiers repeating a nearby name
83.3%
81.7%
81.3%
83.1%
copy
Proof-step conclusions, generated logic
69.1%
46.3%
41.5%
49.4%
rule lookup
Copied values, generated program traces
82.5%
58.3%
62.3%
61.0%
copy
Computed values, program traces
20.5%
10.0%
9.0%
9.7%
computed, skill absent
Table 7
Trigger
Seeds
Final answers
By seed
answer starts (math syntax)
5
38.1%
49.2, 31.4, 8.3, 49.2, 52.4
entropy (the paper’s trigger)
5
5.6%
7.3, 4.8, 4.2, 5.6, 6.1
random positions
5
5.9%
6.1, 5.6, 4.5, 5.2, 8.0
none
3
6.1%
6.2, 6.8, 5.2
denser fixed patches, same compute (8-byte)
5
7.8%
6.7, 12.3, 7.6, 5.2, 7.1
Table 8
Figure 3: BLT-1B exact match on in-line computed results (left) and final answers (right) against forward compute per byte, relative to the default layout, for layouts at 10% and 15% of bytes. Dependence alone is a lookup and pays no entropy-model cost.
Figure 4: Models trained at 10-25% patch budgets (D = 128): exact match against forward compute per byte, relative to entropy at 25%. Points are seed means; bars span the seeds.
Tokenizer-free language models eliminate the tokenizer step of the language modeling pipeline by operating directly on bytes; patch-based variants further aggregate contiguous byte spans into patches for efficiency. However, the average patch size chosen at the model design stage governs a tight trade-off: larger patches reduce compute and KV-cache footprint, but degrade modeling quality. We trace this trade-off to patch lag: until a patch is fully observed, byte predictions within it must rely on a stale representation from the previous patch to preserve causality; this lag widens as patches grow larger. We introduce Scratchpad Patching (SP), which inserts transient scratchpads inside each patch to aggregate the bytes seen so far and refresh patch-level context for subsequent predictions. SP triggers scratchpads using next-byte prediction entropy, selectively allocating compute to information-dense regions and enabling post-hoc adjustment of inference-time compute. Across experiments on natural language and code, SP improves model quality at the same patch size; for example, even at 16 bytes per patch, SP-augmented models match or closely approach the byte-level baseline on downstream evaluations while using a 16× smaller KV cache over patches and 3-4× less inference compute.
Recent byte-level language models (LMs) match the performance of token-level models without relying on subword vocabularies, yet their utility is limited by slow, byte-by-byte autoregressive generation. We address this bottleneck in the Byte Latent Transformer (BLT) through new training and generation techniques. First, we introduce BLT Diffusion (BLT-D), a new model and our fastest BLT variant, trained with an auxiliary block-wise diffusion objective alongside the standard next-byte prediction loss. This enables an inference procedure that generates multiple bytes in parallel per decoding step, substantially reducing the number of forward passes required to generate a sequence. Second, we propose two extensions inspired by speculative decoding that trade some of this speed for higher generation quality: BLT Self-speculation (BLT-S), in which BLT's local decoder continues generating past its normal patch boundaries to draft bytes, which are then verified with a single full-model forward pass; and BLT Diffusion+Verification (BLT-DV), which augments BLT-D with an autoregressive verification step after diffusion-based generation. All methods may achieve an estimated memory-bandwidth cost over 50% lower than BLT on generation tasks. Each approach offers its own unique advantages, together removing key barriers to the practical use of byte-level LMs.
Julie Kallini, Artidoro Pagnoni, Tomasz Limisiewicz +5
FAIR at Meta · Stanford University · University of Washington
Tokenizer-free language models remove the inductive bias of fixed tokenizers by modeling text directly as bytes, but the resulting longer sequences substantially increase computation and eliminate explicit text abstractions. We ask whether this additional computation can be useful, and whether standard Transformers can learn the abstractions that tokenization provides. We study these questions on Transformers without specialized tokenization-related architectures. With token-superposition training and hash embeddings, byte Transformers consistently outperform subword Transformers as model size scales. We further find that byte Transformers build local text abstractions as external tokenizers: a set of segmentation-like positions are used to collect local context representations, and restricting up to 25% of intermediate layers to these local representations preserves downstream performance. Finally, these learned structures induce highly non-uniform generation difficulty, with uncertainty concentrated near local structure boundaries; exploiting them for speculative decoding yields 3.4× more accepted tokens than in subword Transformers.
Jie Wang, Shiwei Luo, Qi Zhang +1
School of Computer Science, East China Normal University, Shanghai, China · School of Computer Science, Fudan University, Shanghai, China