Byte-level language models such as the Byte Latent Transformer (BLT) group bytes into patches and run their large global model once per patch. BLT starts a patch where a small model's next-byte entropy is high, so global compute goes where the next byte is hard to predict. We show that this rule has a systematic blind spot: positions whose type is predictable but whose value must be computed, such as the number after "=" in a worked math solution. Under tight patch budgets, entropy-triggered layouts skip these positions, and accuracy on them collapses. In Meta's BLT-1B with patch starts on 10% of bytes, the entropy rule puts a patch start at 16% of the computed results in GSM8K solutions and gets 19.0% of them exactly right; a boundary after each "=" at the same patch count gets 51.8%, and entropy combined with a label-free boundary-dependence signal gets 67.1% (default layout at 26% of bytes: 76.8%). The gap survives adapting BLT-1B to the budget with low-rank fine-tuning (32.9% vs 72.7%, three runs per rule, paired p < 1e-200) and grows with model size in byte models trained from scratch at a 10% budget: at 1M, 12M and 50M parameters, boundary dependence beats entropy on final answers by -1.6, +10.1 and +19.8 points, and at 50M it gets 35.9% of computed results against 13.9% (3 seeds each). BLT's entropy-jump rule helps neither target at 50M. The entropy trigger of Scratchpad Patching is likewise indistinguishable from random scratchpads on final answers (5.6% vs 5.9%, 5 seeds), while answer-start scratchpads give 38.1%. The effect is specific to computed values: copies and lookups gain little, and values the model cannot compute gain nothing. Boundary dependence, the rise in the model's own loss when a patch start is removed, measured per two-byte context, finds these positions without labels: combined with entropy it beats the hand-written rule on computed results.
Figures & tables
Layout
Patch rate
Results covered
Computed results
Final answers
default (BLT-1B’s own)
0.255
79%
76.8%
73.5%
entropy@15
0.157
48%
47.0%
65.0%
results@15 (hand-written)
0.157
100%
66.1%
67.0%
dep@15
0.153
100%
77.3%
40.8%
entdep@15
0.157
100%
76.5%
67.3%
entropy@10
0.106
16%
19.0%
51.4%
Table 1
Figure 1: BLT-1B patch starts in one GSM8K solution at a 10% budget, with equal patch counts for the two layouts. Bars mark patch starts and shading marks in-line computed results. Entropy starts no patch at any of the six results; entropy plus dependence starts one before each.
Budget
Entropy
Entropy jump (BLT’s monotonic rule)
Dependence (label-free)
Hand-written results rule
10%
7.1% / 10.4% / 1.654
7.0% / 64.1% / 1.639
17.1% / 67.4% / 1.678
24.7% / 72.5% / 1.640
15%
8.3% / 32.5% / 1.552
not run
14.8% / 68.2% / 1.638
22.7% / 74.8% / 1.528
20%
8.4% / 44.0% / 1.475
not run
17.4% / 68.2% / 1.577
23.9% / 75.1% / 1.451
Table 3
Parameters
Rule
Seeds
Final answers (by seed)
Computed results (by seed)
Bits per byte
Eval patch rate
1.1M
entropy
3
7.0% (8.8, 4.8, 7.4)
4.4% (3.6, 4.9, 4.8)
2.093
11.2%
1.1M
dependence
2
5.4% (6.2, 4.5)
4.2% (4.3, 4.1)
2.196
10.9%
1.1M
hand-written
3
45.5% (54.8, 25.3, 56.2)
4.6% (3.9, 4.1, 5.9)
2.109
12.1%
11.7M
entropy
3
46.6% (42.0, 49.1, 48.8)
5.8% (4.8, 6.4, 6.3)
1.657
11.2%
11.7M
dependence
3
56.7% (58.5, 53.9, 57.6)
7.7% (7.2, 5.6, 10.3)
1.673
10.9%
11.7M
hand-written
3
76.3% (76.4, 73.9, 78.5)
12.9% (14.9, 10.3, 13.5)
1.645
12.1%
Table 4
Figure 2: Exact match against parameter count for models trained from scratch at a 10% budget (small markers: seeds; large markers: means). The advantage of the label-free dependence rule over BLT’s entropy rule grows with size, and computed results separate only once the models can compute. BLT’s jump rule and entropy at twice the budget were run at 50M only.
Trained and tested under
Untrained
Runs
Mean
entropy
19.0%
32.8, 32.9, 33.0
32.9%
results (hand-written)
51.8%
62.4, 62.9, 62.7
62.7%
entdep (label-free)
67.1%
72.9, 70.9, 74.4
72.7%
Table 6
Target
Default
Entropy
Entdep
Forced
Kind
GSM8K computed results
76.8%
19.0%
67.1%
51.8%
computed, skill present
Python identifiers repeating a nearby name
83.3%
81.7%
81.3%
83.1%
copy
Proof-step conclusions, generated logic
69.1%
46.3%
41.5%
49.4%
rule lookup
Copied values, generated program traces
82.5%
58.3%
62.3%
61.0%
copy
Computed values, program traces
20.5%
10.0%
9.0%
9.7%
computed, skill absent
Table 7
Trigger
Seeds
Final answers
By seed
answer starts (math syntax)
5
38.1%
49.2, 31.4, 8.3, 49.2, 52.4
entropy (the paper’s trigger)
5
5.6%
7.3, 4.8, 4.2, 5.6, 6.1
random positions
5
5.9%
6.1, 5.6, 4.5, 5.2, 8.0
none
3
6.1%
6.2, 6.8, 5.2
denser fixed patches, same compute (8-byte)
5
7.8%
6.7, 12.3, 7.6, 5.2, 7.1
Table 8
Figure 3: BLT-1B exact match on in-line computed results (left) and final answers (right) against forward compute per byte, relative to the default layout, for layouts at 10% and 15% of bytes. Dependence alone is a lookup and pays no entropy-model cost.
Figure 4: Models trained at 10-25% patch budgets (D = 128): exact match against forward compute per byte, relative to entropy at 25%. Points are seed means; bars span the seeds.