Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on its allowed prefix. On Qwen2.5-14B-Instruct, two fast FP8 realizations repaired to ordinary-looking accuracy still change the answers chosen by likelihood on 5.83% and 10.00% of 240 OpenBookQA items when only the text after the allowed prefix is replaced with the bf16 model's own greedy continuation. Both row-local controls, the bf16 model and a deployed FP8 matrix multiplication kernel, change none. Accuracy thus does not certify prefix invariance, and the stability criteria we analyze cannot tell realizations apart: across all 512 sign variants of two-level Strassen they stay constant while teacher-forced perplexities span a 772.4× range on the same model. We therefore construct certified realizations of two-level Strassen on bounded integer codes that quantize token rows independently, then mix and cancel exactly before rescaling, using 49 block multiplications instead of 64. Our certificate guarantees bitwise equality to a prescribed row-local classical int8 operator at the same quantization specification, so every certified realization inherits its prefix invariance. Certification thus turns realization choice into a pure cost decision: which certified realization runs can no longer change a single scored likelihood.
Figures & tables
Figure 1: A Accuracy on the original input, and answers changed (95% intervals) when the text after the allowed prefix is replaced by the bf16 model’s greedy continuation; #1, #2: two fast algorithms, each with both repairs. B Perplexity of the 512 sign variants of two-level Strassen under one FP8 schedule, sorted, and the criteria we analyze. C A certified realization and classical int8; strip: synthetic row pairs in which a later row changes an earlier output (FP8: our schedule, same coefficients). A, B: Qwen2.5-14B-Instruct.
Figure 2: A One OpenBookQA item on Qwen2.5-14B-Instruct: the answer chosen with the natural text and with the bf16 model’s greedy continuation after an illustrative cut; “produced” is scored at the outlined token. B In two-level Strassen, block sums rounded to FP8 carry row 7 into rows 1, 3 and 5 (repaired #1 only into 3), and masked attention carries those rows on to position 6.
OpenBookQA
ARC-Easy
answers changed
changed
accuracy
continuation
filler
accuracy
filler
bf16 model
34.58
0.00
0.00
84.17
0.00
deployed FP8 kernel
33.75
0.00
0.00
83.33
0.00
repaired #1
35.42
5.83
[2.92, 8.75]
7.08
[4.17, 10.42]
82.50
6.67
[3.75, 10.00]
repaired #2
32.92
10.00
[6.25, 13.75]
7.92
[4.58, 11.67]
80.83
12.08
[8.33, 16.25]
Table 1: Accuracy (%) on the original input, and answers changed (%) when the text after the allowed prefix is replaced by the bf16 model’s greedy continuation or by filler, with 95% bootstrap intervals over items for the answers changed; Qwen2.5-14B-Instruct, the first 240 test items of each task.
Figure 3: A Classical int8 on one group of 128. B One call of two-level Strassen on the same codes; (i)–(iii) mark where the conditions of Proposition 1 bind.
Figure 4: Excess NLL per token over bf16 on eight 2048-token chunks of WikiText-2, at code bound 31 ( A ) and at 127 with the overflow correction ( B ); thick bars raw, thin bars with temperatures fitted on the other seven chunks, bands ending at the deployed FP8 kernel’s value. Certified values are measured on the classical int8 operators the realizations compute (Table 7 ).
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
sign variant
Φ
nonzero coefficients
largest ∣ coefficient ∣
perplexity ratio
smallest ratio
1,024
432
1
1.367
median ratio
1,024
432
1
1.425
Strassen’s original signs
1,024
432
1
63.11
largest ratio
1,024
432
1
1056.0
Appendix
Table 2: The stability criteria take one value on every sign variant of two-level Strassen; the measured perplexity ratio does not. Four of the 512 sign variants on Qwen2.5-14B-Instruct under the FP8 schedule of Section 3.1 ; the ratio is the WikiText-2 perplexity over that of the bf16 model, and the median row is the upper of the two middle ratios. Φ is the measure of expected error that Xie et al. [35] compute from the coefficients; the nonzero coefficients are counted over U , V and W .
model
variants
smallest
median
largest
spread
>2
FP8 kernel
Qwen2.5-3B
512
1.124
1.191
68.10
60.59
41
1.0068
Qwen2.5-7B-Instruct
508
1.099
1.128
1.18
1.07
0
1.0054
Qwen2.5-14B-Instruct
512
1.367
1.425
1056.0
772.4
137
1.0138
Qwen3-8B
509
1.040
1.105
6.15
5.91
14
1.0014
Llama-3.1-8B-Instruct
512
1.262
1.377
2.89
2.29
5
1.0060
OLMo-2-7B-Instruct
512
1.559
1.629
1.73
1.11
0
1.0023
Appendix
Table 3: The spread of the sign variants of two-level Strassen differs by model, while the deployed FP8 kernel stays close to bf16 on all seven. Perplexity ratios under the FP8 schedule of Section 3.1 , each the perplexity of eight 2048-token chunks of WikiText-2 over the bf16 model’s on the same chunks. Runs whose producer checksum changed during execution are omitted on Qwen2.5-7B-Instruct and Qwen3-8B. “Variants” counts the sign variants with a valid measurement; the median is the middle value, the upper of the two middle values when the count is even; the spread is the largest ratio over the smallest; “ >2 ” counts the variants whose ratio exceeds 2; the last column is the deployed FP8 kernel’s ratio.
condition
realization
changed (%)
mean
median
90th pct.
OpenBookQA, continuation
repaired #1
80.14
0.186
0.097
0.491
repaired #2
89.00
0.297
0.179
0.727
OpenBookQA, filler
repaired #1
90.15
0.190
0.101
0.491
repaired #2
99.88
0.310
0.187
0.762
ARC-Easy, filler
repaired #1
99.72
0.459
0.204
1.245
repaired #2
99.72
0.695
0.320
1.884
Appendix
Table 4: Token-level changes on Qwen2.5-14B-Instruct under the replacements of Table 1 . At each scored candidate position, Δ is the original target token’s log probability on the original input minus its log probability after replacement; all 3308 OpenBookQA or 4315 ARC-Easy positions are included per realization and replacement. The last three columns summarize ∣Δ∣ in nats; the 90th percentile is the empirical inverse cumulative distribution function. Both row-local controls have 0 changed target scores and identical full logit rows under all three comparisons.
Wrong before and after the change (9)
correct answer
37
What is used for sensing visual things?
tibia → nerves
cornea
41
What is different about birth in humans and chickens?
Mother → Fertilization
the hard shell
50
Some berries may be eaten by
a bear or wolf → a bear or lion
a bear or person
Appendix
Table 5: The 14 OpenBookQA items on which repaired #1 changes its chosen answer when the text after the allowed prefix is replaced by the bf16 model’s greedy continuation, on Qwen2.5-14B-Instruct, grouped by how the change affects correctness. Each entry gives the item’s number in the test split (counted from 0) and its question, then the answer chosen on the original input → the answer chosen under replacement; ✓ marks a chosen answer that is correct, and the correct answer is given on the right when neither is.
window test (%)
packed test
model
unrepaired
repaired #1
repaired #2
pairs
unrepaired (%)
repaired #2 (%)
Qwen2.5-14B-Instruct
22.1
8.2
12.7
208
18.3
13.0
Qwen2.5-3B
21.3
9.9
14.6
320
20.9
11.9
Llama-3.1-8B-Instruct
24.8
9.4
15.2
–
–
–
OLMo-2-7B-Instruct
40.9
10.0
15.0
240
29.6
12.9
Qwen3-8B
18.2
7.8
12.7
–
–
–
Appendix
Table 6: The dependence on replaced text, per model. Window test: the percentage of next-token predictions in the kept first quarter of a 1,024-token window that change when the rest of the window is replaced. Packed test: two requests in one forward pass, the number of pairs and the percentage of pairs in which the first request’s most likely next token changes when the second request is replaced. The packed test ran on three models, under the unrepaired realization and repaired #2; both row-local controls change no prediction in either test.
bf16 NLL
deployed FP8 kernel
certified, b=31
certified, b=127
(nats)
raw
adjusted
raw
adjusted
raw
adjusted
Qwen2.5-14B-Instruct
1.311
13.7
13.6
60.7
59.8
3.2
3.3
Llama-3.1-8B-Instruct
1.871
5.9
5.9
14.5
14.5
1.8
1.8
OLMo-2-7B-Instruct
2.078
2.3
1.9
11.0
13.3
−0.4
−0.1
Qwen3-8B
2.129
1.4
2.5
−5.5
18.9
0.0
−0.3
Appendix
Table 7: Increase in negative log-likelihood per token over the bf16 model on eight 2048-token chunks of WikiText-2 ( 10−3 nats; lower is better), raw and with every operator and the bf16 model at its own best temperature on the tested grid, chosen for each chunk on the other seven (adjusted). The bf16 NLL is raw. Each certified column is measured on the classical int8 operator its realization computes bit for bit, the one at b=127 with the overflow correction; all three operators use groups of 128 along the inner dimension.
classical int8 reference
model
bf16
deployed FP8
b=31
b=127
Qwen2.5-14B-Instruct
34.58
33.75
35.00
34.17
Llama-3.1-8B-Instruct
35.42
35.00
35.00
35.00
OLMo-2-7B-Instruct
37.92
39.17
37.92
38.75
Qwen3-8B
31.67
33.75
32.08
31.67
Appendix
Table 8: OpenBookQA accuracy (%) on the first 240 test items. The int8 columns are measured on the classical reference operators that the certified realizations compute, with groups of 128 and the quantization specification above. Each answer is chosen by its summed token log probabilities, with the first candidate selected in an exact tie.
results
where
arithmetic
hardware
Tests of prefix invariance on OpenBookQA and ARC-Easy, windows and packed requests
3.1 , 3.2 ; C.1 , C.2 , C.4 , C.5
fast realizations: block sums rounded to FP8 and multiplied by DeepGEMM’s FP8 kernel, products combined in fp32; controls: the bf16 model and DeepGEMM’s FP8 kernel on the whole product
GPU
How a later row reaches an earlier output
3.2 ; C.3
FP8 rounding simulated by casting to FP8 (e4m3) and back; products in fp32
CPU
Sign variants on seven models, four variants rescored with the later text replaced by filler, summation orders and repeated runs
3.3 ; B.2 , B.3
block sums rounded to FP8 and multiplied by DeepGEMM’s FP8 kernel, products combined in fp32
GPU
The outlier ratio of the counterexample
B.2
the bf16 model’s activations, row norms in fp32
GPU
Stability criteria, code bounds and headroom, and the code bounds of the AlphaTensor algorithms
3.3 , 4.2 , 4.3 ; B.1
exact arithmetic on the coefficients
CPU
Synthetic ensembles and the rotation
B.1 , B.3
FP8 rounding simulated by casting to FP8 (e4m3) and back, products in fp32, against a float64 reference
CPU
Appendix
Table 9: The arithmetic and hardware behind each family of results, with the sections and appendices where the results appear. Outside the layers a row names, the models run in bf16, and likelihoods are computed from their logits in fp32.