cs.LGSep 30, 2026

Beyond Accuracy: Prefix-Invariant Realizations of Low-Precision Fast Matrix Multiplication

Authors: Shuxiao Xie, Shuyang Xie, Yuan Cao, Dezhi Ran, Wei Yang, Tao Xie

Abstract

Fast matrix multiplication saves multiplications through exact cancellation, but rounding sums that mix token rows can leave contributions from later tokens in earlier language model outputs. This threatens prefix invariance, which multiple-choice likelihood scoring relies on: a scored likelihood must depend only on its allowed prefix. On Qwen2.5-14B-Instruct, two fast FP8 realizations repaired to ordinary-looking accuracy still change the answers chosen by likelihood on 5.83% and 10.00% of 240 OpenBookQA items when only the text after the allowed prefix is replaced with the bf16 model's own greedy continuation. Both row-local controls, the bf16 model and a deployed FP8 matrix multiplication kernel, change none. Accuracy thus does not certify prefix invariance, and the stability criteria we analyze cannot tell realizations apart: across all 512 sign variants of two-level Strassen they stay constant while teacher-forced perplexities span a 772.4×\times range on the same model. We therefore construct certified realizations of two-level Strassen on bounded integer codes that quantize token rows independently, then mix and cancel exactly before rescaling, using 49 block multiplications instead of 64. Our certificate guarantees bitwise equality to a prescribed row-local classical int8 operator at the same quantization specification, so every certified realization inherits its prefix invariance. Certification thus turns realization choice into a pure cost decision: which certified realization runs can no longer change a single scored likelihood.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. When Good Enough Is Optimal: Multiplication-Only Matrix Inversion Approximation for Quantized Gated DeltaNet

    Jun 4, 2026Luoming Zhang, Yuwei Ren, Kui Zhang +7Matrix MultiplicationKimi Delta Attention

  2. Contraction-Gauge Preconditioning for Quantized Matrix Multiplication

    Jul 21, 2026Piyush Sao, Narasinga Miniskar, Pedro Valero-Lara +2Matrix MultiplicationSpectral Preconditioning

  3. High-Rate Quantized Matrix Multiplication II

    May 13, 2026Or Ordentlich, Yury PolyanskiyMatrix MultiplicationRandom Rotations