cs.ARNov 14, 2025

Bit-Accurate Modeling of GPU Matrix Multiply-Accumulate Units: Demystifying Numerical Discrepancy and Accuracy

Authors: Peichen Xie, Shuotao Xu, Yang Wang, Fan Yang, Mao Yang

Abstract

Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose only instruction-level or API-level interfaces of matrix multiply-accumulate (MMA) operations, while leaving internal floating-point arithmetic behavior undocumented. Consequently, MMAUs across vendors and architectural generations often produce numerical discrepancies for identical inputs, and sometimes exhibit reduced numerical accuracy that can cause training instability. Diagnosing and understanding the root causes of these effects is challenging without white-box models of their arithmetic behavior. This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations. Based on this framework, we analyze all MMA instructions on ten GPU architectures spanning NVIDIA Volta through RTX Blackwell and AMD CDNA1 through CDNA3, and derive the first bit-accurate arithmetic models for these MMAUs. Our models explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, reveal four types of precision bottlenecks and one type of numerical asymmetry, and inform software workarounds as well as design suggestions for future MMAUs. This work is open-source at https://github.com/microsoft/MMA-Sim

Figures & tables

Explore similar work

CardsList
  1. Accelerating the Mitigation of LLM Inference Nondeterminism Across GPU Architectures

    Sep 22, 2026Liam Cooper, Shinnung Jeong, Hyeran Jeon +2LLM Inference OptimizationFloating-Point

  2. Measuring and Reducing Cross-Vendor Mismatch in Language Models

    Oct 4, 2026Erland Hilman Fuadi, Chong Tian, Xiaosong Ma +1Cross-Architecture GeneralizationTraining-Inference Mismatch