Modern AI accelerators rely on matrix multiply-accumulate units (MMAUs), such as NVIDIA Tensor Cores and AMD Matrix Cores, to accelerate deep neural network workloads. MMAUs expose only instruction-level or API-level interfaces of matrix multiply-accumulate (MMA) operations, while leaving internal floating-point arithmetic behavior undocumented. Consequently, MMAUs across vendors and architectural generations often produce numerical discrepancies for identical inputs, and sometimes exhibit reduced numerical accuracy that can cause training instability. Diagnosing and understanding the root causes of these effects is challenging without white-box models of their arithmetic behavior. This paper proposes closed-loop feature probing (CLFP), a generic and systematic framework for constructing bit-accurate arithmetic behavior models of MMA operations. Based on this framework, we analyze all MMA instructions on ten GPU architectures spanning NVIDIA Volta through RTX Blackwell and AMD CDNA1 through CDNA3, and derive the first bit-accurate arithmetic models for these MMAUs. Our models explain previously observed cross-platform numerical discrepancies and accuracy issues, enable white-box numerical error analysis, reveal four types of precision bottlenecks and one type of numerical asymmetry, and inform software workarounds as well as design suggestions for future MMAUs. This work is open-source at https://github.com/microsoft/MMA-Sim
Figures & tables
Fig. 1: The closed-loop feature probing (CLFP) framework for modeling the arithmetic behavior of the MMA operation. The loop of Steps 4 (verification) and 5 (adding new tests and revising the model) ensures completeness and bit-accuracy.
Fig. 2: Examples of summation trees (left) and corresponding values of d(i,j)/v (right), where (i,j) denotes the input with pi=U , pj=−U , and all other summands set to v .
GPU
MMA-Sim
Human
Total
Step 1
5 minutes
0
0
5 minutes
Step 2
< 1 minute
0
0
< 1 minute
Step 3
< 1 minute
0
< 1 hour
< 1 hour
Step 4
5 minutes
1 hour*
0
1 hour
Step 5
< 1 minute
< 1 minute
4 hours
4 hours
Total
10-20 minutes
1-3 hours
1-9 hours
2-12 hours
TABLE I: Approximate time of running CLFP with 0-2 revision loops (typical). *MMA-Sim execution time has been reduced from 10-20 hours to 1 hour through optimization.
TABLE II: Categories of our bit-accurate models for GPU matrix multiply-accumulate units.
Architecture
Instruction
Model
Input Type
Output Type
Lmax
F
ρ
Volta
HMMA
ΦT-FDPA
FP16
FP32
4
23
RZ FP32
FP16
FP16
4
23
RNE FP16
Turing
HMMA
ΦT-FDPA
FP16
FP32
8
24
RZ FP32
FP16
FP16
8
24
RNE FP16
Ampere
DMMA
ΦFMA
FP64
FP64
1
N/A
Configurable
HMMA
ΦT-FDPA
TF32
FP32
4
24
RZ FP32
TABLE III: Models and parameters for all NVIDIA and AMD GPU MMA instructions. In addition, G=16 for all instructions modeled by ΦGST-FDPA .
Architecture
TF32/BF16 Instr.
FP16 Instr.
FP8 Instr.
Volta
N/A
0.0
N/A
Turing
N/A
−0.5
N/A
Ampere
−0.5
−0.5
N/A
Ada Lovelace
−0.5
−0.5
0.0
Hopper
−0.75
−0.75
0.0
Blackwell
−0.75
−0.75
−0.75
TABLE IV: The divergent results of different MMA instructions for the same input in Equation 10 . In addition, all FP64/FP32 instructions produce d0,0=−0.875 .
Elementary Op.
Error Source
Error Bound
FTZ-Add/Mul
Input FTZ
2−126 or 2−14
Add/Mul
0.5 ulp
Output FTZ
2−126
FMA, E-FDPA
Output rounding
0.5 ulp or 1 ulp
T-FDPA, others
Fused summation
(L+1)2emax−F
Output rounding
0.5 ulp or 1 ulp
TABLE V: Sources and upper bounds of numerical error.
Architectures and Instructions
Concerns
AMD CDNA2, FP16-input
Input FTZ
NVIDIA Ada Lovelace and Hopper, FP8-input
Small F
NVIDIA Ada Lovelace and Hopper, FP8-input
ρ=RZE8M13
All NVIDIA architectures, FP16-output
ρ=RNEFP16
AMD CDNA3, TF32/BF16/FP16/FP8-input
Asymmetry
TABLE VI: Concerns regarding numerical precision (top) and accuracy (bottom).
Fig. 3: Distributions of δRD , the numerical deviation of the AMD CDNA3 FP16 MMA instruction, which uses the “round down” (RD) mode internally, and δRZ , the numerical deviation of a hypothetical variant that replaces the internal RD operations with “round toward zero” (RZ) operations.