cs.DCSep 27, 2026

Validating Memory-Optimal Transformer Kernels on Real Hardware: From Formal Derivation to Measured Performance Across Two HPC Clusters

Authors: Lenore M. Mullin, Gaetan Hains

Organizations: Professor Emerita, College of Nanotechnology, Science, and Engineering, University at Albany (SUNY), Albany, NY, USA.

Abstract

We validate memory-optimal cost functions for transformer kernels derived via the Mathematics of Arrays (MoA). Companion Papers I-IV formally derive kernels for attention forward, backward, fused forward+backward, decode, and the complete block (RMSNorm, gated MLP) as a hardware-independent specification (DNF) transformed to a machine-specific realization (ONF) via gamma, with verification to machine precision against PyTorch. This paper checks those predictions against measured performance on two HPC clusters (Purdue Anvil, NCSA Delta) across CPU and GPU. Three results stand out. (1) We identify and fix a GPU regression: fusing forward+backward, proven to avoid materializing an O(n^2) intermediate, initially ran slower than naive on GPU due to atomic contention. Profiling confirmed 2.00x more atomic instructions; a targeted ONF rewrite reversed it, yielding up to 2.5x speedup. (2) Identical derivations produce markedly different real costs by topology: 535x NUMA-locality penalty on one cluster vs <3x oversubscription on another, showing optimal deployment is a function of the machine's array structure. (3) We report a partially resolved anomaly: identical denotational computations run faster in C than Fortran on CPU but faster in Fortran than C on GPU, narrowed to one dominant kernel and one memory-latency stall mechanism (3.17x time gap matches 3.35x stall gap). We treat hardware-specific optimization as a routine ONF rewrite with fixed, verified DNF, a candidate methodology for scaling AI onto evolving hardware without re-deriving correctness.

Figures & tables

Explore similar work

CardsList
  1. Attention at the Theoretical Minimum: A Mathematics of Arrays Framework for Memory-Optimal Transformer Kernels

    Jun 5, 2026Lenore Mullin, Gaetan HainsMatrix MultiplicationTransformer Architectures

  2. MoA-Structured Decode Attention DNF Derivation, KV-Cache Accumulation, GQA/MQA, and OpenACC Kernel

    Jul 21, 2026Lenore Mulin, Gaetan HainsMatrix MultiplicationTransformer Attention

  3. CODA: Rewriting Transformer Blocks as GEMM-Epilogue Programs

    May 19, 2026Han Guo, Jack Zhang, Arjun Menon +4Graphics Processing Unit KernelsMatrix Multiplication