Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention
Abstract
How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row- approximation rank , exactly the least rank achieving uniform error over all bounded vector-valued values. Row softmax exposes the intrinsic interaction , whereas invertible gauges leave fixed while changing the Euclidean geometry of a chosen query/key factorization. We replace that coordinate-dependent description by a projective residual and an attained factor-radius size . For every rank- retained interaction with , we prove with the same unknown dimension constant as the underlying weighted Gibbs-row cover. The profile is gauge invariant, termwise no worse than native retained-subspace bounds at the same declared dimension, and has a worst-case sharp size exponent at fixed and . We then measure directly on learned attention using 9,978 certified brackets across BERT, GPT-2, Qwen2.5, and two ViT checkpoints; where certificates do not close, the optimum remains interval-valued. A pre-specified 2,302-cell held-out study further shows that the historical native-coordinate geometry block contains coarse, mostly head-level information but no detectable incremental information beyond a strong calibrated baseline. The new intrinsic descriptor is not evaluated in that study. Together, the theory and measurements distinguish an operator-intrinsic complexity control from a stronger empirical explanation that the learned-head evidence does not support.