cs.LGAug 28, 2026

Intrinsic Interaction Geometry Controls the Low-Rank Complexity of Softmax Attention

Authors: Yuhe Sui, Jianing Zhang, Yingzhi Tang

Abstract

How much matrix rank is required to preserve every bounded value output of normalized softmax attention? We study the unrestricted maximum-row-ℓ1\ell_1 approximation rank rε(A)r_\varepsilon(A), exactly the least rank achieving uniform error over all bounded vector-valued values. Row softmax exposes the intrinsic interaction C=Pm(log⁡A)PNC=P_m(\log A)P_N, whereas invertible Q/KQ/K gauges leave AA fixed while changing the Euclidean geometry of a chosen query/key factorization. We replace that coordinate-dependent description by a projective residual q(C−T)q(C-T) and an attained factor-radius size κ(T)κ(T). For every rank-rr retained interaction with τ(T)<ετ(T)<\varepsilon, we prove rε(A)≤min⁡{N,  Cr(1+κ(T)(ε−τ(T))2)r/2},r_\varepsilon(A)\le \min\left\{ N,\; C_r\left( 1+\frac{κ(T)} {(\varepsilon-τ(T))^2} \right)^{r/2} \right\}, with the same unknown dimension constant as the underlying weighted Gibbs-row cover. The profile is gauge invariant, termwise no worse than native retained-subspace bounds at the same declared dimension, and has a worst-case sharp r/2r/2 size exponent at fixed rr and ε\varepsilon. We then measure rε(A)r_\varepsilon(A) directly on learned attention using 9,978 certified brackets across BERT, GPT-2, Qwen2.5, and two ViT checkpoints; where certificates do not close, the optimum remains interval-valued. A pre-specified 2,302-cell held-out study further shows that the historical native-coordinate geometry block contains coarse, mostly head-level information but no detectable incremental information beyond a strong calibrated baseline. The new intrinsic descriptor is not evaluated in that study. Together, the theory and measurements distinguish an operator-intrinsic complexity control from a stronger empirical explanation that the learned-head evidence does not support.

Explore similar work

CardsList