stat.MLOct 7, 2026
SaveGaussian Equivalence for Multi-Head Self-Attention
Organizations: Artificial Intelligence Research Center (AIRC), AIST · RIKEN AIP
Abstract
A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.
Figures & tables
Figure 1: Gaussian equivalence through concatenation and projections. The blue dashed and orange solid families are independent of each other.
| Model | Transformation | Control used |
|---|---|---|
| Centered multi-head self-attention | Initial model | |
| Replace the row denominators by a constant | Operator-norm error ; rank at most | |
| Subtract the conditional row mean | Rank at most one; norms | |
| Gaussian rows with matching covariance | Quadratic-form concentration; self-averaging | |
| Approximate by | Normalized Frobenius error | |
| Use the original queries and independent noise | Same conditional law; self-averaging |
Table 1: Equivalent random matrix models, defined in Section 4 . Each step preserves all fixed spectral moments after right multiplication under Assumption 4.1 . The third column states the control used.
| Dimensions of | Sharing | Rel. (%) | ||
|---|---|---|---|---|
| BERT-base | None | 512 | 12 | |
| ModernBERT-L | None | 1,024 | 16 | |
| DINOv3-L/16 | None | 261 | 16 | |
| LLaDA-8B | None | 4,096 | 32 | |
| GPT-3 175B | None | 2,048 | 96 | |
| Falcon-7B | MQA | 2,048 | 71 |
Table 2: Finite-size MHA–GE spectral agreement at dimensions taken from publicly available models. Means with 95% CIs. Full dimensions and sources: Section D.1 .
Figure 2: Spectral predictions for readout learning and width allocation. (a) Training loss for three head counts. (b) Coefficient-error map with fixed-output (blue) and fixed-budget (orange) paths; dotted lines mark width boundaries. (c) Errors along these paths. (d) Coefficient-error ratio, , at fixed ; above the dashed unit line, equal widths give lower error; below it, wider V gives lower error. Curves: GE; markers: MHA; bars: 95% CIs. Maximum error bars: (a) ; (c) ; (d) . Details: Sections D.2 , D.3 and D.4 .
Figure 3: Parameter sharing. (a) Across-head sharing; diamond: unshared. (b) GQA versus within-head K=V at equal widths and parameter count, . Lines: GE; markers: MHA; bars: 95% CIs. Maximum error bars: (a) ; (b) . Settings in Sections D.5 and D.6 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Name | Definition |
|---|---|---|
| Sets and linear algebra | ||
| Natural numbers | ; zero is excluded. | |
| , | Real and complex numbers | The real and complex number fields. |
| Upper half-plane | . | |
| Identity matrix | identity. | |
| All-ones vector | . | |
Table A.1: General mathematical notation.
| Symbol | Name | Definition or range |
|---|---|---|
| Context length | . | |
| Input width | , . | |
| Query/key width per head | . | |
| Value width per head | . | |
| Output width | . | |
| Number of heads | . |
Table A.2: Dimensions and parameters.
| Symbol | Size | Name | Definition |
|---|---|---|---|
| Input and projection matrices | |||
| Input matrix | ; Section 2.2 . | ||
| Query weights | Section 2.2 . | ||
| Key weights | Section 2.2 . | ||
| Value weights | Section 2.2 . | ||
| Output weights | Section 2.2 . | ||
Table A.3: Matrices and their dimensions.
| Symbol | Name | Definition |
|---|---|---|
| Exponential feature function | . | |
| Row softmax | . | |
| Cauchy transform | ( C.1 ). | |
| -function | ( C.2 ). | |
| Inverse -function | . | |
| -transform | ( C.3 ). |
Table A.4: Model functions and spectral transforms.
| Symbol | Name | Definition |
|---|---|---|
| Gaussian law | Mean and covariance ; in one dimension the second parameter is the variance. | |
| Dirac measure | Unit point mass at . | |
| Empirical eigenvalue law | . | |
| Squared-singular-value law | . | |
| Value covariance law | . | |
| Output covariance law | . |
Table A.5: Probability laws.
| Figure | Quantity or path | Absolute | Relative (%) |
|---|---|---|---|
| 2 (a) | Relative training loss | ||
| 2 (c) | Coefficient error | ||
| 2 (d) | Coefficient-error ratio | ||
| 3 (a) | Arrival time, fixed widths | ||
| 3 (a) | Arrival time, fixed parameter count | ||
| 3 (b) | Coefficient error |
Table D.1: Maximum error bars in the learning figures. Loss, coefficient-error, and error-ratio distances are dimensionless; arrival-time distances use the gradient-flow time units.
| Dimensions of | Sharing | Rel. (%) | KS | ||||
|---|---|---|---|---|---|---|---|
| BERT-base | None | 512 | 768 | 12/12 | 64 | ||
| ModernBERT-large | None | 1,024 | 1,024 | 16/16 | 64 | ||
| DINOv3-L/16 | None | 261 | 1,024 | 16/16 | 64 | ||
| LLaDA-8B | None | 4,096 | 4,096 | 32/32 | 128 | ||
| GPT-3 175B | None | 2,048 | 12,288 | 96/96 | 128 | ||
| Falcon-7B | MQA | 2,048 | 4,544 | 71/1 | 64 |
Table D.2: Finite-size GE at dimensions taken from publicly available models, with . Centered MHA–GE spectral distances at , with . ∗ Kayyam uses the 2048-token evaluation setting.
| Experiment | |||||||
|---|---|---|---|---|---|---|---|
| 2 (a), learning curves | 4096 | 4096 | 32–128 | 64 | 64 | 2048–8192 | 8192 |
| 2 (c), fixed output | 4096 | 4096 | 16–120 | 64 | 64 | 1024–7680 | 2048 |
| 2 (c), fixed budget | 4096 | 4096 | 32–120 | 64 | 64 | 2048–7680 | 819–36864 |
| 2 (d), temperature | 8192 | 12288 | 96 | 64–128 | 128–192 | 12288–18432 | 12288 |
| 3 (a), across-head sharing | 4096 | 4096 | 64 | 64–106 | 64–106 | 4096–6784 | 8192 |
| 3 (b), sharing schemes | 4096 | 8192 | 2–32 | 64–1024 | 64–1024 | 2048 | 2048 |
Table D.3: Dimensions used in the learning experiments. Ranges cover the sampled softmax configurations.
Figure D.1: Head splitting concentrates the attention bulk (a), output width changes the MHA spectral scale and shape (b), and within-head key–value tying changes the limit while across-head sharing preserves it (c). Histograms show original-model spectra; black curves show the theoretical limits. Experimental details are in Section D.7 .
Explore similar work
We develop a rigorous statistical theory of multi-head attention (MHA) as an ensemble of Nadaraya-Watson (NW) kernel regression estimators. Building on the algebraic identity between single-head softmax attention and the NW estimator, we prove that MHA is a structured ensemble of H NW estimators, each operating in a distinct learned projection subspace of the key space. We derive an explicit Bias-Variance-Covariance decomposition of the MHA mean squared error, showing that variance reduction depends not merely on the number of heads H but fundamentally on the decorrelation of head outputs. Decorrelation is governed by the principal angles between learned projection subspaces: orthogonal projections yield maximum variance reduction; aligned projections yield none. We introduce the Head Diversity Index (HDI), a computable spectral measure of inter-head decorrelation, and prove that MHA mean squared error is monotonically decreasing in HDI. This provides the first rigorous theoretical explanation for the empirically observed specialization of attention heads. Under a fixed total-dimension budget D = H * d_k, we solve the optimal head-dimension allocation problem, deriving the MSE-minimizing pair (H*, d_k*) from data distribution and regression smoothness. The solution yields a new architectural scaling law: the optimal per-head dimension grows logarithmically with training set size, while the optimal number of heads grows nearly linearly with the total budget D. Our framework unifies three strands of prior work: the NW theory of single-head attention, the general weighting theory for ensemble learning, and the decorrelation-variance-reduction isomorphism between biological and computational ensembles. Multi-head attention is the Transformer's instantiation of a universal principle: identical agents plus diversity-enforcing mechanisms yields emergent optimality.
Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention
Transformer self-attention can be interpreted as a gradient flow on the unit sphere, in which tokens evolve under softmax interaction potentials and tend to form clusters. While prior work has established clustering behavior for single-head attention, the multi-head setting remains less understood due to geometric interference between heads, which invalidates standard monotonicity arguments. In this work, we develop a theoretical framework for multi-head self-attention dynamics and resolve several open questions. We show that, under suitable conditions on the score matrices, a natural multi-head energy functional is non-decreasing along both flat and spherical dynamics. We identify the key obstruction to per-head monotonicity as radial shadow terms, which are projections of each head's output onto token directions, persisting even under orthogonality assumptions. We introduce a sufficient condition ensuring monotonicity and establish robustness to approximate orthogonality. In a simplified scalar-head regime with equiangular token configurations, we derive a closed-form expression for the critical inverse temperature governing clustering behavior, and show that heterogeneous heads exhibit super-additive clustering rates. In this regime, we also prove a separation in clustering time between ReLU and softmax attention in the linearized dynamics. Finally, we establish an entropy production identity and show that attention entropy increases monotonically toward equilibrium as clustering progresses. Our results provide a unified perspective on the dynamics of multi-head attention and clarify the mechanisms underlying clustering and stability in transformer models.
Provably Learning Multi-Head Attention with Queries
We study the problem of learning multi-head softmax attention from black-box input-output access. The learner may query arbitrary real-valued token sequences and observe only the scalar output at the final token. Recent work gives an algorithm using value queries to recover the single-head parameters . For multiple heads, the same work establishes identifiability under the assumption that the heads occupy pairwise orthogonal subspaces. Applying the single-head recovery algorithm separately to the heads additionally requires bases for these subspaces to be known. We recover a canonical representation by merging heads with the same , summing their corresponding , and discarding a merged head when this sum is zero, without these subspace assumptions. By varying the number of copies of a token, our algorithm obtains samples of a rational function whose interpolation separates the canonical heads. Additional queries formed by adding selected token vectors then match the same head across different queries. When the oracle outputs and all subsequent computations are exact, the learner chooses its query vectors at random and recovers the canonical pairs up to permutation with probability one. When is known, it uses exactly value queries of maximum length . If only a known upper bound is available, the algorithm uses value queries of maximum length . For approximate oracle outputs, we give conditions under which the parameter error is at most a model- and query-dependent constant multiple of the output error. Finally, we extend our result to a one-layer Transformer with multi-head attention followed by a bias-free ReLU feed-forward network. Under additional conditions, we recover a functionally equivalent Transformer without relying on a separate algorithm for learning the feed-forward network.