stat.MLOct 7, 2026
SaveGaussian Equivalence for Multi-Head Self-Attention
Organizations: Artificial Intelligence Research Center (AIRC), AIST · RIKEN AIP
Abstract
A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.
Figures & tables
Figure 1: Gaussian equivalence through concatenation and projections. The blue dashed and orange solid families are independent of each other.
| Model | Transformation | Control used |
|---|---|---|
| Centered multi-head self-attention | Initial model | |
| Replace the row denominators by a constant | Operator-norm error ; rank at most | |
| Subtract the conditional row mean | Rank at most one; norms | |
| Gaussian rows with matching covariance | Quadratic-form concentration; self-averaging | |
| Approximate by | Normalized Frobenius error | |
| Use the original queries and independent noise | Same conditional law; self-averaging |
Table 1: Equivalent random matrix models, defined in Section 4 . Each step preserves all fixed spectral moments after right multiplication under Assumption 4.1 . The third column states the control used.
| Dimensions of | Sharing | Rel. (%) | ||
|---|---|---|---|---|
| BERT-base | None | 512 | 12 | |
| ModernBERT-L | None | 1,024 | 16 | |
| DINOv3-L/16 | None | 261 | 16 | |
| LLaDA-8B | None | 4,096 | 32 | |
| GPT-3 175B | None | 2,048 | 96 | |
| Falcon-7B | MQA | 2,048 | 71 |
Table 2: Finite-size MHA–GE spectral agreement at dimensions taken from publicly available models. Means with 95% CIs. Full dimensions and sources: Section D.1 .
Figure 2: Spectral predictions for readout learning and width allocation. (a) Training loss for three head counts. (b) Coefficient-error map with fixed-output (blue) and fixed-budget (orange) paths; dotted lines mark width boundaries. (c) Errors along these paths. (d) Coefficient-error ratio, , at fixed ; above the dashed unit line, equal widths give lower error; below it, wider V gives lower error. Curves: GE; markers: MHA; bars: 95% CIs. Maximum error bars: (a) ; (c) ; (d) . Details: Sections D.2 , D.3 and D.4 .
Figure 3: Parameter sharing. (a) Across-head sharing; diamond: unshared. (b) GQA versus within-head K=V at equal widths and parameter count, . Lines: GE; markers: MHA; bars: 95% CIs. Maximum error bars: (a) ; (b) . Settings in Sections D.5 and D.6 .
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbol | Name | Definition |
|---|---|---|
| Sets and linear algebra | ||
| Natural numbers | ; zero is excluded. | |
| , | Real and complex numbers | The real and complex number fields. |
| Upper half-plane | . | |
| Identity matrix | identity. | |
| All-ones vector | . | |
Table A.1: General mathematical notation.
| Symbol | Name | Definition or range |
|---|---|---|
| Context length | . | |
| Input width | , . | |
| Query/key width per head | . | |
| Value width per head | . | |
| Output width | . | |
| Number of heads | . |
Table A.2: Dimensions and parameters.
| Symbol | Size | Name | Definition |
|---|---|---|---|
| Input and projection matrices | |||
| Input matrix | ; Section 2.2 . | ||
| Query weights | Section 2.2 . | ||
| Key weights | Section 2.2 . | ||
| Value weights | Section 2.2 . | ||
| Output weights | Section 2.2 . | ||
Table A.3: Matrices and their dimensions.
| Symbol | Name | Definition |
|---|---|---|
| Exponential feature function | . | |
| Row softmax | . | |
| Cauchy transform | ( C.1 ). | |
| -function | ( C.2 ). | |
| Inverse -function | . | |
| -transform | ( C.3 ). |
Table A.4: Model functions and spectral transforms.
| Symbol | Name | Definition |
|---|---|---|
| Gaussian law | Mean and covariance ; in one dimension the second parameter is the variance. | |
| Dirac measure | Unit point mass at . | |
| Empirical eigenvalue law | . | |
| Squared-singular-value law | . | |
| Value covariance law | . | |
| Output covariance law | . |
Table A.5: Probability laws.
| Figure | Quantity or path | Absolute | Relative (%) |
|---|---|---|---|
| 2 (a) | Relative training loss | ||
| 2 (c) | Coefficient error | ||
| 2 (d) | Coefficient-error ratio | ||
| 3 (a) | Arrival time, fixed widths | ||
| 3 (a) | Arrival time, fixed parameter count | ||
| 3 (b) | Coefficient error |
Table D.1: Maximum error bars in the learning figures. Loss, coefficient-error, and error-ratio distances are dimensionless; arrival-time distances use the gradient-flow time units.
| Dimensions of | Sharing | Rel. (%) | KS | ||||
|---|---|---|---|---|---|---|---|
| BERT-base | None | 512 | 768 | 12/12 | 64 | ||
| ModernBERT-large | None | 1,024 | 1,024 | 16/16 | 64 | ||
| DINOv3-L/16 | None | 261 | 1,024 | 16/16 | 64 | ||
| LLaDA-8B | None | 4,096 | 4,096 | 32/32 | 128 | ||
| GPT-3 175B | None | 2,048 | 12,288 | 96/96 | 128 | ||
| Falcon-7B | MQA | 2,048 | 4,544 | 71/1 | 64 |
Table D.2: Finite-size GE at dimensions taken from publicly available models, with . Centered MHA–GE spectral distances at , with . ∗ Kayyam uses the 2048-token evaluation setting.
| Experiment | |||||||
|---|---|---|---|---|---|---|---|
| 2 (a), learning curves | 4096 | 4096 | 32–128 | 64 | 64 | 2048–8192 | 8192 |
| 2 (c), fixed output | 4096 | 4096 | 16–120 | 64 | 64 | 1024–7680 | 2048 |
| 2 (c), fixed budget | 4096 | 4096 | 32–120 | 64 | 64 | 2048–7680 | 819–36864 |
| 2 (d), temperature | 8192 | 12288 | 96 | 64–128 | 128–192 | 12288–18432 | 12288 |
| 3 (a), across-head sharing | 4096 | 4096 | 64 | 64–106 | 64–106 | 4096–6784 | 8192 |
| 3 (b), sharing schemes | 4096 | 8192 | 2–32 | 64–1024 | 64–1024 | 2048 | 2048 |
Table D.3: Dimensions used in the learning experiments. Ranges cover the sampled softmax configurations.
Figure D.1: Head splitting concentrates the attention bulk (a), output width changes the MHA spectral scale and shape (b), and within-head key–value tying changes the limit while across-head sharing preserves it (c). Histograms show original-model spectra; black curves show the theoretical limits. Experimental details are in Section D.7 .