stat.MLOct 7, 2026

Gaussian Equivalence for Multi-Head Self-Attention

Authors: Tomohiro Hayase, Ryo Karakida

Organizations: Artificial Intelligence Research Center (AIRC), AIST · RIKEN AIP

Abstract

A theoretical understanding of multi-head self-attention is fundamental to the study of modern neural networks. Using random matrix theory, we establish Gaussian equivalence for multi-head self-attention: replacing softmax attention with rescaled scores plus Gaussian noise preserves the limiting spectral law of the centered output. This equivalence also covers value and output projections that depend on the keys. The resulting laws separate the effects of head allocation and projection widths, and distinguish spectrum-preserving across-head sharing from within-head key--value dependence.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity

    May 18, 2026Ernest FokouéMulti-Head AttentionIntrinsic Dimensionality

  2. Provably Learning Multi-Head Attention with Queries

    Aug 4, 2026Sunyeop Kim, Insung Kim, Jian GuoMulti-Head AttentionRectified Linear Unit