stat.MLOct 8, 2026

Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must

Authors: Simon Gabet, Etienne Boursier, Claire Boyer

Organizations: LMO · Université Paris-Saclay, CNRS, Inria, Laboratoire de mathématiques d’Orsay, 91405 Orsay, France · LMO, CELESTE · LMO, IUF · Institut Universitaire de France

Abstract

Softmax attention, at the heart of Transformers, has demonstrated remarkable capabilities. Yet its underlying mechanisms remain only partially understood. Recent theoretical work studies Gaussian prompts, where the infinite-prompt limit reduces softmax attention to a linear map, but also removes the query-dependent selection that distinguishes it from linear attention. This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies. We show that softmax attention can represent and learn, via gradient-based methods, optimal solutions to a range of statistical tasks, including supervised classification and denoising. Our results highlight two complementary capabilities of softmax attention: it can recover linear tasks as effectively as its simpler linear counterpart, while also exploiting query-dependent context selection to solve nonlinear tasks beyond the reach of linear attention.

Explore similar work

Jun 9, 2026cs.LG

Gaussian Mixture Attention: Linear-Time Sequence Mixing via Probabilistic Latent Routing

The dense token-to-token interaction pattern of standard dot-product attention remains a central bottleneck in scaling Transformer architectures to long contexts. We introduce \textbf{Gaussian Mixture Attention (GMA)}, a probabilistic attention-style sequence mixer that replaces explicit pairwise query--key comparison with routing through KK learned Gaussian mixture components. Queries and keys are mapped to posterior \textit{responsibility} vectors over a shared latent routing space; their overlap defines an implicit responsibility-space affinity, while values are written into and read from a KK-slot latent memory. By exploiting the associativity of matrix multiplication, GMA avoids materializing the induced N×NN\times N affinity matrix and instead uses two responsibility matrices whose dominant activation storage scales as O(NK)\mathcal{O}(NK) rather than O(N2)\mathcal{O}(N^2) for fixed KK. We formulate bidirectional and causal variants of GMA, provide an end-to-end differentiable parameterization of the Gaussian mixture components, and analyze its responsibility-modulated gradient structure, constrained non-negative low-rank affinity interpretation, and local routing stability. Empirically, GMA exhibits the intended fixed-KK linear memory scaling and is competitive with attention-style baselines on long-context classification, while causal GMA improves over tested linear/random-feature attention variants on WikiText-103 but remains behind optimized causal SDPA and Mamba in the current implementation. Analysis of learned responsibilities further shows broad component usage and moderate alignment with surface-form token categories, supporting GMA as a probabilistic, interpretable, fixed-KK linear-time attention-style alternative rather than a universal replacement for optimized softmax attention or state-space models.
Apr 27, 2026cs.LG

Transformer Approximations from ReLUs

We provide a systematic recipe for translating ReLU approximation results to softmax attention mechanism. This recipe covers many common approximation targets. Importantly, it yields target-specific, economic resource bounds beyond universal approximation statements. We showcase the recipe on multiplication, reciprocal computation, and min/max primitives. These results provide new analytical tools for analyzing softmax transformer models.
Oct 7, 2026stat.ML

What can linear attention learn from nonlinear teachers in-context?

Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, y=f(x⊤w)+εy=f(x^\top w)+\varepsilon . Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of ff, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.