stat.MLOct 8, 2026

Softmax Attention on Gaussian Mixtures: Linear When It Can, Selective When It Must

Authors: Simon Gabet, Etienne Boursier, Claire Boyer

Organizations: LMO · Université Paris-Saclay, CNRS, Inria, Laboratoire de mathématiques d’Orsay, 91405 Orsay, France · LMO, CELESTE · LMO, IUF · Institut Universitaire de France

Abstract

Softmax attention, at the heart of Transformers, has demonstrated remarkable capabilities. Yet its underlying mechanisms remain only partially understood. Recent theoretical work studies Gaussian prompts, where the infinite-prompt limit reduces softmax attention to a linear map, but also removes the query-dependent selection that distinguishes it from linear attention. This work studies the infinite-prompt limit of softmax attention on Gaussian mixtures, which retain the tractability of Gaussian data while introducing latent structure, multimodality, and nonlinear dependencies. We show that softmax attention can represent and learn, via gradient-based methods, optimal solutions to a range of statistical tasks, including supervised classification and denoising. Our results highlight two complementary capabilities of softmax attention: it can recover linear tasks as effectively as its simpler linear counterpart, while also exploiting query-dependent context selection to solve nonlinear tasks beyond the reach of linear attention.

Explore similar work

CardsList
  1. Transformer Approximations from ReLUs

    Apr 27, 2026Jerry Yao-Chieh Hu, Mingcheng Lu, Yi-Chen Lee +1Softmax AttentionNeural Network Approximation Theory

  2. What can linear attention learn from nonlinear teachers in-context?

    Oct 7, 2026Mary Letey, Arman Rysmakhanov, Yue M. Lu +2In-Context LearningLinear Attention