cs.LGMay 15, 2026

Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows

Authors: Alex MassuccoLeonardo Del GrandeMarcello CarioniChristoph BruneCarola-Bibiane Schönlieb

Organizations: Department of Applied Mathematics and Theoretical Physics, University of Cambridge, Cambridge, UK · Department of Mathematics, University of Twente, Enschede, Netherlands

Abstract

In recent years, transformer architectures have revolutionized the field of language processing, opening the door to previously unforeseen possibilities. However, from a theoretical point of view, the mathematical models proposed in the literature often lack direct contact with the actual architectures and depend on strong simplifying assumptions. In this paper, we reduce this gap by modelling the data flow in multi-headed transformer architectures as time-dependent gradient flows for a suitable interaction energy capturing the design of the attention mechanism. The explicit dependence on time allows us to consider different weights for each head and for each layer, without imposing constraints on the initialization method. Moreover, we prove that, under a suitable integrability assumption on the evolution of the weights, each element of the ωω-limit set of the gradient flows is a stationary point of the interaction energy at a limiting weight distribution. Finally, we analyse the stability of the gradient flows considering perturbations of both the initial data and the weights. Specifically, on the one hand, we study the robustness of the proposed models with respect to noisy inputs, establishing a continuous dependence of the gradient flows on the initial data and uniqueness of the flows. On the other hand, we prove the ΓΓ-convergence of the perturbed interaction energy to the unperturbed one, leading to the convergence of the corresponding gradient flows. We complement these theoretical results with numerical experiments that confirm the predicted energy-dissipation identity and clarify the asymptotic behavior of the dynamics in both the autonomous-like (Ornstein--Uhlenbeck) and the genuinely non-autonomous (oscillating-weights) regimes.

Explore similar work

CardsList
  1. Training Infinitely Deep and Wide Transformers

    May 17, 2026Raphaël Barboni, Maarten V. de Hoop, Takashi Furuya +1Transformer ArchitecturesMean-Field Limit