cs.LGMay 8, 2026

Convergent Stochastic Training of Attention and Understanding LoRA

Authors: Zhengkai SunDibyakanti KumarAlejandro F FrangiAnirbit MukherjeeMingfei Sun

Abstract

Transformers have revolutionized machine learning and deploying attention layers in the model is increasingly standard across a myriad of applications. Further, for large models, it is common to implement Low Rank Adaptation (LoRA), whereby a factorized parameterization of them is trained, to achieve a surprisingly beneficial accuracy-size trade-off. In this work, via a unified framework we rigorously establish trainability of such models under stochastic methods. We prove that for any mild regularization, the empirical regression loss on a attention layer and LoRA on a shallow neural net, both induce Poincaré inequality for the corresponding Gibbs' measure. Then it follows via invoking recent results that a certain SDE, which mimics the SGD, minimizes the corresponding losses. In both the cases, our first-of-its-kind results of trainability on attention and nets, do not rely on any assumptions on the data or the size of the architecture.

Explore similar work

Jun 4, 2026cs.LG

High-Dimensional Theory of LoRA Fine-Tuning in a Solvable Attention Model

We develop a high-dimensional statistical theory of low-rank adaptation (LoRA) in attention models, capturing the interplay between pre-training and fine-tuning. We introduce a solvable framework in which a single-head attention layer is first pre-trained on a data-abundant task and subsequently adapted via a rank-one LoRA update on limited data. In the high-dimensional limit, both stages admit a sharp asymptotic characterization in terms of a finite set of order parameters, yielding explicit predictions for test errors and representation alignment. Our analysis shows that the impact of pre-training on LoRA is summarized by an effective noise term, from which we derive prescriptions for the optimal pre-training procedure. We also demonstrate a regime with a mismatch between the value of the test error and representation quality, and propose an application of our theory to active fine-tuning.
O. Duranthon, F. Boncoraglio, L. Zdeborová
Jul 24, 2026cs.LG

On the Convergence of Stochastic Low-Rank Adaptation

Low-rank adaptation (LoRA) optimizes J(B,A)=L(Wbase+sBA)J(B,A)=\mathcal L(W_\mathrm{base}+sBA) over two adapters BRm×rB \in \mathbb{R}^{m \times r} and ARr×nA \in \mathbb{R}^{r \times n} that form a low-rank update to a frozen pretrained weight matrix WbaseRm×nW_\mathrm{base} \in \mathbb{R}^{m \times n}. The prior analysis shows LoRA-GD takes exp{O(ε2)}\exp\{\mathcal{O}(ε^{-2})\} oracle calls to find an εε-stationary point such that J(B,A)ε\|\nabla J(B,A)\|\leq ε in the deterministic setting. We sharpen the analysis and show that O(ε4)\mathcal{O}(ε^{-4}) full-gradient evaluations suffice for the same first-order criterion. We further study stochastic LoRA under unbiased gradient estimates and finite variance. We propose LoRA-NSGDM, which finds an εε-stationary point with O(ε8)\mathcal{O}(ε^{-8}) stochastic oracle complexity. Under the additional mean-square smoothness condition, we use variance reduction strategy and propose LoRA-STORM, which improves the stochastic oracle complexity to O(ε6)\mathcal{O}(ε^{-6}).
Ru Wang, Chengchang Liu, John C. S. Lui
May 29, 2026cs.LG

Balanced LoRA: Removing Parameter Invariance to Accelerate Convergence

Low-Rank Adaptation (LoRA) is the most widely adopted method for fine-tuning large language models. Notably, LoRA is inherently overparameterized: multiple pairs of low-rank factors can yield the same adapted weight matrix. We show--both theoretically and empirically--that these pairs exhibit significantly different condition numbers. As a result, converging to different loss minimizers directly impacts the convergence rate of LoRA. Building on this observation, we introduce Balanced Low-Rank Adaptation (BaLoRA), a variant of LoRA that projects iterates onto a balanced manifold. This manifold improves the conditioning of the loss landscape while preserving the adapted matrix. The projection step is computationally lightweight and integrates seamlessly into existing fine-tuning pipelines. Empirically, BaLoRA converges faster than standard LoRA and achieves superior performance across a range of fine-tuning tasks.
Valérie Castin, Kimia Nadjahi, Pierre Ablin +1