What can linear attention learn from nonlinear teachers in-context?
Authors: Mary Letey, Arman Rysmakhanov, Yue M. Lu, Cengiz Pehlevan, Jacob Zavatone-Veth
Organizations: Applied Mathematics, Harvard University Kempner Institute, Harvard University · Williams College · Center for Brain Science, Harvard University Society of Fellows, Harvard University
Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, y=f(x⊤w)+ε. Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of f, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.
Figures & tables
Figure 1 : We simulate ICL test loss directly from data to verify Result 3 . The linear target is y=w⊤x , while the nonlinear target is sin(w⊤x) . d=64 , α=2.0 , κ=2.0 , Ctrain=Id=Ctest .
Figure 2 : We see that nonlinear functions raise the limiting ICL test error, even in the best possible data case of α,κ,τ→∞ due to nonlearnable signal captured as noise. Here d=32,τ=α/2,ρ=0.1 , f(x)=tanh(x).
Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.
Alexander Hsu, Zhaiming Shen, Wenjing Liao +1
Department of Mathematics, Purdue University · School of Mathematics, Georgia Institute of Technology
In-context learning (ICL) enables a pretrained model to infer a task from demonstrations without updating its parameters. While much of the existing theory focuses on linear target functions, in this paper we study nonlinear cases by comparing two one-layer attention architectures on the same family of single-index tasks. A kernel learner first maps inputs through a fixed nonlinear feature map and then applies linear attention, whereas a feature learner applies attention to the original input, followed by a learned nonlinear readout. We derive predictions for their memorization and generalization errors using the replica method, retaining the effects of pretraining size, task-pool diversity, and training and inference context lengths. The resulting predictions closely match numerical experiments across a broad range of regimes. Our analysis yields phase diagrams that characterize when each architecture is advantageous as the amount of pretraining data, task diversity, and context lengths vary. We further identify qualitatively different context-length scalings for the two learners. Together, these results clarify how architectural choices interact with the dataset and govern nonlinear in-context learning.
Haotian Gu, Yizhou Xu, Lenka Zdeborová
Statistical Physics of Computation Laboratory, École Polytechnique Fédérale de Lausanne (EPFL) · University of Chinese Academy of Sciences · Information, Learning and Physics Laboratory, École Polytechnique Fédérale de Lausanne (EPFL)
Transformer-based large models have demonstrated remarkable generalization abilities across different tasks by leveraging a context-aware attention module for in-context learning. With richer context, transformers adapt more effectively to the current use case without any parameter updates. However, the quadratic computational and memory complexity with respect to context length significantly slows data processing in softmax transformers. Linear transformers were proposed to address this issue by reducing the complexity to linear dependence on context length, but the design and understanding of the feature mapping in linear attention, from a theoretical viewpoint, remain unclear. In this paper, we investigate the approximation and generalization abilities of linear transformers under a two-staged sampling process from domain generalization. We show that linear transformers perform in-context learning as learning a mapping from context distributions to response functions. A dimension-independent convergence rate is obtained for our generalization analysis, which also exhibits the tradeoff between the regularities of data distributions and latent features. Guided by our theoretical framework, we propose a new perspective on activation and loss design for linearizing pretrained softmax large language models.
Peilin Liu, Ding-Xuan Zhou
School of Mathematics and Statistics, University of Sydney, NSW Australia