stat.MLOct 7, 2026

What can linear attention learn from nonlinear teachers in-context?

Authors: Mary Letey, Arman Rysmakhanov, Yue M. Lu, Cengiz Pehlevan, Jacob Zavatone-Veth

Organizations: Applied Mathematics, Harvard University Kempner Institute, Harvard University · Williams College · Center for Brain Science, Harvard University Society of Fellows, Harvard University

Abstract

Linear attention is a tractable model for understanding the mechanisms governing in-context learning in transformers. For linear regression tasks, recent asymptotic analyses have characterised its learning and generalisation behaviour. We extend this theory to nonlinear single-index targets, y=f(x⊤w)+εy=f(x^\top w)+\varepsilon . Our main result establishes a nonlinearity-noise equivalence: linear attention extracts only the linear Hermite component of ff, while the remaining nonlinear structure contributes to the generalisation error as effective noise. This reduction allows results from the corresponding linear theory to be transferred to nonlinear tasks. We illustrate its implications for finite pretraining data and for the transition from task memorisation to task generalisation as task diversity increases. These results identify a limitation of the reduced linear-attention model and provide a tractable starting point for studying nonlinear in-context learning.

Figures & tables

Explore similar work

CardsList
  1. Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

    May 6, 2026Alexander Hsu, Zhaiming Shen, Wenjing Liao +1Transformer AttentionIn-Context Learning

  2. In-context Learning of Single-index Targets: Comparing Kernel and Feature Learners

    Oct 1, 2026Haotian Gu, Yizhou Xu, Lenka ZdeborováRepresentation LearningSingle-Index Models

  3. Ghost in the Kernel: In-Context Learning with Efficient Transformers via Domain Generalization

    Jul 1, 2026Peilin Liu, Ding-Xuan ZhouLinear AttentionIn-Context Learning