cs.LGApr 13, 2025

Ordinary Least Squares as an Attention Mechanism

Authors: Philippe Goulet Coulombe

Organizations: Département des Sciences Économiques, Université du Québec à Montréal

Abstract

I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a learned embedding space. In this representation, least squares does not estimate coefficients per se. Instead, it selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products. This maps directly onto the query-key-value structure of attention mechanisms. I then discuss extensions to dimensionality reduction, nonlinearity, and time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that nonlinear Attention Regression performs competitively against standard machine learning baselines. In the reverse direction, I replace the attention sublayer of a transformer for tabular data with an explicit regression on polynomial features. The resulting model performs comparably to the standard transformer at a fraction of its parameter count.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

Jul 17, 2026cs.LG

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention

In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regression problems, in which the most of them implement gradient descent algorithm. However, it is still unclear whether those implementations have actually been acquired through training. In this paper, we construct a transformer with linear self-attention, which in-context learns the least squares estimate in a simple regression task. The point here is that the closed form (analytical) solution is approximately obtained by using layer normalization rather than an approximate solution based on gradient descent algorithm. Then, we show an experimental example, in which our implementation is mainly used in the transformer trained with l1 regularization when the target output is the least squares estimate.
May 6, 2026cs.LG

Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.
May 27, 2026cs.LG

PLS in the Mirror of Self-Attention

This note provides an interesting observation on casting partial least square (PLS) as a linearized self-attention so that PLS may be studied within the neural network paradigm. On the other hand, the dimensionality reduction and selection of predictors in PLS may indicate that self-attention includes certain degree of dimensionality normalization toward improved learning.