cs.LGApr 13, 2025

Ordinary Least Squares as an Attention Mechanism

Authors: Philippe Goulet Coulombe

Organizations: Département des Sciences Économiques, Université du Québec à Montréal

Abstract

I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a learned embedding space. In this representation, least squares does not estimate coefficients per se. Instead, it selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products. This maps directly onto the query-key-value structure of attention mechanisms. I then discuss extensions to dimensionality reduction, nonlinearity, and time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that nonlinear Attention Regression performs competitively against standard machine learning baselines. In the reverse direction, I replace the attention sublayer of a transformer for tabular data with an explicit regression on polynomial features. The resulting model performs comparably to the standard transformer at a fraction of its parameter count.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention

    Jul 17, 2026Katsuyuki HagiwaraIn-Context LearningTransformer Architectures

  2. Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

    May 6, 2026Alexander Hsu, Zhaiming Shen, Wenjing Liao +1In-Context LearningTransformer Architectures