I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a learned embedding space. In this representation, least squares does not estimate coefficients per se. Instead, it selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products. This maps directly onto the query-key-value structure of attention mechanisms. I then discuss extensions to dimensionality reduction, nonlinearity, and time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that nonlinear Attention Regression performs competitively against standard machine learning baselines. In the reverse direction, I replace the attention sublayer of a transformer for tabular data with an explicit regression on polynomial features. The resulting model performs comparably to the standard transformer at a fraction of its parameter count.
Figures & tables
Figure 1: Roadmap: Three equivalent views of OLS predictions. Left: Coefficient-based formulation. Middle: Similarity-based interpretation—predictions as weighted averages of training outcomes in a transformed space. Right: Restricted attention module with identity activation.
Figure 2: Monte Carlo: average out-of-sample R2 across N∈{500,1000,2500,5000} and SNR ∈{0.5,1,2,3} for each DGP. Error bars are ±1 standard deviation across the 16 conditions per DGP. Full per-condition table in Appendix B .
Baselines
Contributions
Dataset
N
P
OLS
RF
MLP
FT-T
Att. Reg
Reg. Block
California
5000
8
0.597
0.777
0.763
0.764
0.738
0.766
Yacht
308
6
0.562
0.979
0.970
0.988
0.958
0.989
Energy
768
8
0.913
0.996
0.994
0.994
0.995
0.995
Concrete
1030
8
0.624
0.892
0.895
0.898
0.837
0.907
Airfoil
1503
5
0.497
0.910
0.917
0.924
0.757
0.927
Table 1: Out-of-sample R2 on UCI/OpenML tabular regression benchmarks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Models
DGP
N
SNR
OLS
RF
MLP
GBM
Att Reg
Linear
500
0.5
0.32
0.27
0.28
0.26
0.27
1
0.49
0.43
0.45
0.44
0.45
2
0.66
0.59
0.63
0.61
0.64
3
0.75
0.67
0.72
0.70
0.73
1000
0.5
0.32
0.27
0.31
0.28
0.29
Appendix
Table 2 : Model Performance Comparison (Out-of-Sample R2 ) Across Data Generating Processes
Dataset
FT-T
Reg. Block
Difference
SE
p
California
0.764
0.766
+0.0020
0.0062
0.76
Yacht
0.988
0.989
+0.0013
0.0064
0.85
Energy
0.994
0.995
+0.0002
0.0013
0.86
Concrete
0.898
0.907
+0.0091
0.0071
0.27
Airfoil
0.924
0.927
+0.0030
0.0104
0.78
Abalone
0.324
0.329
+0.0054
0.0048
0.32
Appendix
Table 3: Paired comparison of the Regression Block and the FT-Transformer.
Dataset
A0
A1
A2
A3
A4
A5
A6
A7
RB
T
California
0.764
0.769
0.701
0.757
0.728
0.772
0.736
0.768
0.766
0.759
Yacht
0.988
0.981
0.572
0.990
0.538
0.982
0.968
0.989
0.989
0.989
Energy
0.994
0.994
0.941
0.977
0.922
0.980
0.991
0.995
0.995
0.988
Concrete
0.899
0.898
0.806
0.874
0.814
0.873
0.888
0.908
0.907
0.880
Airfoil
0.924
0.926
0.524
0.908
0.865
0.913
0.933
0.921
0.929
0.912
Abalone
0.324
0.314
0.286
0.324
0.328
0.324
0.310
0.330
0.329
0.327
Appendix
Table 4: Ablation grid, out-of-sample R2 .
Dataset
Ntrain
P
OLS
RF
FT-T
Reg. Block
TabPFN
TabICL
Panel A: Full sample size
California
16,512
8
0.601
0.817
0.807
0.805
0.872
0.878
Kin8nm
6,553
8
0.410
0.691
0.924
0.931
0.935
0.935
Protein
36,584
9
0.275
0.678
0.617
0.613
0.773
0.779
Panel B: Higher dimensions (capped at N=5000 )
CPU_act
4,000
21
0.708
0.978
0.977
0.973
0.985
0.985
Appendix
Table 5: Scope of the regression comparison, out-of-sample R2 .
Dataset
LogReg
RF
FT-T
Reg. Block
TabPFN
TabICL
BreastCancer
0.977
0.949
0.963
0.967
0.984
0.981
Diabetes
0.779
0.781
0.774
0.775
0.781
0.781
Phoneme
0.752
0.903
0.872
0.857
0.901
0.914
Wine
0.967
0.972
0.961
0.967
0.983
0.994
Vehicle
0.794
0.748
0.782
0.846
0.881
0.892
Segment
0.938
0.973
0.968
0.974
0.988
0.990
Appendix
Table 6: Classification accuracy on six OpenML datasets.
In-context learning is a remarkable property of transformers and has recently received a lot of interest. In many studies of in-context learning, it has been shown that transformers are capable of implementing solver for linear and non-linear regression problems, in which the most of them implement gradient descent algorithm. However, it is still unclear whether those implementations have actually been acquired through training. In this paper, we construct a transformer with linear self-attention, which in-context learns the least squares estimate in a simple regression task. The point here is that the closed form (analytical) solution is approximately obtained by using layer normalization rather than an approximate solution based on gradient descent algorithm. Then, we show an experimental example, in which our implementation is mainly used in the transformer trained with l1 regularization when the target output is the least squares estimate.
Katsuyuki Hagiwara
Faculty of Education, Mie University · Kurima-Machiya-cho, Tsu, 514-8507, Japan
Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, we study ICL in the nonlinear regression setting. Through the interaction mechanism in attention, we explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, we establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features. Our theory provides finite-sample generalization error bounds in terms of context length and training set size. We numerically validate the theory on synthetic regression tasks.
Alexander Hsu, Zhaiming Shen, Wenjing Liao +1
Department of Mathematics, Purdue University · School of Mathematics, Georgia Institute of Technology
This note provides an interesting observation on casting partial least square (PLS) as a linearized self-attention so that PLS may be studied within the neural network paradigm. On the other hand, the dimensionality reduction and selection of predictors in PLS may indicate that self-attention includes certain degree of dimensionality normalization toward improved learning.