I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. The connection comes from viewing OLS as a similarity-based prediction rule in a learned embedding space. In this representation, least squares does not estimate coefficients per se. Instead, it selects an embedding that minimizes squared prediction error by matching training and test vectors through inner products. This maps directly onto the query-key-value structure of attention mechanisms. I then discuss extensions to dimensionality reduction, nonlinearity, and time series econometrics. Monte Carlo simulations and real-data experiments on UCI/OpenML benchmarks show that nonlinear Attention Regression performs competitively against standard machine learning baselines. In the reverse direction, I replace the attention sublayer of a transformer for tabular data with an explicit regression on polynomial features. The resulting model performs comparably to the standard transformer at a fraction of its parameter count.
Figures & tables
Figure 1: Roadmap: Three equivalent views of OLS predictions. Left: Coefficient-based formulation. Middle: Similarity-based interpretation—predictions as weighted averages of training outcomes in a transformed space. Right: Restricted attention module with identity activation.
Figure 2: Monte Carlo: average out-of-sample R2 across N∈{500,1000,2500,5000} and SNR ∈{0.5,1,2,3} for each DGP. Error bars are ±1 standard deviation across the 16 conditions per DGP. Full per-condition table in Appendix B .
Baselines
Contributions
Dataset
N
P
OLS
RF
MLP
FT-T
Att. Reg
Reg. Block
California
5000
8
0.597
0.777
0.763
0.764
0.738
0.766
Yacht
308
6
0.562
0.979
0.970
0.988
0.958
0.989
Energy
768
8
0.913
0.996
0.994
0.994
0.995
0.995
Concrete
1030
8
0.624
0.892
0.895
0.898
0.837
0.907
Airfoil
1503
5
0.497
0.910
0.917
0.924
0.757
0.927
Table 1: Out-of-sample R2 on UCI/OpenML tabular regression benchmarks.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Models
DGP
N
SNR
OLS
RF
MLP
GBM
Att Reg
Linear
500
0.5
0.32
0.27
0.28
0.26
0.27
1
0.49
0.43
0.45
0.44
0.45
2
0.66
0.59
0.63
0.61
0.64
3
0.75
0.67
0.72
0.70
0.73
1000
0.5
0.32
0.27
0.31
0.28
0.29
Appendix
Table 2 : Model Performance Comparison (Out-of-Sample R2 ) Across Data Generating Processes
Dataset
FT-T
Reg. Block
Difference
SE
p
California
0.764
0.766
+0.0020
0.0062
0.76
Yacht
0.988
0.989
+0.0013
0.0064
0.85
Energy
0.994
0.995
+0.0002
0.0013
0.86
Concrete
0.898
0.907
+0.0091
0.0071
0.27
Airfoil
0.924
0.927
+0.0030
0.0104
0.78
Abalone
0.324
0.329
+0.0054
0.0048
0.32
Appendix
Table 3: Paired comparison of the Regression Block and the FT-Transformer.
Dataset
A0
A1
A2
A3
A4
A5
A6
A7
RB
T
California
0.764
0.769
0.701
0.757
0.728
0.772
0.736
0.768
0.766
0.759
Yacht
0.988
0.981
0.572
0.990
0.538
0.982
0.968
0.989
0.989
0.989
Energy
0.994
0.994
0.941
0.977
0.922
0.980
0.991
0.995
0.995
0.988
Concrete
0.899
0.898
0.806
0.874
0.814
0.873
0.888
0.908
0.907
0.880
Airfoil
0.924
0.926
0.524
0.908
0.865
0.913
0.933
0.921
0.929
0.912
Abalone
0.324
0.314
0.286
0.324
0.328
0.324
0.310
0.330
0.329
0.327
Appendix
Table 4: Ablation grid, out-of-sample R2 .
Dataset
Ntrain
P
OLS
RF
FT-T
Reg. Block
TabPFN
TabICL
Panel A: Full sample size
California
16,512
8
0.601
0.817
0.807
0.805
0.872
0.878
Kin8nm
6,553
8
0.410
0.691
0.924
0.931
0.935
0.935
Protein
36,584
9
0.275
0.678
0.617
0.613
0.773
0.779
Panel B: Higher dimensions (capped at N=5000 )
CPU_act
4,000
21
0.708
0.978
0.977
0.973
0.985
0.985
Appendix
Table 5: Scope of the regression comparison, out-of-sample R2 .
Dataset
LogReg
RF
FT-T
Reg. Block
TabPFN
TabICL
BreastCancer
0.977
0.949
0.963
0.967
0.984
0.981
Diabetes
0.779
0.781
0.774
0.775
0.781
0.781
Phoneme
0.752
0.903
0.872
0.857
0.901
0.914
Wine
0.967
0.972
0.961
0.967
0.983
0.994
Vehicle
0.794
0.748
0.782
0.846
0.881
0.892
Segment
0.938
0.973
0.968
0.974
0.988
0.990
Appendix
Table 6: Classification accuracy on six OpenML datasets.