Deep randomized models fix hidden-layer parameters through random initialization and learn only closed-form readouts, typically adding depth by stacking random trans formations without target-aware control of hidden-state evolution. We propose LAIR Net, the Leaky Alignment-Impulse Residual Network, which mixes a shallow learned anchor into each hidden state through a leaky residual transition. We derive a depth uniform bound on input-perturbation sensitivity and use controlled simulations to attribute gains over a randomized baseline to the anchor rather than recursion or added capacity. Benefits emerge when a nonlinear target structure is learnable at the available noise level and diminish for nearly linear targets or dominant noise. Across 23 benchmark datasets, LAIR-Net achieves the best average rank among eight randomized networks and twelve conventional models, with relative performance associated with the same nonlinear-structure and noise quantities identified in simulation.
Figures & tables
Figure 1: A single LAIR-Net block. The layer input contains the bias, the observed inputs, and the previous hidden state. A fixed random transformation produces H(ℓ) . The leaky transition combines this representation with H(ℓ−1) using γ . The alignment impulse then combines the resulting state with S using α .
Figure 2: The LAIR-Net architecture. The SLAN produces the anchor S , which is supplied to every randomized layer. Each LAIR-Net block applies a fixed random transformation, a leaky transition, and an alignment impulse. A ridge readout is fitted at every depth, and the resulting layerwise predictions are combined using A .
Figure 3: The region in which the propagation factor of ( 11 ) is below one, for four values of the state gain c1 . The condition holds above each curve. Shading marks the region for c1=8 , the most demanding case shown. The curves are the identity ρ=(1−α)(1+γ(c1−1)) rather than a measured quantity, so the figure describes the scope of Theorem 4.4 and not the behaviour of any fitted model.
Dataset
RVFL
dRVFL
edRVFL
edRVFL-SC
Fuzzy RVFL
RVFL-Bag
RVFL-Boost
LAIR-Net
Airfoil
7.943 ± 2.342
4.128 ± 1.060
5.589 ± 1.397
3.391 ± 0.952
17.553 ± 3.613
10.315 ± 1.574
6.716 ± 1.462
2.208 ± 0.485
Auto MPG
7.368 ± 1.936
7.096 ± 1.885
7.613 ± 1.886
7.407 ± 2.537
8.520 ± 2.138
7.160 ± 2.107
6.975 ± 1.662
6.919 ± 2.074
Autos
0.020 ± 0.006
0.021 ± 0.005
0.019 ± 0.006
0.019 ± 0.006
0.025 ± 0.007
0.017 ± 0.007
0.020 ± 0.008
0.021 ± 0.010
Breast Cancer
869.556 ± 310.974
896.296 ± 291.201
917.961 ± 268.373
901.050 ± 317.573
877.582 ± 279.775
896.660 ± 290.495
908.803 ± 307.653
871.580 ± 265.293
Challenger
0.495 ± 0.470
0.466 ± 0.490
0.975 ± 1.160
0.510 ± 0.543
0.418 ± 0.433
0.480 ± 0.415
0.329 ± 0.252
0.389 ± 0.448
Concrete
56.708 ± 10.841
56.480 ± 11.082
48.850 ± 9.629
50.208 ± 11.808
92.994 ± 13.458
55.531 ± 11.564
54.718 ± 11.492
40.199 ± 10.883
Table 1: Test mean squared error against the randomized-network family, mean ± standard deviation over ten folds. Lower is better and the best entry in each row is in bold. The final rows give the average rank within this family and the Holm-adjusted Wilcoxon signed-rank p -value against LAIR-Net.
Figure 4: Nemenyi comparison within the randomized-network family. Each point is a mean rank with its interval, and the shaded band marks the region not separated from the best-ranked model at the 0.05 level, with CD=2.19 over 23 datasets. Lower is better.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Where the advantage appears. Each curve is the median ratio of edRVFL to LAIR-Net test error over ten seeds at one noise level, with the alignment coefficient fixed at α=0.5 , and the band is a 95% percentile bootstrap over those seeds. The dashed line at one is the crossover. At σ=0.1 the curve crosses between κ=0 and κ=0.25 , at σ=0.5 it reaches one only at κ=0.75 , and at σ=1.0 it does not cross within the grid.
κ
σ=0.1
σ=0.5
σ=1.0
0.00
0.852
0.614
0.591
0.25
2.106
0.751
0.625
0.50
2.522
0.875
0.661
0.75
2.639
1.005
0.693
1.00
2.991
1.122
0.729
Appendix
Table 2: Median ratio of edRVFL to LAIR-Net test mean squared error over ten seeds, with α=0.5 fixed. A ratio above one favours LAIR-Net. Rows vary the nonlinear share κ of Equation ( 13 ); columns vary the noise level σ .
Figure 6: Decomposition of the edRVFL-to-LAIR-Net error ratio across the (κ,σ) grid. The architecture term compares edRVFL with an α=0 arm that keeps the residual layer input and the leaky transition; the anchor term compares that arm with the full model. Values above one favour the model on the right of each comparison.
Figure 7: Effect of selecting α on a held-out validation split, with all other hyperparameters fixed. Left, the edRVFL-to-LAIR-Net error ratio with α selected; right, the ratio of the selected model’s error to that of the best α in the grid. A ratio near one on the right indicates that validation recovered the best available setting.
Figure 8: What validation selection of α recovers, and what it costs. Left, the selected α against the best available one, with marker area proportional to the number of seeds landing on that pair and the dashed line marking exact agreement. Right, the ratio of the selected model’s error to the error at the best available α , by nonlinear share, with 95% percentile bootstrap bands over seeds. Points on the dashed line on the right lose nothing by selecting.
Figure 9: Alignment selection at N=250 . Left, the error of the selected α relative to the best in the grid; right, the edRVFL-to-LAIR-Net ratio at the best α . Effect sizes are smaller than at N=1000 and the low-noise cells retain the clearest separation.
Figure 10: Relative performance of LAIR-Net against the best randomized baseline on each benchmark dataset, plotted against the nonlinearity proxy κ . The vertical axis is log(MSELAIR−Net/MSEbest randomized) , so points below the horizontal line at zero favour LAIR-Net. The gas dataset is marked as an off-scale point and is discussed in the text.
Dataset
Acronym
N
d
Provenance note
Airfoil Self-Noise
Airfoil
1503
5
( Brooks et al., 1989 )
Auto MPG
Auto MPG
392
7
( Quinlan, 1993a ) ; local rows reflect preprocessing
Automobile
Autos
159
25
( Kibler et al., 1989 ) ; local shape reflects encoding
Breast Cancer Prognostic
Breast Cancer
194
33
( Street et al., 1995 )
Challenger O-Ring
Challenger
23
4
( Draper, 1995 )
Concrete Strength
Concrete
1030
8
( Yeh, 1998 )
Appendix
Table 3: Dataset provenance targets for the benchmark suite. The local shape is the shape used by the preprocessed benchmark files.
Figure 11: Nemenyi comparison against conventional models, drawn as in Figure 4 , with CD=3.47 over 23 datasets. Lower is better.
Dataset
MLP
DT
RF
AdaBoost
GBM
SVR
KNN
LR
Ridge
Lasso
ElasticNet
LAIR-Net
Airfoil
5.315 ± 1.512
10.705 ± 2.487
7.501 ± 1.896
21.337 ± 4.179
16.551 ± 2.630
9.192 ± 1.916
3.868 ± 0.786
23.378 ± 5.534
23.378 ± 5.532
23.378 ± 5.532
23.375 ± 5.530
2.208 ± 0.485
Auto MPG
7.483 ± 1.514
12.560 ± 2.468
8.538 ± 2.757
11.207 ± 3.339
9.825 ± 2.377
6.734 ± 1.975
7.880 ± 2.700
11.266 ± 2.480
11.252 ± 2.453
11.399 ± 2.619
11.400 ± 2.525
6.919 ± 2.074
Autos
0.027 ± 0.013
0.034 ± 0.011
0.029 ± 0.015
0.031 ± 0.015
0.027 ± 0.018
0.028 ± 0.010
0.031 ± 0.022
0.027 ± 0.010
0.024 ± 0.007
0.027 ± 0.013
0.027 ± 0.012
0.021 ± 0.010
Breast Cancer
939.197 ± 333.499
1.06e3 ± 253.065
832.325 ± 208.431
839.905 ± 229.339
741.678 ± 162.746
911.205 ± 290.348
1.02e3 ± 236.385
944.958 ± 364.476
896.263 ± 308.580
914.506 ± 276.342
874.418 ± 257.062
871.580 ± 265.293
Challenger
0.375 ± 0.436
0.407 ± 0.365
0.407 ± 0.370
0.480 ± 0.556
0.415 ± 0.358
0.317 ± 0.314
0.428 ± 0.397
0.365 ± 0.408
0.359 ± 0.391
0.414 ± 0.451
0.395 ± 0.432
0.389 ± 0.448
Concrete
43.674 ± 12.915
74.354 ± 26.260
46.642 ± 9.922
79.306 ± 15.171
65.111 ± 10.004
42.791 ± 9.954
58.822 ± 11.054
110.468 ± 12.515
110.461 ± 12.455
110.905 ± 11.686
110.962 ± 11.709
40.199 ± 10.883
Appendix
Table 4: Test mean squared error against conventional neural, tree-based, kernel, instance-based and linear models, mean ± standard deviation over ten folds. Lower is better and the best entry in each row is in bold. The final rows give the average rank within this family and the Holm-adjusted Wilcoxon signed-rank p -value against LAIR-Net.
Model
Avg. rank
p
pHolm
LAIR-Net
4.22
–
–
MLP
6.70
0.0078
0.0389
edRVFL-SC
7.17
0.0037
0.0221
RVFL-Boost
7.35
0.0010
0.0106
RVFL-Bag
8.00
0.0022
0.0157
edRVFL
8.35
0.0001
0.0022
Appendix
Table 5: Average rank and paired Wilcoxon signed-rank tests against LAIR-Net with all nineteen models pooled, over the 23 datasets. Lower rank is better. p -values are Holm-adjusted over the eighteen comparisons. Ranks here are not comparable with the within-family tables, since each model is ranked against a different set.
Model
Average rank
Dataset wins
LAIR-Net
2.565
3
LAIR-SLANInit
3.043
7
LAIR-Ortho
3.783
4
LAIR-OrthoBMA
3.913
4
LAIR-NoAlign
4.652
3
LAIR-BestState
4.696
2
Appendix
Table 6: Average ranks and per-dataset wins within the variant family. Lower average rank and higher win count are better.
Hyperparameter
Search space
Shared by all randomized models, including LAIR-Net
m
{10,20,…,200}
Input scale
U(0,1)
λ
logU(10−8,102)
Activation
{ReLU,tanh}
LAIR-Net-specific
Appendix
Table 7: Hyperparameter search spaces for the randomized models.
Model
Parameter
Range
MLP
hidden units
{8,16,32,64}
activation
{identity,logistic,tanh,ReLU}
learning-rate schedule
{constant,invscaling,adaptive}
initial learning rate
{10−1,10−2,10−3,10−4}
Decision tree
criterion
{squared,absolute,Friedman}
splitter
{best,random}
Appendix
Table 8: Hyperparameter search space for the conventional models. Ridge and lasso regression are fitted by their own internal cross-validation over the regularisation path rather than by the outer search, and linear regression has no hyperparameter, so none of the three enters this table.