Accurate forecasting models are usually large, expensive to update online, and fixed in architecture once trained. We apply ONE-NAS, an online neuroevolutionary architecture search that evolves a population of small recurrent networks as each window of data arrives, to daily cross-sectional stock return prediction, and pilot it on a host and endpoint pipeline: the host runs the search and ships each generation's champion genomes over TCP/IP to a Raspberry Pi 4B, which predicts online. On the Pi a single champion predicts a 50-stock window in 24.6ms and the ensemble of 40 island champions in 556ms, far inside the daily decision cycle. On four panels of US mid-cap equities over 2022--2024, reading the population as a rank-mean ensemble of island champions returns +27.5% net of realised transaction costs, against +11.3 to +14.8% for online LSTM, online GRU and monthly-retrained LSTM baselines and +4.5% for the single best genome used in prior ONE-NAS work.
Figures & tables
Figure 1: Deployment pipeline: ONE-NAS search on the host, online prediction on the Raspberry Pi endpoint.
Figure 2: On-device cost per pass of the two deployment options, on log axes; the line is the mean over ten runs and the band the min–max envelope. The island-champion ensemble runs all 40 island champions, 39× the weights of the single champion.
2022
2023
2024
2022–24
tc
ONE-NAS ( N=60 )
+6.9
+13.7
+7.0
+30.1±0.9
−1.63
ONE-NAS ( N=40 )
+6.1
+12.4
+6.3
+27.5±1.1
–
ONE-NAS ( N=20 )
+5.1
+13.3
+6.3
+27.4±1.1
0.10
Online LSTM
+1.2
+4.5
+3.8
+11.3±1.8
7.32
Online GRU
+2.1
+3.1
+5.0
+11.7±2.2
6.59
Periodic LSTM (mo.)
+1.4
+3.9
+6.9
+14.8±1.0
6.66
Table 1: Net return (%) by trade year, mean over V1 – V4 and seeds (10 for ONE-NAS; 48, 40 and 40 for the online LSTM, online GRU and periodic LSTM, Appendix A.3 ). The 2022–24 column is one run over the window, not the sum of the yearly cells.
Figure 3: Cumulative net return, 2022–2024 (books restarted at window start). Buy and hold follows the prior-study benchmark convention.
Fleet
Single
Ensemble
Δ net ( t )
Δ Sharpe ( t )
60 isl.
+1.0 / 0.09
+30.1 / 0.93
+29.1(18.3)
+0.84(14.1)
40 isl.
+4.5 / 0.21
+27.5 / 0.87
+23.0(18.2)
+0.66(14.7)
20 isl.
+4.7 / 0.25
+27.4 / 0.87
+22.6(14.9)
+0.62(14.9)
40 isl. a
+8.9 / 0.30
+36.7 / 1.22
+27.8(7.5)
+0.92(6.6)
Table 2: Prediction rule, within-run paired: single global-best genome vs. island-champion rank-mean ensemble from the same runs (Overlapping rule). Net % / Sharpe; Δ = ensemble − single, paired t over seeds.
Arm
1×
2×
3×
5×
B/E a
ONE-NAS ( N=60 )
+30.1
+27.2
+24.4
+18.8
11.6×
ONE-NAS ( N=40 )
+27.5
+24.7
+21.8
+16.0
10.6×
ONE-NAS ( N=20 )
+27.4
+24.4
+21.5
+15.6
10.3×
Periodic LSTM
+14.8
+11.4
+8.0
+1.2
5.4×
Online GRU
+11.7
+8.3
+4.9
−1.8
4.5×
Online LSTM
+11.3
+8.1
+4.9
−1.5
4.5×
Table 3: Cost sensitivity: net return (%) on 2022–2024 as the realised per-name transaction cost is scaled by the stated multiple.
Figure 4: Performance vs. island count on the evaluation span (2022–2024, Overlapping rule; 10 seeds × 4 panels at every width). Left: pooled net return (mean ± SE). Right: Sharpe, mean with ±1 seed-SD band, and the worst seed. The 40-island headline was fixed on the tuning span (Table 5 , Appendix B.1 ) before this curve was measured.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
Islands N
40 (headline); 8 (reg.)
Elites / island
8
Offspring / island / generation
5
Mut. / intra- / inter-island crossover
0.7 / 0.2 / 0.1
Repopulation
every 50 generations
Fitness
validation MSE
Appendix
Table 4: ONE-NAS configuration used in Section 5 , in the notation of Sections 3.2 and 4.2 , selected on the 2016–2019 tuning span (Appendix B.1 ). “reg.”: the registered 8-island configuration; abbreviations in parentheses as in Tables and 11 .
Islands
n
rank IC
Net%
Sharpe
MDD%
8
20
+0.0132
+26.0
0.92
11.6
16
20
+0.0149
+31.5
1.11
10.2
20
20
+0.0157
+36.0
1.24
10.4
40
40
+0.0169
+36.2
1.21
10.4
50
40
+0.0170
+36.4
1.21
10.4
60
40
+0.0168
+35.7
1.19
10.8
Appendix
Table 5: Island count on the tuning span, 2016–2019, Overlapping rule, island-champion ensemble: the span on which the headline width was chosen. The evaluation-span curve is Figure 4 .
Seed
Gen.
Weights
Infer.
Energy
Power
Idle
MSE
(end)
(ms)
(mJ)
(mW)
(mW)
42
300
143
22.8
56.9
2507
2103
1.046
43
300
145
23.5
58.5
2498
2126
1.026
44
300
138
23.6
58.3
2495
2102
1.071
45
300
145
25.3
62.9
2501
2125
1.011
46
300
155
24.2
59.4
2479
2122
1.025
Appendix
Table 6: On-device cost of the single champion (global-best genome), per seed. Each run streams 300 generations to the Raspberry Pi 4B, which evaluates the genome on that generation’s 50-stock test window. Times and energies are per pass, averaged over the run; energy is whole-board energy, and idle power is metered at the start of each run (Section 5.1 ).
Generations
Weights
Infer. (ms)
Energy (mJ)
Power (mW)
0–74
57
15.2
37.1
2453
75–149
111
23.8
58.7
2484
150–224
135
28.0
69.2
2495
225–299
152
31.5
77.8
2490
Growth
2.7 ×
2.1 ×
2.1 ×
1.02 ×
Appendix
Table 7: Single-champion cost against search progress, pooled over the ten runs of Table 6 . Each row averages one quarter of the campaign. Inference time and energy grow with the evolved network; board power is flat, so the energy growth is time, not draw (Section 5.1 ).
Seed
Gen.
Weights
Infer.
Energy
MSE
MSE
(ms)
(mJ)
(ens.)
(memb.)
45
268
4,528
527
1,589 (371)
0.975
1.018
46
300
4,317
444
1,378 (371)
0.975
1.011
47
300
4,584
618
1,796 (411)
0.975
1.019
48
300
4,521
634
1,845 (223)
0.975
1.020
49
300
4,651
602
1,746 (213)
0.975
1.016
Appendix
Table 8: On-device cost of the island-champion ensemble, per seed. A pass runs all 40 island champions and averages their predictions, so weights and costs are totals over those networks; energy is whole-board energy, with the marginal cost above idle in brackets. MSE (ens.) is the ensemble’s error and MSE (memb.) the mean error of its members. Seeds that stopped before generation 300 are averaged over the generations they completed (Section 5.1 ).
Option
Runs
Weights
Infer.
Energy
MSE
(ms)
(mJ)
Single champion
10
114
24.6
61 (7)
1.031
Island-champion ens.
6
4,492
556.2
1,652 (328)
0.975
Ratio
–
39 ×
23 ×
27 × (45 × )
0.946 ×
Appendix
Table 9: Cost of one pass for the two deployment options, pooled over all generations of all runs; weights are averaged over the campaign (the per-quarter values of Table 7 average to 114). Energy is whole-board energy, with the marginal cost above idle in brackets (Section 5.1 ).
Fleet
N
ρ
Member m
Ensemble
Predicted
20 islands
20
0.166
+0.0068
+0.0145
+0.0149
40 islands
40
0.150
+0.0065
+0.0153
+0.0156
60 islands
60
0.163
+0.0067
+0.0157
+0.0160
Appendix
Table 10: Ensemble diversity mechanism (Section 5.3 ): member rank IC m , pairwise member prediction correlation ρ , measured ensemble rank IC, and the equicorrelated averaging prediction mN/(1+(N−1)ρ) , 2022–2024, Overlapping rule.
Configuration
Δ net
Configuration
Δ net
C0 (registered)
+3.9
NTS 600
+12.8
Select: IC-gated
+0.5
Rounds 2
+11.0
Select: IC
+2.8
PER α 0.4
+4.0
Target: 5-day CS
+4.5
PER α 0.8
+11.7
5-day CS + IC-gated
+1.8
PER λ×3
+20.6
Elite 5 / gen 10
+9.6
BP epochs 20
+12.5
Appendix
Table 11: Ensembling gain (ensemble − single champion, within-run, net % on 2016–2019, Overlapping rule) inside every trained hyperparameter configuration of the tuning screen; positive in 16/16 cells (Section 5.3 ).
Rule
2022
2023
2024
2022–24
Overlapping
+6.1±0.8
+12.4±0.7
+6.3±0.5
+27.5±1.1
Banded a
+12.5±1.2
+14.6±1.3
+3.4±0.9
+33.9±1.8
Cond. L/S b
+1.3±1.6
+12.8±1.0
+10.8±1.2
+25.0±1.8
Appendix
Table 12: Trading rule: net return (%) of the ONE-NAS ensemble (40 islands) under the three rules on the same predictions and costs, mean over V1 – V4 , 10 seeds; the Overlapping row is the 40-island row of Table 1 (Section 5.4 ).
Arm
bps/day
%/yr
NW t
Beta
ONE-NAS ( N=60 )
+4.43
+11.2
+2.73
+0.120
ONE-NAS ( N=40 )
+4.05
+10.2
+2.56
+0.109
ONE-NAS ( N=20 )
+3.99
+10.0
+2.77
+0.114
Periodic LSTM (monthly)
+2.39
+6.0
+2.01
+0.046
Online LSTM
+1.93
+4.9
+1.61
+0.069
Online GRU
+1.75
+4.4
+1.91
+0.033
Appendix
Table 13: Factor-adjusted alphas: pooled daily book returns of each arm of Table 1 regressed on the Fama–French three factors plus momentum, 2022–2024, Overlapping rule, Newey–West t with lag 10 (Section 5.4 ).
Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream. A parallel debate asks whether modern transformer architectures forecast better than recurrent and other lightweight models. We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy. All models are fit on a pooled panel, one network trained across the whole universe. Across four mid-cap portfolios and three trading years, the evolved networks rank first on both forecast accuracy and net trading performance, while the second most accurate model loses money once positions are formed and costs are charged. The advantage tracks a horizon match, since rank IC for the evolved networks rises from a one-day to a ten-day scoring horizon while every model above 300 parameters declines. They are also the cheapest end to end: a CPU-only search of 16 minutes yields 66-weight networks that predict in 10.8~μs on a Raspberry Pi Zero, against transformer baselines of up to 817,153 parameters that require GPU training.
Jonathan Chang, Zimeng Lyu
Union County Magnet High School Scotch Plains, New Jersey, USA · Kean University Union, New Jersey, USA
Accurate prediction of equity returns remains a major challenge in computational finance due to the non-stationary, nonlinear, and low signal-to-noise ratio nature of financial time series. This paper proposes a hybrid two-stage architecture that combines a long short-term memory (LSTM) network with an XGBoost gradient-boosted regressor for multi-horizon stock return prediction across a diversified panel of 14 U.S. equities spanning six industry sectors. The LSTM component, comprising two stacked layers with 64 hidden units, processes 60-day sliding windows of five sequential market features to produce 64-dimensional temporal embeddings that encode learned sequential market dynamics. These embeddings are concatenated with 14 hand-crafted technical indicators to form a 78-dimensional hybrid feature vector, which is subsequently passed to an XGBoost regressor tuned via 3-fold cross-validation grid search. The framework is trained on a multi-stock pooled corpus using strict chronological splits and per-stock MinMaxScaling to prevent look-ahead bias, and evaluated across four prediction horizons of 30, 90, 252, and 365 trading days. Experimental results demonstrate that the hybrid model achieves a test RMSE of 0.0949 on the 30-day horizon, roughly one-third that of the standalone LSTM baseline, while marginally matching or surpassing the XGBoost-Only baseline across the majority of stocks. Directional accuracy rises with horizon length, reaching 97.6% at 365 days; we show, however, that this largely tracks the high base rate of positive long-horizon returns in the sample, and we therefore benchmark directional accuracy against a naive always-positive predictor and treat the above-base-rate gap at short horizons as the more informative signal. A composite investment scoring framework derived from multi-horizon predictions is further proposed to support portfolio ranking and decision support.
Prediction-based approaches are widely used in neural architecture search (NAS), where a predictor estimates the performance of candidate architectures to guide selection. However, existing predictors are typically trained via supervised regression on limited samples, leading to overfitting and poor generalization to unseen architectures. In this work, we propose a fundamentally different formulation that models performance prediction as a conditional function inference problem using a Convolutional Neural Process (ConvNP) with meta-learning capabilities. Instead of fitting a fixed mapping to limited samples, our approach meta-learns to infer performance from partial observations by training with context-target splits across a group of synthesized tasks, explicitly optimizing for generalization under data scarcity and aligning the training procedure with the deployment setting in NAS. We further design simple yet effective meta-features for cell-based architectures and evaluate our method on NAS-Bench-101 and NAS-Bench-201. Extensive experiments show that our approach consistently improves top-K ranking quality and achieves the state-of-the-art architecture selection using limited samples.
Liping Deng, MingQing Xiao
Department of Mathematics University of California, Riverside Riverside, CA 92507 USA · School of Mathematical and Statistical Sciences Southern Illinois University Carbondale Carbondale, IL 62901 USA