Time series forecasting models are typically compared on pointwise error, which scores a prediction in isolation from the decision it is produced for, and a lower forecast error does not imply a better decision downstream. A parallel debate asks whether modern transformer architectures forecast better than recurrent and other lightweight models. We compare linear, fixed recurrent, transformer, and mixing based architectures against recurrent networks evolved by neuroevolutionary architecture search, evaluating each on forecast accuracy and on the net return of a daily long/short strategy. All models are fit on a pooled panel, one network trained across the whole universe. Across four mid-cap portfolios and three trading years, the evolved networks rank first on both forecast accuracy and net trading performance, while the second most accurate model loses money once positions are formed and costs are charged. The advantage tracks a horizon match, since rank IC for the evolved networks rises from a one-day to a ten-day scoring horizon while every model above 300 parameters declines. They are also the cheapest end to end: a CPU-only search of 16 minutes yields 66-weight networks that predict in 10.8~μs on a Raspberry Pi Zero, against transformer baselines of up to 817,153 parameters that require GPU training.
Figures & tables
Training
Validation
Test
Fold
Rows
Period
Rows
Period
Rows
Period
Test 2022
3,290
2007-12-07–2020-12-31
252
2021
251
2022
Test 2023
3,542
2007-12-07–2021-12-31
251
2022
250
2023
Table 1. Baseline Portfolio Walk-Forward Folds
Training
Validation
Test
Fold
Rows
Period
Rows
Period
Rows
Period
Test 2022
4,038 – 4,252
2004–2020-12-31
252
2021
251
2022
Test 2023
4,290 – 4,504
2004–2021-12-31
251
2022
250
2023
Test 2024
4,541 – 4,755
2004–2022-12-30
250
2023
252
2024
Table 2. Validation Portfolio Walk-Forward Folds
Test 2022
Test 2023
Construction
Rank IC
MSE
Net(%)
Rank IC
MSE
Net(%)
Individual
+0.0001
4.726e-4
+10.47
−0.0050
3.037e-4
−4.17
Concatenated
−0.0081
4.739e-4
−28.48
−0.0020
3.113e-4
−17.03
Pooled
+0.0030
4.633e-4
+12.73
+0.0195
2.926e-4
+30.24
Table 3. Pooling ablation for EXAMM on Baseline Portfolio
Other
Fixed RNNs
Transformer-based
Mixing
Panel
Test yr
EXAMM
DLinear
LSTM-i
GRU-i
PatchTST
iTransf.
DeformT.
TimeMixer
V1
2022
+0.0000
+0.0015
+0.0230
+0.0229
+0.0149
+0.0138
+0.0029
-0.0079
V1
2023
+0.0145
+0.0079
+0.0050
+0.0026
+0.0076
-0.0017
-0.0022
+0.0164
V1
2024
+0.0257
+0.0117
-0.0093
-0.0138
+0.0118
-0.0093
+0.0025
+0.0041
V2
2022
+0.0041
-0.0010
+0.0150
+0.0159
+0.0047
+0.0096
+0.0005
-0.0041
V2
2023
+0.0213
+0.0083
-0.0047
-0.0032
+0.0063
-0.0042
-0.0061
+0.0285
Table 4. Pearson information coefficient (IC) of Forecasting Performance across All Models
Benchmarks
Other
Fixed RNNs
Transformer-based
Mixing
Panel
Test yr
B&H
EW
EXAMM
DLinear
LSTM-i
GRU-i
PatchTST
iTransf.
DeformT.
TimeMixer
V1
2022
-7.16
-14.26
-4.63
+6.15
+7.13
+6.65
+3.54
+0.80
+2.48
-0.18
V1
2023
+5.18
-1.67
+21.34
-4.27
-5.20
+1.48
-10.61
-0.94
-4.64
+23.46
V1
2024
-4.73
-9.92
+17.19
-15.48
-15.49
-15.57
-23.11
-17.28
+1.73
+13.39
V2
2022
-11.42
-18.91
+35.22
+5.66
-4.05
+0.45
-1.85
+13.72
-1.61
-25.23
V2
2023
+15.37
+7.40
+16.92
+3.29
-22.25
-16.90
-10.35
-10.34
-15.06
+28.56
Table 5. Net return (%) using Model’s Predicted Return for Portfolio Trading
Figure 1. Annualized Net Alpha with 95% Confidence Intervals Alpha forest plot showing annualized net alpha with 95% confidence intervals. Only EXAMM's interval excludes zero.
Specification
α (%/yr)
t
p
MKT
+11.27
+2.32
0.021
MKT + STR1
+10.96
+2.24
0.025
MKT + STR1 + STR5
+10.64
+2.16
0.031
Table 6. EXAMM Annualized Net Alpha
Benchmarks
Other
Fixed RNNs
Transformer-based
Mixing
Test yr
B&H
EW
EXAMM
DLinear
LSTM-i
GRU-i
PatchTST
iTransf.
DeformT.
TimeMixer
2022
-10.05
-17.39
+15.46
+4.63
+7.89
+9.25
+6.62
+14.19
+0.11
+1.74
2023
+11.18
+3.99
+9.07
+6.09
-14.58
-13.22
-9.31
-4.80
-10.13
+4.31
2024
+6.80
-0.93
+8.52
-2.46
-10.18
-9.70
-1.95
-10.26
-6.80
+6.23
Mean
+2.64
-4.78
+11.02
+2.75
-5.62
-4.56
-1.55
-0.29
-5.61
+4.09
Table 7. Average Net return (%) across V1 – V4
Model
Params
IC
Net (%)
Sharpe
MDD (%)
EXAMM
66
+0.0106
+11.02
+0.74
-10.76
Other
DLinear
194
+0.0067
+2.75
+0.42
-12.49
Fixed RNNs
LSTM-i
679
+0.0035
-5.62
-0.15
-18.93
GRU-i
511
+0.0024
-4.56
-0.10
-19.39
Transformer-based
PatchTST
16,836
+0.0069
-1.55
+0.12
-15.48
iTransf.
817,153
+0.0046
-0.29
+0.17
-15.15
Table 8. Average Model Summary over Validation Portfolios.
Figure 2. Rank IC as a function of scoring horizon Plot showing Rank IC as a function of scoring horizon.
Architecture search (CPU only, 10,000 genomes)
Ranks
Workers
Wall-clock (min)
Core-h
8
7
72.8±15.6
9.70
16
15
37.1±7.7
9.88
32
31
16.2±2.8
8.65
64
63
8.1±1.9
8.59
128
127
3.9±1.0
8.26
Table 9. EXAMM cost end to end: architecture search on CPU, and on-device inference