Are We Really Benchmarking Forecasting Models? The Impact of Preprocessing on Time Series Performance
Authors: Guilherme Afonso Galindo Padilha, Paulo Salgado Gomes de Mattos Neto, Rafael Menelau Oliveira e Cruz
Organizations: LIVIA, Department of Software and IT Engineering École de technologie supérieure (ÉTS) Montreal, QC, Canada · Center for Informatics (CIn) Universidade Federal de Pernambuco (UFPE) Recife, PE, Brazil
While established literature underscores the pivotal role of preprocessing in forecasting accuracy, this stage remains largely overlooked in current research. Modern benchmarks typically resort to simple scaling, failing to account for critical transformations required to address nonstationarity, such as differencing. This omission creates a significant structural preprocessing bias that favors models with built-in data treatments while obscuring the true potential of simpler architectures. We study this effect through a preprocessing-aware benchmark that evaluates 11 forecasting models across 16 reversible preprocessing pipelines on 29,000 M4 time series. Our results identify preprocessing as a key driver of forecasting performance. Optimizing preprocessing per series yields gains of approximately 27% to 87% across all evaluated models, with architectures lacking internalized preprocessing experiencing the most substantial improvements. This allows simpler architectures to become highly competitive with complex, state-of-the-art models in modern forecasting benchmarks. All resources and experimental results from this benchmark are stored in a comprehensive metadataset to support future metalearning tasks.
Figures & tables
Figure 1: Forecasts from an MLP and an ARMA without preprocessing. (a) Stationary AR(1): both models track the series well. (b) Trend nonstationarity: the rising level leads forecasts to lag the mean learned in training. (c) Variance nonstationarity: the mean is captured but larger fluctuations inflate absolute errors. (d) Random walk: persistent level changes cause drift and error accumulation. (e) Seasonality: without seasonal modeling, forecasts revert to the global mean and miss cyclic fluctuations. (f) Multi: multiple interacting forms of nonstationarity amplify drift and mismatch, causing the largest forecast degradation.
Figure 2: Preprocessing-Aware Time Series Benchmark framework and Metadataset.
Dataset
Freq.
# Series
Avg. Length
# Obs.
Lag
Horizon
M4 Quarterly
Quarterly
24,000
100
2,406,108
4
8
M4 Weekly
Weekly
359
1,035
371,579
52
13
M4 Daily
Daily
4,227
2,371
10,023,836
7
14
M4 Hourly
Hourly
414
902
373,372
24
48
Table 2: Main characteristics of the M4 datasets used in this study. Lag and horizon are reported in the number of time steps.
Pipeline
ARIMA
ETS
THETA
LR
XGBst
MLP
NBTS
TNet
PTST
Informer
Chronos
X
2.35
2.25
2.22
2.87
2.75
31.88
4.65
2.65
2.77
106.75
2.29
V
3.18
2.38
2.36
3.75
2.61
4.27
4.78
2.45
2.48
46.12
2.56
U
2.23
2.40
2.60
2.27
2.34
3.10
5.91
3.08
3.16
3.12
2.09
S
2.38
3.19
3.19
2.71
2.49
4.06
5.78
3.10
3.42
4.45
2.46
Q
2.30
2.18
2.22
2.87
3.71
4.86
4.86
2.54
2.66
5.51
2.29
VU
2.23
2.45
2.84
2.45
2.18
2.87
1.48e24
3.35
3.47
3.21
2.17
Table 3: Forecasting performance (MASE) across preprocessing pipelines and models. Best per model in bold. Columns in grey are IP-Models. Pipelines are composed of: X (None), V (Boxcox transformation), U (First-order differencing), S (Seasonal differencing), and Q (Standard scaling).
Figure 3: Barplot MASE of best and oracle pipelines. Lower loss values mean better results.
Figure 6
Trend
Seasonality
Stationarity
Shifting
Entropy
Model
High
Low
High
Low
Stat.
Non-Stat.
High
Low
High
Low
ARIMA
VUQ
VS
VUQ
VUQ
VUQ
VUQ
VUQ
VUQ
VQ
VUQ
ETS
X
Q
Q
X
V
VQ
Q
VQ
V
V
THETA
X
X
X
X
X
X
X
VQ
Q
Q
LR
VU
V
VU
VU
UQ
VU
VU
Q
V
VUQ
XGBoost
VUQ
VQ
VSQ
VUQ
SQ
VUQ
VUQ
VUQ
VQ
VUQ
Table 5: Best performing preprocessing pipelines (MASE) for each model under different statistical properties. IP-Models are highlighted in gray.
Figure 5: Mean MASE vs. mean training time for local models. Marker size indicates peak GPU memory (baselined at 250MB for CPU-only models).
Appendix figures & tables17 assets
Supplementary material from the paper’s appendix.
Appendix
#
Feature Name
Description
1
mean
Mean value of the time series.
2
var
Variance of the time series.
3
max_kl_shift
Maximum Kullback-Leibler divergence detected across windows.
4
time_kl_shift
Relative time index where the maximum KL shift occurs.
5
max_level_shift
Maximum change in mean level between windows.
6
time_level_shift
Relative time index where the maximum level shift occurs.
Appendix
Table 6: Description of Time Series Features extracted via tsfeatures.
Hyperparameter
Linear Regression
XGBoost
Library
scikit-learn
XGBoost
Max Depth
N/A
6
Estimators
N/A
1000
Subsample
N/A
0.8
Colsample bytree
N/A
0.8
L2 Regularization ( λ )
1.0
1.0
Appendix
Table 7: Hyperparameters for Machine Learning Baselines
Hyperparameter
MLP
N-BEATS
TimesNet
Architecture
2 Hidden Layers
6 Stacks / 4 Layers
2 Encoder Layers
Hidden Dim ( dmodel )
512
256 (Layer Size)
16
FFN Dim ( dff )
N/A
N/A
32
Activation
ReLU
ReLU
GeLU
Appendix
Table 8: Hyperparameters for Deep Learning Models
Hyperparameter
PatchTST
Informer
Layers
3 Encoder
2 Encoder / 1 Decoder
Attention Heads
16
4
Model Dim ( dmodel )
128
256
FFN Dim ( dff )
256
512
Activation
GeLU
GeLU
Normalization
RevIN
Standard
Appendix
Table 9: Hyperparameters for Transformer-based Models
Pipeline
ARIMA
ETS
THETA
LR
XGBst
MLP
NBTS
TNet
PTST
Informer
X
2.35
2.25
2.22
2.89
4.55
2.94
2.48
2.29
2.27
63.59
V
3.19
2.38
2.36
3.77
4.54
6.42
2.49
2.26
2.26
18.23
U
2.23
2.40
2.60
2.28
2.73
3.01
3.02
3.27
3.08
2.35
S
2.38
3.19
3.19
2.71
3.26
3.47
3.59
3.35
3.10
3.04
Q
2.30
2.18
2.22
2.89
4.55
3.97
3.76
2.27
2.27
4.95
VU
2.23
2.45
2.84
2.45
2.87
2.58
3.19
3.83
3.39
4.74
Appendix
Table 10: Forecasting performance (MASE) across preprocessing pipelines and models when trained on each series locally. Best per model in bold. Columns in grey are IP-Models. Pipelines are composed of: X (Baseline/None), V (Boxcox transformation), U (First-order differencing), S (Seasonal differencing), and Q (Standard scaling).
Figure 6: Barplot MASE of best and oracle pipelines of local models. Lower loss values mean better results.
Pipeline
ARIMA
ETS
THETA
LR
XGBst
MLP
NBTS
TNet
PTST
Informer
Chronos
X
1.94e6
9.96e6
1.63e6
7.07e6
2.21e6
3.23e7
3.10e6
1.95e6
1.93e6
5.01e7
1.60e6
V
2.35e6
2.14e6
1.91e6
9.16e8
2.81e6
1.19e7
3.14e9
1.78e6
1.92e6
6.26e7
1.73e6
U
1.78e6
5.17e6
4.76e6
2.40e6
1.79e6
2.76e6
8.88e7
4.63e6
6.29e6
2.62e6
1.96e6
S
2.00e6
3.82e6
3.99e6
1.90e7
1.83e6
2.18e6
5.23e6
3.65e6
4.85e6
2.10e6
2.24e6
Q
1.71e6
3.53e6
1.64e6
7.07e6
2.85e6
3.04e6
4.07e6
1.77e6
1.95e6
3.81e6
1.60e6
VU
2.01e6
2.93e6
3.01e7
5.45e6
1.72e6
2.18e6
1.18e58
7.98e8
9.21e8
2.53e6
2.14e6
Appendix
Table 11: Forecasting performance (MSE) across preprocessing pipelines and models. Best per model in bold. Columns in grey are IP-Models. Pipelines are composed of: X (Baseline/None), V (Boxcox transformation), U (First-order differencing), S (Seasonal differencing), and Q (Standard scaling).
Pipeline
ARIMA
ETS
THETA
LR
XGBst
MLP
NBTS
TNet
PTST
Informer
Chronos
X
618.42
648.65
591.69
721.70
661.90
1423.34
724.41
670.77
689.61
3891.09
593.42
V
674.92
610.30
615.76
949.83
685.92
786.98
1243.25
642.18
652.60
4349.75
624.72
U
602.66
661.07
736.04
621.58
615.27
680.22
2082.37
882.41
942.62
680.39
593.82
S
624.96
804.39
821.39
733.01
629.87
718.31
887.38
806.01
919.89
706.19
657.58
Q
610.05
608.39
594.12
721.61
854.76
923.15
1025.95
652.68
672.63
1026.64
593.26
VU
607.93
667.51
815.84
668.15
604.06
685.63
6.40e26
1446.61
1473.95
711.34
614.74
Appendix
Table 12: Forecasting performance (RMSE) across preprocessing pipelines and models. Best per model in bold. Columns in grey are IP-Models. Pipelines are composed of: X (Baseline/None), V (Boxcox transformation), U (First-order differencing), S (Seasonal differencing), and Q (Standard scaling).
Pipeline
ARIMA
ETS
THETA
LR
XGBst
MLP
NBTS
TNet
PTST
Informer
Chronos
X
526.24
553.24
504.76
614.33
569.67
1324.94
620.81
574.60
591.64
3833.70
502.59
V
581.59
518.54
524.29
731.96
589.24
694.80
852.16
550.80
559.18
4296.38
531.81
U
512.02
564.21
633.66
526.75
525.10
580.70
1842.06
766.06
823.37
581.06
505.46
S
531.41
686.34
702.94
618.88
537.41
618.67
767.10
699.94
812.81
603.85
557.67
Q
518.92
517.77
507.36
614.21
745.04
826.66
927.45
558.95
578.99
932.19
502.46
VU
514.37
570.43
689.56
562.35
515.75
589.72
1.18e26
1134.40
1154.18
612.51
523.73
Appendix
Table 13: Forecasting performance (MAE) across preprocessing pipelines and models. Best per model in bold. Columns in grey are IP-Models. Pipelines are composed of: X (Baseline/None), V (Boxcox transformation), U (First-order differencing), S (Seasonal differencing), and Q (Standard scaling).
Pipeline
ARIMA
ETS
THETA
LR
XGBst
MLP
NBTS
TNet
PTST
Informer
Chronos
X
9.79
9.52
9.29
10.93
10.18
23.49
11.83
10.39
10.68
69.50
9.21
V
12.43
9.21
9.38
10.73
9.88
12.15
10.92
9.71
9.98
81.52
9.43
U
9.49
10.29
12.25
9.73
9.80
11.20
28.96
14.35
15.59
11.46
9.43
S
9.78
13.09
13.58
10.95
9.85
12.18
15.68
13.77
17.16
13.27
10.40
Q
9.54
9.26
9.44
10.92
12.82
14.64
15.93
10.07
10.38
15.95
9.21
VU
9.28
9.84
11.15
9.67
9.32
10.48
28.47
13.19
13.41
10.84
9.22
Appendix
Table 14: Forecasting performance (sMAPE) across preprocessing pipelines and models. Best per model in bold. Columns in grey are IP-Models. Pipelines are composed of: X (Baseline/None), V (Boxcox transformation), U (First-order differencing), S (Seasonal differencing), and Q (Standard scaling).
Figure 7: Relative preference intensity for preprocessing pipelines achieving Top-3 MASE performance across time series subsets. Values are row-normalized per model to highlight the preferred pipelines for each architecture. Darker blue indicates a higher frequency of Top-3 performance.
Model
MASE (No Pipeline)
MASE (Oracle)
Pct Gain (%)
ARIMA
2.35 ± 3.36
1.38 ± 1.49
33.13 ± 26.34
ETS
2.25 ± 3.61
1.44 ± 1.69
27.45 ± 25.93
THETA
2.22 ± 2.11
1.41 ± 1.46
29.17 ± 27.34
LR
2.87 ± 5.96
1.48 ± 1.65
33.98 ± 27.71
XGBoost
2.75 ± 3.77
1.24 ± 1.24
41.48 ± 28.68
MLP
31.88 ± 116.60
1.13 ± 1.04
59.21 ± 31.51
Appendix
Table 15: MASE for each model across "No Pipeline" and "Oracle" configurations, and percentage gain. Positive gain means an increase in performance when using "Oracle". Grey rows indicate IP-Models.
Figure 8: Heatmaps illustrating the distribution of each preprocessing pipeline per Rank across all models.
Model
Oracle Mean (s)
X Mean (s)
Δ Mean (s)
Δ Std (s)
Δ (%)
ARIMA
264.16
157.74
-106.41
531.22
-40.28
ETS
0.196
0.181
-0.015
0.054
-7.75
THETA
0.120
0.120
-0.00003
0.0021
-0.03
LR
0.139
0.139
-0.00009
0.0064
-0.07
XGBoost
3.393
3.115
-0.277
1.161
-8.17
MLP
18.78
19.12
0.338
1.956
1.80
Appendix
Table 16: Training time comparison between Oracle and No-Pipeline configurations for local models. Δ is computed as No Pipeline (X) minus Oracle. Grey rows indicate IP-Models.
Figure 9: Mean MASE vs. mean training time for global models. Marker size indicates peak GPU memory (baselined at 250MB for CPU-only models).
Model
Oracle Mean (s)
X Mean (s)
Δ Mean (s)
Δ Std (s)
Δ (%)
ARIMA
7.56
4.61
-2.94
88.68
-63.83
ETS
0.15
0.17
0.02
0.08
11.95
THETA
0.13
0.13
0.00
0.08
2.56
LR
0.12
0.12
-0.00
0.00
-0.15
XGBoost
0.07
0.06
-0.02
0.01
-28.89
MLP
0.10
0.10
0.01
0.02
5.95
Appendix
Table 17: Training time comparison between Oracle and No-Pipeline configurations for global models. Δ is computed as No Pipeline (X) minus Oracle. Grey rows indicate IP-Models.
Column
Description
series_id
Unique identifier of the time series
model
Forecasting model (e.g., ARIMA, MLP, PatchTST)
pipeline
Encoded preprocessing pipeline (e.g., X, VUQ)
transformation
Variance-stabilizing transformation (V)
diff
First-order differencing (U) indicator
sdiff
Seasonal differencing (S) indicator
Appendix
Table 18: Schema of the Preprocessing-Aware Time Series Metadataset (PATM).
Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models. Existing benchmarks often have limited domain coverage or overlook real-world settings such as tasks with covariates. Their aggregation procedures frequently lack statistical rigor, making it unclear whether observed performance differences reflect true improvements or random variation. Many benchmarks lack consistent evaluation infrastructure or are too rigid for integration into existing pipelines. To address these gaps, we propose fev-bench, a benchmark of 100 forecasting tasks across seven domains, including 46 with covariates. Supporting the benchmark, we introduce fev, a lightweight Python library for forecasting evaluation emphasizing reproducibility and integration with existing workflows. Using fev, fev-bench employs principled aggregation with bootstrapped confidence intervals to report performance along two dimensions: win rates and skill scores. We report results on fev-bench for pretrained, statistical, and baseline models and identify promising future research directions.
Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen +5
Time-series forecasting research has been moving steadily toward larger architectures, from specialized transformers to general-purpose foundation models, on the assumption that capacity is what unlocks accuracy. We take the opposite position: most of the gap can be closed at far lower cost by tuning preprocessing rather than scaling models. We use Ridge regression as the testbed, since it has a closed-form solution and interpretable weights, which let the optimal hyperparameters be read off the search directly. We search over context length, local normalization, regularization, and augmentation on eight standard benchmarks and find three patterns. (1) Optimal lookback is strongly series-specific and often non-monotonic in forecast horizon, with fitted power-law exponents ranging from +0.46 on ETTm2 to −0.19 on Exchange and Traffic, challenging the convention that longer horizons need longer history. (2) Normalizing over a learned trailing fraction of the context, rather than its entirety, is almost universally preferred. (3) Series within the same dataset often disagree on hyperparameters; the optimal degree of cross-series sharing varies from fully shared to fully per-series. The resulting models beat prior linear forecasters on most dataset-horizon entries and exceed Transformer, MLP, and CNN baselines on six of eight benchmarks. The optimized hyperparameters also serve as a diagnostic on the data itself, revealing structures that larger models absorb silently into their learned parameters. We provide an accompanying interactive online demonstration and the code at https://sakanaai.github.io/SearchCast/.
Lang Huang, Jinglue Xu, Luke Darlow
1Sakana AI, Tokyo, Japan · 2National Institute of Informatics, Japan
Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promise of TSFMs, the field lacks a comprehensive and community-accepted model evaluation framework. We see at least four major issues impeding progress on the development of such a framework. First, existing evaluation frameworks comprise benchmark forecasting tasks derived from often outdated datasets (e.g., M3), many of which lack clear metadata and overlap with the corpora used to pre-train TSFMs. Second, these frameworks evaluate models along a narrowly defined set of benchmark forecasting tasks, such as forecast horizon length or domain, but overlook core statistical properties such as non-stationarity and seasonality. Third, domain-specific models (e.g., XGBoost) are often compared unfairly, as existing frameworks do not enforce a systematic and consistent hyperparameter tuning convention for all models. Fourth, visualization tools for interpreting comparative performance are lacking. To address these issues, we introduce TempusBench, an open-source evaluation framework for TSFMs. TempusBench consists of 1) new datasets which are not included in existing TSFM pretraining corpora, 2) a set of novel benchmark tasks that go beyond existing ones, 3) a model evaluation pipeline with a standardized hyperparameter tuning protocol, and 4) a tensorboard-based visualization interface. We provide access to our code on GitHub: https://github.com/Smlcrm/TempusBench and maintain a live leaderboard at https://smlcrm.com/tempusbench.