Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promise of TSFMs, the field lacks a comprehensive and community-accepted model evaluation framework. We see at least four major issues impeding progress on the development of such a framework. First, existing evaluation frameworks comprise benchmark forecasting tasks derived from often outdated datasets (e.g., M3), many of which lack clear metadata and overlap with the corpora used to pre-train TSFMs. Second, these frameworks evaluate models along a narrowly defined set of benchmark forecasting tasks, such as forecast horizon length or domain, but overlook core statistical properties such as non-stationarity and seasonality. Third, domain-specific models (e.g., XGBoost) are often compared unfairly, as existing frameworks do not enforce a systematic and consistent hyperparameter tuning convention for all models. Fourth, visualization tools for interpreting comparative performance are lacking. To address these issues, we introduce TempusBench, an open-source evaluation framework for TSFMs. TempusBench consists of 1) new datasets which are not included in existing TSFM pretraining corpora, 2) a set of novel benchmark tasks that go beyond existing ones, 3) a model evaluation pipeline with a standardized hyperparameter tuning protocol, and 4) a tensorboard-based visualization interface. We provide access to our code on GitHub: https://github.com/Smlcrm/TempusBench and maintain a live leaderboard at https://smlcrm.com/tempusbench.
Figures & tables
Property
Monash [ 18 ]
TFB [ 19 ]
LTSF [ 20 ]
BasicTS+ [ 21 ]
ProbTS [ 22 ]
GIFT-Eval [ 23 ]
TempusBench
Frequency Range
Second to Year
Minute to Year
Minute to Week
Minute to Day
Minute to Week
Second to Year
Second to Year
Num. Domains
7
6
5
3
5
7
10
Train/Test data leak
✓
✓
✓
✓
✓
✓
✗
Variate Types
Uni
Uni/Multi
Multi
Multi
Multi
Uni/Multi
Uni/Multi
Prediction Length
Short
Short
Long
Short/Long
Short/Long
Short/Long
Short/Long
Stat. Benchmarks
✗
✗
✗
✗
✗
✗
✓
Table 1: Property comparisons of various forecasting benchmarks.
Category
Benchmark Tasks
Movement
Stationary, Non-Stationary
Data Quality
Noisy data, Data with measurement error
Frequency
Seconds, Minutes, Hours, Days, Weeks, Months, Quarterly, Years
Table 2: Taxonomy of all univariate, multivariate, and covariate tasks included in TempusBench.
Rank
Model
Params (M)
Stochastic
Deterministic
CRPS
QS
WIS
MAE
MAPE
MASE
RMSE
1
TiRex (NX-AI)
300
70.4%
70.4%
70.4%
75.1%
70.5%
75.1%
73.4%
2
TiRex 1.1 (NX-AI)
300
68.5%
68.5%
68.5%
74.3%
69.7%
74.3%
71.5%
3
TimesFM 2.5 (Google)
200
66.4%
66.4%
66.4%
73.3%
69.4%
73.4%
71.6%
4
Moirai 2.0 (Salesforce)
12
64.0%
64.0%
64.0%
69.6%
67.1%
69.6%
68.8%
5
Chronos-2 (Amazon)
120
64.6%
64.6%
64.6%
66.8%
63.0%
66.8%
68.5%
Table 3 : Univariate tasks – Average win rates by metric for deterministic & probabilistic models
Rank
Model
Params (M)
Stochastic
Deterministic
CRPS
QS
WIS
MAE
MAPE
MASE
RMSE
1
TiRex (NX-AI)
300
78.1%
78.1%
78.1%
86.2%
80.6%
86.2%
85.9%
2
TiRex 1.1 (NX-AI)
300
78.3%
78.3%
78.3%
85.3%
78.9%
85.3%
84.8%
3
TimesFM 2.5 (Google)
200
73.3%
73.3%
73.3%
83.6%
76.7%
83.5%
82.5%
4
Chronos-2 (Amazon)
120
77.8%
77.8%
77.8%
81.3%
64.5%
81.1%
83.6%
5
Chronos-2-Small (Amazon)
40
74.8%
74.8%
74.8%
75.9%
66.7%
75.7%
81.0%
Table 4 : Multivariate tasks – Average win rates by metric for deterministic & probabilistic models
Rank
Model
Params (M)
Stochastic
Deterministic
CRPS
QS
WIS
MAE
MAPE
MASE
RMSE
1
Chronos-2 (Amazon)
120
83.0%
83.0%
83.0%
79.1%
72.1%
79.3%
80.7%
2
TabPFN-TS (Prior Labs)
11
77.0%
77.0%
77.0%
82.0%
80.4%
82.2%
72.3%
3
Chronos-2-Small (Amazon)
40
80.3%
80.3%
80.3%
74.1%
62.3%
74.2%
79.7%
4
TimesFM 200M (Google)
200
—
—
—
72.2%
71.0%
72.1%
65.4%
5
Chronos-Bolt-Base (Amazon)
205
61.1%
61.1%
61.1%
67.2%
69.3%
67.2%
67.9%
Table 5 : Covariate tasks – Average win rates by metric for deterministic & probabilistic models
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Rank
Model
MAE
MAPE
MASE
RMSE
1
TiRex (NX-AI)
0.639
0.660
0.644
0.596
2
TiRex 1.1 (NX-AI)
0.625
0.635
0.631
0.579
3
TimesFM 2.5 (Google)
0.622
0.640
0.628
0.584
4
Moirai 2.0 (Salesforce)
0.621
0.639
0.627
0.578
5
Chronos-2 (Amazon)
0.637
0.644
0.642
0.608
6
TimesFM 500M (Google)
0.525
0.563
0.532
0.482
Appendix
Table 6: Per-metric averaged skill scores across univariate tasks
Rank
Model
MAE
MAPE
MASE
RMSE
1
TiRex (NX-AI)
0.629
0.462
0.639
0.617
2
TiRex 1.1 (NX-AI)
0.624
0.494
0.635
0.605
3
TimesFM 2.5 (Google)
0.653
0.681
0.663
0.644
4
Chronos-2 (Amazon)
0.524
0.154
0.537
0.523
5
Chronos-2-Small (Amazon)
0.573
0.417
0.585
0.576
6
TimesFM 500M (Google)
0.577
0.264
0.589
0.563
Appendix
Table 7: Per-metric averaged skill scores across multivariate tasks
Rank
Model
MAE
MAPE
MwASE
RMSE
1
Chronos-2 (Amazon)
0.630
0.688
0.638
0.628
2
TabPFN-TS (Prior Labs)
0.757
0.824
0.762
0.742
3
Chronos-2-Small (Amazon)
0.589
0.607
0.598
0.599
4
TimesFM 200M (Google)
0.716
0.775
0.722
0.704
5
Chronos-Bolt-Base (Amazon)
0.385
0.548
0.399
0.384
6
Chronos-Bolt-Tiny (Amazon)
0.417
0.575
0.430
0.411
Appendix
Table 8: Per-metric averaged skill scores across covariate tasks
Category
Included Models
Core Characteristics
Foundation Models
Moirai, Moirai-MoE, TimesFM, TimesFM-2.0, Chronos, Lag-Llama, Toto, MOMENT
Paradigm: Universal, zero-shot/few-shot forecasting. A single large model is pre-trained on massive, diverse datasets and generalizes to new tasks without redataset-specific training. Architecture: Primarily based on Transformers or other deep learning structures like MLP-Mixers. They process raw time series data via patching or novel tokenization schemes. I/O: Natively handle univariate, multivariate, and covariate data. Often produce probabilistic forecasts.
Classic Machine Learning
LSTM, Random Forest (RF), XGBoost, SVR, TabPFN-TS, Tiny Time Mixers (TTM)
Paradigm: Supervised learning models trained per-dataset. They excel at capturing complex, non-linear relationships but require specific training for each task. Architecture: Diverse. E.g., RNNs (for sequence memory), Tree Ensembles (for interaction effects), and Kernel Methods. I/O: Typically require explicit feature (e.g., lags, calendar variables) engineering to create a tabular format. Most often produce point forecasts.
Paradigm: Assume the time series is generated by an underlying statistical process or can be decomposed into simpler, interpretable components like trend and seasonality. Architecture: An explicit mathematical formula is fit directly to an individual time series. I/O: Highly interpretable point forecasts. Often specialized for particular data patterns (e.g., intermittency with Croston’s ).
Appendix
Table 9 : Summary of forecasters included in TempusBench.
Category
Task
l
h
n
m
Tot. pts.
Trend
Non-stationary
Software job postings [ 41 ]
512
64
1
0
1,827
Decomposition
Additive
Synthetic additive ( Section D.1 )
1024
64
1
0
3,000
Multiplicative
Synthetic multiplicative ( Section D.2 )
1024
64
1
0
3,000
Frequency
Appendix
Table 11 : Summary of univariate benchmark tasks
Category
Task
l
h
n
m
Tot. pts.
Frequency
Days
Gold India [ 58 ]
1024
64
5
0
20,120
Hours
Madrid transport pollution [ 59 ]
2048
64
14
0
2,544,542
Minutes
U.S. stocks2003–24 [ 60 ]
2048
64
6
0
732,660
Minutes
U.S. stocks 2003–24 (longest) [ 60 ]
2048
64
6
0
732,660
Months
Airlines baggage [ 61 ]
32
8
4
0
336
Appendix
Table 12 : Summary of multivariate benchmark tasks
Benchmark quality is critical for meaningful evaluation and sustained progress in time series forecasting, particularly with the rise of pretrained models. Existing benchmarks often have limited domain coverage or overlook real-world settings such as tasks with covariates. Their aggregation procedures frequently lack statistical rigor, making it unclear whether observed performance differences reflect true improvements or random variation. Many benchmarks lack consistent evaluation infrastructure or are too rigid for integration into existing pipelines. To address these gaps, we propose fev-bench, a benchmark of 100 forecasting tasks across seven domains, including 46 with covariates. Supporting the benchmark, we introduce fev, a lightweight Python library for forecasting evaluation emphasizing reproducibility and integration with existing workflows. Using fev, fev-bench employs principled aggregation with bootstrapped confidence intervals to report performance along two dimensions: win rates and skill scores. We report results on fev-bench for pretrained, statistical, and baseline models and identify promising future research directions.
Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen +5
High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to reveal complex temporal patterns and align poorly with human intuitive preferences. While the ''LLM-as-a-Judge'' paradigm has revolutionized text evaluation by providing flexible, human-aligned judgment, its application to time series remains largely unexplored. In this paper, we leverage Vision-Language Models (VLMs) as judges for time series forecasting, harnessing their ability to comprehend time series plots grounded in textual information. Specifically, we propose a novel framework integrating micro- and macro-level judgments informed by contextual information to evaluate time series forecasting. To this end, we introduce TimeVista, a comprehensive VLM-as-a-Judge benchmark comprising 5563 time series samples paired with detailed evaluation rubrics. Extensive meta-evaluations demonstrate that VLMs are highly reliable judges, achieving significantly higher consistency with human preferences than conventional metrics. Building upon our benchmark, we comprehensively assess recent Time Series Foundation Models (TSFMs) under the VLM-as-a-Judge paradigm. Our results demonstrate that VLMs serve as robust and interpretable judges, providing a comprehensive, human-aligned standard for evaluating time series models.
Zhi Chen, Yuxuan Wang, Jialong Wu +5
School of Software, BNRist, Tsinghua University, Beijing 100084, China
Despite the success of Time Series Foundation Models (TSFMs) on broad benchmarks, their ability to internalize basic temporal logic, especially in settings supported by exogenous covariates, remains under-examined. We introduce SimpleTimeBench, a diagnostic univariate and multivariate "unit test" suite for primitives such as monotonic trends, periodic signals and leading indicator covariates, scenarios where near-perfect forecasts should be trivial. Surprisingly, prominent multivariate TSFMs (Chronos-2, Moirai and Toto) frequently produce suboptimal zero-shot forecasts for these inputs. While fine-tuning Chronos-2 improves its behaviour on specific tasks, we show that this adaptation degrades performance on other fundamental patterns rather than enhancing its generalizable foundational capabilities. This reveals a gap between pre-training scale and basic temporal reasoning, suggesting that current TSFMs could potentially lack the inductive biases needed to capture simple predictable functions. We further demonstrate that these failures are not merely synthetic curiosities: they persist in real-world sensor forecasting, where TSFMs consistently underutilize leading indicators available in observed covariates. This inability to capture simple relationships limits the practical utility and reliability of current multivariate models.