TempusBench: An Evaluation Framework for Time-Series Forecasting
Organizations: Simulacrum New York City, NY, USA
Abstract
Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promise of TSFMs, the field lacks a comprehensive and community-accepted model evaluation framework. We see at least four major issues impeding progress on the development of such a framework. First, existing evaluation frameworks comprise benchmark forecasting tasks derived from often outdated datasets (e.g., M3), many of which lack clear metadata and overlap with the corpora used to pre-train TSFMs. Second, these frameworks evaluate models along a narrowly defined set of benchmark forecasting tasks, such as forecast horizon length or domain, but overlook core statistical properties such as non-stationarity and seasonality. Third, domain-specific models (e.g., XGBoost) are often compared unfairly, as existing frameworks do not enforce a systematic and consistent hyperparameter tuning convention for all models. Fourth, visualization tools for interpreting comparative performance are lacking. To address these issues, we introduce TempusBench, an open-source evaluation framework for TSFMs. TempusBench consists of 1) new datasets which are not included in existing TSFM pretraining corpora, 2) a set of novel benchmark tasks that go beyond existing ones, 3) a model evaluation pipeline with a standardized hyperparameter tuning protocol, and 4) a tensorboard-based visualization interface. We provide access to our code on GitHub: https://github.com/Smlcrm/TempusBench and maintain a live leaderboard at https://smlcrm.com/tempusbench.
Figures & tables
| Property | Monash [ 18 ] | TFB [ 19 ] | LTSF [ 20 ] | BasicTS+ [ 21 ] | ProbTS [ 22 ] | GIFT-Eval [ 23 ] | TempusBench |
| Frequency Range | Second to Year | Minute to Year | Minute to Week | Minute to Day | Minute to Week | Second to Year | Second to Year |
| Num. Domains | 7 | 6 | 5 | 3 | 5 | 7 | 10 |
| Train/Test data leak | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✗ |
| Variate Types | Uni | Uni/Multi | Multi | Multi | Multi | Uni/Multi | Uni/Multi |
| Prediction Length | Short | Short | Long | Short/Long | Short/Long | Short/Long | Short/Long |
| Stat. Benchmarks | ✗ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ |
| Category | Benchmark Tasks |
| Movement | Stationary, Non-Stationary |
| Data Quality | Noisy data, Data with measurement error |
| Frequency | Seconds, Minutes, Hours, Days, Weeks, Months, Quarterly, Years |
| Context Length | 12, 18, 32, 45, 64, 128, 200, 256, 450, 512, 1024, 1536, 2048 |
| Forecast Horizon | 4, 5, 6, 8, 16, 17, 22, 30, 33, 35, 39, 42, 55, 64 |
| Seasonality | Cyclical, Non-Stationary cyclical, Regressive, Irregular, Additive, Multiplicative |
| Rank | Model | Params (M) | Stochastic | Deterministic | |||||
| CRPS | QS | WIS | MAE | MAPE | MASE | RMSE | |||
| 1 | TiRex (NX-AI) | 300 | 70.4% | 70.4% | 70.4% | 75.1% | 70.5% | 75.1% | 73.4% |
| 2 | TiRex 1.1 (NX-AI) | 300 | 68.5% | 68.5% | 68.5% | 74.3% | 69.7% | 74.3% | 71.5% |
| 3 | TimesFM 2.5 (Google) | 200 | 66.4% | 66.4% | 66.4% | 73.3% | 69.4% | 73.4% | 71.6% |
| 4 | Moirai 2.0 (Salesforce) | 12 | 64.0% | 64.0% | 64.0% | 69.6% | 67.1% | 69.6% | 68.8% |
| 5 | Chronos-2 (Amazon) | 120 | 64.6% | 64.6% | 64.6% | 66.8% | 63.0% | 66.8% | 68.5% |
| Rank | Model | Params (M) | Stochastic | Deterministic | |||||
| CRPS | QS | WIS | MAE | MAPE | MASE | RMSE | |||
| 1 | TiRex (NX-AI) | 300 | 78.1% | 78.1% | 78.1% | 86.2% | 80.6% | 86.2% | 85.9% |
| 2 | TiRex 1.1 (NX-AI) | 300 | 78.3% | 78.3% | 78.3% | 85.3% | 78.9% | 85.3% | 84.8% |
| 3 | TimesFM 2.5 (Google) | 200 | 73.3% | 73.3% | 73.3% | 83.6% | 76.7% | 83.5% | 82.5% |
| 4 | Chronos-2 (Amazon) | 120 | 77.8% | 77.8% | 77.8% | 81.3% | 64.5% | 81.1% | 83.6% |
| 5 | Chronos-2-Small (Amazon) | 40 | 74.8% | 74.8% | 74.8% | 75.9% | 66.7% | 75.7% | 81.0% |
| Rank | Model | Params (M) | Stochastic | Deterministic | |||||
| CRPS | QS | WIS | MAE | MAPE | MASE | RMSE | |||
| 1 | Chronos-2 (Amazon) | 120 | 83.0% | 83.0% | 83.0% | 79.1% | 72.1% | 79.3% | 80.7% |
| 2 | TabPFN-TS (Prior Labs) | 11 | 77.0% | 77.0% | 77.0% | 82.0% | 80.4% | 82.2% | 72.3% |
| 3 | Chronos-2-Small (Amazon) | 40 | 80.3% | 80.3% | 80.3% | 74.1% | 62.3% | 74.2% | 79.7% |
| 4 | TimesFM 200M (Google) | 200 | — | — | — | 72.2% | 71.0% | 72.1% | 65.4% |
| 5 | Chronos-Bolt-Base (Amazon) | 205 | 61.1% | 61.1% | 61.1% | 67.2% | 69.3% | 67.2% | 67.9% |
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
| Rank | Model | MAE | MAPE | MASE | RMSE |
|---|---|---|---|---|---|
| 1 | TiRex (NX-AI) | 0.639 | 0.660 | 0.644 | 0.596 |
| 2 | TiRex 1.1 (NX-AI) | 0.625 | 0.635 | 0.631 | 0.579 |
| 3 | TimesFM 2.5 (Google) | 0.622 | 0.640 | 0.628 | 0.584 |
| 4 | Moirai 2.0 (Salesforce) | 0.621 | 0.639 | 0.627 | 0.578 |
| 5 | Chronos-2 (Amazon) | 0.637 | 0.644 | 0.642 | 0.608 |
| 6 | TimesFM 500M (Google) | 0.525 | 0.563 | 0.532 | 0.482 |
| Rank | Model | MAE | MAPE | MASE | RMSE |
|---|---|---|---|---|---|
| 1 | TiRex (NX-AI) | 0.629 | 0.462 | 0.639 | 0.617 |
| 2 | TiRex 1.1 (NX-AI) | 0.624 | 0.494 | 0.635 | 0.605 |
| 3 | TimesFM 2.5 (Google) | 0.653 | 0.681 | 0.663 | 0.644 |
| 4 | Chronos-2 (Amazon) | 0.524 | 0.154 | 0.537 | 0.523 |
| 5 | Chronos-2-Small (Amazon) | 0.573 | 0.417 | 0.585 | 0.576 |
| 6 | TimesFM 500M (Google) | 0.577 | 0.264 | 0.589 | 0.563 |
| Rank | Model | MAE | MAPE | MwASE | RMSE |
|---|---|---|---|---|---|
| 1 | Chronos-2 (Amazon) | 0.630 | 0.688 | 0.638 | 0.628 |
| 2 | TabPFN-TS (Prior Labs) | 0.757 | 0.824 | 0.762 | 0.742 |
| 3 | Chronos-2-Small (Amazon) | 0.589 | 0.607 | 0.598 | 0.599 |
| 4 | TimesFM 200M (Google) | 0.716 | 0.775 | 0.722 | 0.704 |
| 5 | Chronos-Bolt-Base (Amazon) | 0.385 | 0.548 | 0.399 | 0.384 |
| 6 | Chronos-Bolt-Tiny (Amazon) | 0.417 | 0.575 | 0.430 | 0.411 |
| Category | Included Models | Core Characteristics |
| Foundation Models | Moirai, Moirai-MoE, TimesFM, TimesFM-2.0, Chronos, Lag-Llama, Toto, MOMENT | Paradigm: Universal, zero-shot/few-shot forecasting. A single large model is pre-trained on massive, diverse datasets and generalizes to new tasks without redataset-specific training. Architecture: Primarily based on Transformers or other deep learning structures like MLP-Mixers. They process raw time series data via patching or novel tokenization schemes. I/O: Natively handle univariate, multivariate, and covariate data. Often produce probabilistic forecasts. |
| Classic Machine Learning | LSTM, Random Forest (RF), XGBoost, SVR, TabPFN-TS, Tiny Time Mixers (TTM) | Paradigm: Supervised learning models trained per-dataset. They excel at capturing complex, non-linear relationships but require specific training for each task. Architecture: Diverse. E.g., RNNs (for sequence memory), Tree Ensembles (for interaction effects), and Kernel Methods. I/O: Typically require explicit feature (e.g., lags, calendar variables) engineering to create a tabular format. Most often produce point forecasts. |
| Statistical & Decomposable | ARIMA, Holt-Winters, Prophet, Theta Method, Croston’s Method, Seasonal Naive | Paradigm: Assume the time series is generated by an underlying statistical process or can be decomposed into simpler, interpretable components like trend and seasonality. Architecture: An explicit mathematical formula is fit directly to an individual time series. I/O: Highly interpretable point forecasts. Often specialized for particular data patterns (e.g., intermittency with Croston’s ). |
| Category | Task | Tot. pts. | ||||
|---|---|---|---|---|---|---|
| Trend | ||||||
| Non-stationary | Software job postings [ 41 ] | 512 | 64 | 1 | 0 | 1,827 |
| Decomposition | ||||||
| Additive | Synthetic additive ( Section D.1 ) | 1024 | 64 | 1 | 0 | 3,000 |
| Multiplicative | Synthetic multiplicative ( Section D.2 ) | 1024 | 64 | 1 | 0 | 3,000 |
| Frequency | ||||||
| Category | Task | Tot. pts. | ||||
|---|---|---|---|---|---|---|
| Frequency | ||||||
| Days | Gold India [ 58 ] | 1024 | 64 | 5 | 0 | 20,120 |
| Hours | Madrid transport pollution [ 59 ] | 2048 | 64 | 14 | 0 | 2,544,542 |
| Minutes | U.S. stocks2003–24 [ 60 ] | 2048 | 64 | 6 | 0 | 732,660 |
| Minutes | U.S. stocks 2003–24 (longest) [ 60 ] | 2048 | 64 | 6 | 0 | 732,660 |
| Months | Airlines baggage [ 61 ] | 32 | 8 | 4 | 0 | 336 |
| Category | Task | Tot. pts. | ||||
|---|---|---|---|---|---|---|
| Trend | ||||||
| Non-stationary | NIFTY-50 minute covariates (non-stationary) [ 65 ] | 2048 | 64 | 1 | 4 | 2,001,355 |
| Noisy data | U.S. macro covariates (high-variation panel) [ 66 ] | 18 | 5 | 1 | 6 | 336 |
| Frequency | ||||||
| Days | U.S. equities (daily covariates) [ 67 ] | 2048 | 64 | 1 | 4 | 15,095 |
| Hours | SoCal energy (hourly covariates) [ 68 ] | 2048 | 64 | 1 | 34 | 1,840,475 |
| Era | Primary Data Challenge(s) | Dominant Model Paradigm | Key Models | Inherent Limitations |
| Statistical | Trends, Seasonality, Stationarity | Time-Domain Statistical Models | ARIMA, Holt-Winters, Theta | Struggle with non-linearity and complex dependencies |
| Machine Learning | Non-Linearity, Complex Interactions | Non-parametric | Ensemble Models (SVR, Random Forest, XGBoost) | Limited handling of long-range temporal dependencies |
| Deep Learning | Long-Range Dependencies, Sequential Patterns | Recurrent, Attention-based Networks | LSTMs, Transformers | Data-hungry, computationally intensive, task-specific |
| Foundation Models | Data Heterogeneity, Scale, Task Generalization | Large Pre-trained Models | MOMENT, TOTO, Chronos | Reliance on massive, curated datasets; evaluation bottleneck |