cs.LGApr 13, 2026

TempusBench: An Evaluation Framework for Time-Series Forecasting

Authors: Denizalp Goktas, Gerardo Riaño-Briceño, Omkar Tekawade, Alif Abdullah, Md Yasif Jamal, Vinesha Shaik, Arsh Singh, Aryan Nair, +9 more

Organizations: Simulacrum New York City, NY, USA

Abstract

Foundation models have transformed natural language processing and computer vision, and a rapidly growing literature on time-series foundation models (TSFMs) seeks to replicate this success in forecasting. While recent open-source models demonstrate the promise of TSFMs, the field lacks a comprehensive and community-accepted model evaluation framework. We see at least four major issues impeding progress on the development of such a framework. First, existing evaluation frameworks comprise benchmark forecasting tasks derived from often outdated datasets (e.g., M3), many of which lack clear metadata and overlap with the corpora used to pre-train TSFMs. Second, these frameworks evaluate models along a narrowly defined set of benchmark forecasting tasks, such as forecast horizon length or domain, but overlook core statistical properties such as non-stationarity and seasonality. Third, domain-specific models (e.g., XGBoost) are often compared unfairly, as existing frameworks do not enforce a systematic and consistent hyperparameter tuning convention for all models. Fourth, visualization tools for interpreting comparative performance are lacking. To address these issues, we introduce TempusBench, an open-source evaluation framework for TSFMs. TempusBench consists of 1) new datasets which are not included in existing TSFM pretraining corpora, 2) a set of novel benchmark tasks that go beyond existing ones, 3) a model evaluation pipeline with a standardized hyperparameter tuning protocol, and 4) a tensorboard-based visualization interface. We provide access to our code on GitHub: https://github.com/Smlcrm/TempusBench and maintain a live leaderboard at https://smlcrm.com/tempusbench.

Figures & tables

Appendix figures & tables8 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. fev-bench: A Realistic Benchmark for Time Series Forecasting

    Sep 30, 2025Oleksandr Shchur, Abdul Fatir Ansari, Caner Turkmen +5Time Series ForecastingDense Baseline

  2. TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting

    Jun 15, 2026Zhi Chen, Yuxuan Wang, Jialong Wu +5Time Series ForecastingTime Series