cs.LGSep 30, 2026

When, Not How Much: Evaluating Time-Series Foundation Models on Sparse Events

Authors: Daniel Schoess, Florian von Wangenheim

Organizations: ETH Zurich

Abstract

Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on which future periods contain activity. Standard benchmarks do not assess this. On five sparse datasets, we rank positions within forecast windows that contain both events and zeros. The released point forecasts of 12 TSFMs improve chance-corrected average precision over training-free references by at most 0.031, and in chance-corrected AUC the median TSFM falls below them on every dataset. With event supervision, linear probes of six frozen backbones improve on their backbone's point forecast in 29 of 30 backbone--dataset pairs. Averaging the predicted quantiles instead of taking their median improves the ranking of most TSFMs that forecast the median, and on two datasets the strongest such outputs rival the probes. The probes' advantage over raw-context learners depends on the dataset, and under the same probe, pretrained features outperform randomly initialized ones for five of six backbones. For sparse-event ranking, released point forecasts thus add little over simple references, whereas lightweight event heads on frozen TSFMs rank events better than these forecasts, and the best of them exceed gradient-boosted trees trained on the raw context on three of the five datasets. More broadly, assessing pretrained forecasters on tasks beyond value forecasting requires reporting their outputs, supervised probes of their representations, and raw-context and randomized controls side by side, since each supports a different conclusion.

Figures & tables

Appendix figures & tables30 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. TempusBench: An Evaluation Framework for Time-Series Forecasting

    Apr 13, 2026Denizalp Goktas, Gerardo Riaño-Briceño, Omkar Tekawade +14Time Series Foundation Models

  2. Evaluating Accuracy and Probabilistic Reliability of Zero-Shot Time Series Foundation Models

    Sep 22, 2026Panagiotis Michael, Moysis Symeonides, Demetris TrihinasZero-Shot ForecastingCnn-Lstm