When, Not How Much: Evaluating Time-Series Foundation Models on Sparse Events
Organizations: ETH Zurich
Abstract
Pretrained time-series foundation models (TSFMs) are evaluated as forecasters of future values, yet for sparse series many decisions depend only on which future periods contain activity. Standard benchmarks do not assess this. On five sparse datasets, we rank positions within forecast windows that contain both events and zeros. The released point forecasts of 12 TSFMs improve chance-corrected average precision over training-free references by at most 0.031, and in chance-corrected AUC the median TSFM falls below them on every dataset. With event supervision, linear probes of six frozen backbones improve on their backbone's point forecast in 29 of 30 backbone--dataset pairs. Averaging the predicted quantiles instead of taking their median improves the ranking of most TSFMs that forecast the median, and on two datasets the strongest such outputs rival the probes. The probes' advantage over raw-context learners depends on the dataset, and under the same probe, pretrained features outperform randomly initialized ones for five of six backbones. For sparse-event ranking, released point forecasts thus add little over simple references, whereas lightweight event heads on frozen TSFMs rank events better than these forecasts, and the best of them exceed gradient-boosted trees trained on the raw context on three of the five datasets. More broadly, assessing pretrained forecasters on tasks beyond value forecasting requires reporting their outputs, supervised probes of their representations, and raw-context and randomized controls side by side, since each supports a different conclusion.
Figures & tables
| Question | Compared scores | Event labels | Models |
| Q1 Output | point forecast vs. training-free reference | neither | 12 TSFMs |
| Q2 Probe | linear probe vs. point forecast | probe | 6 backbones |
| linear probe vs. raw-context logistic regression | both | 6 backbones | |
| linear probe vs. raw-context LightGBM | both | 6 backbones | |
| Q3 Pretraining | pretrained vs. randomized linear probe | both | 6 backbones |
| Dataset | Cadence | Context | Split design | Windows | Series | Zeros (%) |
| Amazon | daily | 364 | 5 series-disjoint | 22,600 | 1,135 | 87.5 |
| Favorita | daily | 364 | 5 series-disjoint | 321,747 | 26,427 | 45.7 |
| M5 | daily | 364 | 5 series-disjoint | 120,063 | 5,979 | 64.1 |
| NYC taxi | hourly | 336 | 1 calendar | 19,207 | 164 | 73.7 |
| Web traffic | daily | 364 | 5 series-disjoint | 41,816 | 1,597 | 65.6 |
| AP skill | AUC skill | |||||||
| Dataset | Ref. | Median | Max | Best probe | Ref. | Median | Max | Best probe |
| Amazon | 0.008 | 0.002 | 0.002 | 0.006 | 0.024 | 0.017 | 0.006 | 0.009 |
| Favorita | 0.103 | 0.007 | 0.023 | 0.092 | 0.114 | 0.012 | 0.005 | 0.080 |
| M5 | 0.043 | 0.005 | 0.013 | 0.033 | 0.062 | 0.022 | 0.001 | 0.040 |
| NYC taxi | 0.185 | 0.013 | 0.031 | 0.049 | 0.310 | 0.053 | 0.020 | 0.042 |
| Web traffic | 0.015 | 0.003 | 0.006 | 0.057 | 0.016 | 0.003 | 0.004 | 0.071 |
| Contrast | Amazon | Favorita | M5 | NYC taxi | Web traffic |
| Probe point forecast | 0.006 (5/1/0) | 0.082 (6/0/0) | 0.032 (6/0/0) | 0.015 (6/0/0) | 0.047 (6/0/0) |
| Probe raw logistic | 0.002 (3/3/0) | 0.008 (4/0/2) | 0.002 (3/0/3) | 0.087 (6/0/0) | 0.033 (6/0/0) |
| Probe raw LightGBM | 0.000 (2/3/1) | 0.020 (0/0/6) | 0.002 (3/0/3) | 0.001 (2/3/1) | 0.115 (0/0/6) |
| Pretrained randomized | 0.005 (5/1/0) | 0.070 (5/0/1) | 0.031 (5/0/1) | 0.094 (6/0/0) | 0.021 (5/0/1) |
Appendix figures & tables30 assets
Supplementary material from the paper’s appendix.
Appendix
| Dataset | Source | Series | Value per position | Period | Source series | Retained |
| Amazon | ( Berke et al., 2024 ) | respondent | daily purchase amount | 2018-01-01 to 2024-08-15 | 5,027 | 5,024 |
| Favorita | ( Corporación Favorita et al., 2017 ) | store–item pair | daily unit sales | 2013-01-01 to 2017-08-15 | 174,685 | 121,619 |
| M5 | ( Makridakis et al., 2022 ) | item–store pair | daily unit sales | 2011-01-29 to 2016-05-22 | 30,490 | 26,334 |
| NYC taxi | ( NYC Taxi and Limousine Commission, 2024 ) | pickup location | hourly pickups | 2022-01-01 to 2024-12-31 | 262/263/263 | 185/188/168 |
| Web traffic | ( Godahewa et al., 2021 ; Godahewa et al., 2022b ) | page | daily page views | 2015-07-01 to 2022-06-30 | 145,063 | 7,048 |
| Windows | Series | Target zeros (%) | ||||||||
| Dataset | Cadence | Context | Splits | Candidate | Context-eligible | Mixed-label | Candidate | Mixed-label | Candidate | Mixed-label |
| Amazon | daily | 364 | 5 | 30,240 | 26,914 | 22,600 | 1,154 | 1,135 | 90.6 | 87.5 |
| Favorita | daily | 364 | 5 | 608,100 | 375,861 | 321,747 | 27,500 | 26,427 | 67.9 | 45.7 |
| M5 | daily | 364 | 5 | 158,040 | 130,032 | 120,063 | 5,979 | 5,979 | 72.0 | 64.1 |
| NYC taxi | hourly | 336 | 1 | 22,008 | 20,440 | 19,207 | 168 | 164 | 77.0 | 73.7 |
| Web traffic | daily | 364 | 5 | 60,010 | 50,612 | 41,816 | 1,615 | 1,597 | 70.1 | 65.6 |
| Benchmark | Zero threshold | Windows | Valid points | Absolute target mass |
| GIFT-Eval | 40% | 24.19 | 12.30 | 6.74 |
| GIFT-Eval | 80% | 12.02 | 3.96 | 0.28 |
| GIFT-Eval | 90% | 8.69 | 3.27 | 0.11 |
| BOOM | 40% | 1.55 | 0.87 | 0.30 |
| BOOM | 80% | 1.55 | 0.86 | 0.30 |
| BOOM | 90% | 1.55 | 0.86 | 0.30 |
| TSFM | Amazon | Favorita | M5 | NYC taxi | Web traffic |
| Chronos-2 | E | E | E | O | O |
| Chronos-Bolt | U | U | E | U | U |
| Moirai-1.1 | U | U | U | U | U |
| Moirai-2.0 | E | O | O | O | O |
| Moirai-MoE | E | O | O | O | O |
| Sundial | U | O | E | O | O |
| AP skill | AUC skill | |||||
| Contrast | E (12) | O (13) | All (30) | E (12) | O (13) | All (30) |
| Pretrained randomized | 0.019 (11/1/0) | 0.035 (10/0/3) | 0.027 (26/1/3) | 0.028 (10/1/1) | 0.041 (10/0/3) | 0.031 (25/1/4) |
| Probe point forecast | 0.034 (11/1/0) | 0.046 (13/0/0) | 0.034 (29/1/0) | 0.052 (11/1/0) | 0.063 (13/0/0) | 0.057 (29/1/0) |
| Probe raw logistic | 0.005 (8/2/2) | 0.035 (10/0/3) | 0.009 (22/3/5) | 0.008 (9/2/1) | 0.039 (10/0/3) | 0.015 (22/4/4) |
| Labels | Probability | AP | AUROC |
| 1100 | |||
| 0110 | |||
| 0011 | |||
| 1001 |
| TSFM | Checkpoint | Ranked point forecast | Generation interface | Forward passes |
| Chronos-2 | amazon/chronos-2 | median | direct | 1 |
| Chronos-Bolt | amazon/chronos-bolt-base | median | direct | 1 |
| Moirai-1.1 | Salesforce/moirai-1.1-R-large | mean of 20 samples | sampling | 1 |
| Moirai-2.0 | Salesforce/moirai-2.0-R-small | median | autoregressive | 1 |
| Moirai-MoE | Salesforce/moirai-moe-1.0-R-base | mean of 20 samples | autoregressive sampling | 4 |
| Sundial | thuml/sundial-base-128m | mean of 20 samples | sampling | 1 |
| Backbone | Captured tensor | Dimension | Anchors | Positions/anchor |
| Chronos-Bolt | output-patch-head input | 768 | 1 | 64 |
| Chronos-2 | output-patch query token | 768 | 1 | 64 |
| TimeMoE | last-token LM-head input | 768 | 1 | 64 |
| Timer | last-token LM-head input | 1,024 | 1 | 64 |
| Sundial | flow-head conditioning state | 768 | 1 | 64 |
| Toto-2.0 | output-head input | 2,048 | 2 | 32 |
| Procedure | Input | Standardization | Training sample | Validation selection | Output organization |
| Raw logistic | 364/336-step context | train mean/std | 100k cap; 20 strata | L2 | one linear map |
| Representation linear | frozen representation | train mean/std | 100k cap; 20 strata | L2 | one map |
| Raw LightGBM | raw context | none | 100k cap; 20 strata | per-position early stop | 64 classifiers |
| Representation GBM | frozen representation | none | 100k cap; 20 strata | per-position early stop | 64 classifiers |
| AP skill | AUC skill | |||||
| Skill | reference | Skill | reference | |||
| Training-free reference | 0.0078 | – | 0.0238 | – | ||
| Chronos-2 | 0.0057 | 0.0021 | [ 0.0042, 0.0003] | 0.0100 | 0.0137 | [ 0.0203, 0.0072] |
| Chronos-Bolt | 0.0055 | 0.0023 | [ 0.0041, 0.0004] | 0.0116 | 0.0122 | [ 0.0192, 0.0050] |
| Moirai-1.1 | 0.0004 | 0.0074 | [ 0.0094, 0.0054] | 0.0009 | 0.0247 | [ 0.0310, 0.0181] |
| Moirai-2.0 | 0.0050 | 0.0028 | [ 0.0045, 0.0011] | 0.0050 | 0.0188 | [ 0.0255, 0.0120] |
| AP skill | AUC skill | |||||
| Skill | reference | Skill | reference | |||
| Training-free reference | 0.1026 | – | 0.1136 | – | ||
| Chronos-2 | 0.1160 | 0.0133 | [0.0123, 0.0144] | 0.1106 | 0.0030 | [ 0.0043, 0.0018] |
| Chronos-Bolt | 0.0859 | 0.0167 | [ 0.0180, 0.0155] | 0.0816 | 0.0320 | [ 0.0334, 0.0306] |
| Moirai-1.1 | 0.0482 | 0.0544 | [ 0.0558, 0.0530] | 0.0478 | 0.0658 | [ 0.0673, 0.0643] |
| Moirai-2.0 | 0.1231 | 0.0204 | [0.0195, 0.0215] | 0.1163 | 0.0027 | [0.0014, 0.0039] |
| AP skill | AUC skill | |||||
| Skill | reference | Skill | reference | |||
| Training-free reference | 0.0434 | – | 0.0619 | – | ||
| Chronos-2 | 0.0453 | 0.0019 | [0.0006, 0.0032] | 0.0518 | 0.0101 | [ 0.0121, 0.0080] |
| Chronos-Bolt | 0.0329 | 0.0105 | [ 0.0119, 0.0092] | 0.0290 | 0.0329 | [ 0.0350, 0.0307] |
| Moirai-1.1 | 0.0173 | 0.0262 | [ 0.0277, 0.0247] | 0.0237 | 0.0382 | [ 0.0404, 0.0361] |
| Moirai-2.0 | 0.0561 | 0.0127 | [0.0114, 0.0140] | 0.0626 | 0.0007 | [ 0.0015, 0.0028] |
| AP skill | AUC skill | |||||
| Skill | reference | Skill | reference | |||
| Training-free reference | 0.1853 | – | 0.3095 | – | ||
| Chronos-2 | 0.2117 | 0.0265 | [0.0171, 0.0363] | 0.2572 | 0.0523 | [ 0.0672, 0.0366] |
| Chronos-Bolt | 0.2164 | 0.0311 | [0.0226, 0.0395] | 0.2710 | 0.0385 | [ 0.0541, 0.0224] |
| Moirai-1.1 | 0.0035 | 0.1817 | [ 0.2058, 0.1607] | 0.0051 | 0.3044 | [ 0.3292, 0.2799] |
| Moirai-2.0 | 0.1609 | 0.0244 | [ 0.0339, 0.0150] | 0.2173 | 0.0922 | [ 0.1076, 0.0768] |
| AP skill | AUC skill | |||||
| Skill | reference | Skill | reference | |||
| Training-free reference | 0.0154 | – | 0.0155 | – | ||
| Chronos-2 | 0.0187 | 0.0033 | [0.0002, 0.0064] | 0.0175 | 0.0020 | [ 0.0020, 0.0062] |
| Chronos-Bolt | 0.0149 | 0.0005 | [ 0.0036, 0.0027] | 0.0113 | 0.0042 | [ 0.0085, 0.0002] |
| Moirai-1.1 | 0.0043 | 0.0111 | [ 0.0139, 0.0085] | 0.0046 | 0.0109 | [ 0.0148, 0.0070] |
| Moirai-2.0 | 0.0214 | 0.0059 | [0.0029, 0.0091] | 0.0192 | 0.0037 | [ 0.0004, 0.0077] |
| AP skill | AUC skill | |||||||||
| Backbone | Dataset | Probe | Pretr. | Rand. | Difference [95% CI] | Pretr. | Rand. | Difference [95% CI] | ||
| Chronos-2 | Amazon | linear | 0.0130 | 0.0022 | 0.0108 | [0.0086, 0.0130] | 0.0327 | 0.0035 | 0.0292 | [0.0231, 0.0356] |
| Chronos-2 | Amazon | GBM | 0.0044 | 0.0001 | 0.0045 | [0.0026, 0.0066] | 0.0106 | 0.0012 | 0.0094 | [0.0037, 0.0153] |
| Chronos-2 | Favorita | linear | 0.1932 | 0.0801 | 0.1131 | [0.1113, 0.1149] | 0.1938 | 0.0817 | 0.1121 | [0.1105, 0.1136] |
| Chronos-2 | Favorita | GBM | 0.1811 | 0.0661 | 0.1150 | [0.1132, 0.1168] | 0.1854 | 0.0692 | 0.1163 | [0.1146, 0.1178] |
| Chronos-2 | M5 | linear | 0.0762 | 0.0148 | 0.0614 | [0.0593, 0.0635] | 0.1016 | 0.0224 | 0.0792 | [0.0769, 0.0814] |
| Method | Amazon | Favorita | M5 | NYC taxi | Web traffic |
| Random ordering | 1.00 | 4.34 | 2.87 | 2.11 | 2.75 |
| Training-free reference | 1.07 | 4.80 | 3.18 | 3.37 | 2.78 |
| Best point forecast | 1.06 | 4.79 | 3.16 | 3.54 | 2.86 |
| Raw logistic | 1.06 | 4.89 | 3.24 | 3.13 | 2.89 |
| Raw LightGBM | 1.07 | 5.01 | 3.22 | 3.51 | 3.65 |
| Probe: Chronos-2 | 1.09 | 4.98 | 3.27 | 3.50 | 3.10 |
| Dataset | Weighting | Reference | Median ref. | Max ref. | Raw LightGBM ref. [95% CI] |
| AP skill | |||||
| Amazon | window | 0.0078 | 0.0025 | 0.0022 | 0.0013 [ 0.0006, 0.0034] |
| series | 0.0070 | 0.0019 | 0.0042 | 0.0016 [ 0.0013, 0.0047] | |
| Favorita | window | 0.1026 | 0.0067 | 0.0228 | 0.0931 [0.0911, 0.0952] |
| series | 0.0923 | 0.0069 | 0.0217 | 0.0981 [0.0953, 0.1008] | |
| M5 | window | 0.0434 | 0.0052 | 0.0127 | 0.0207 [0.0192, 0.0223] |
| AP skill, linear GBM | AUC skill, linear GBM | |||
| Dataset | equal-window | equal-series | equal-window | equal-series |
| Amazon | 0.005 to 0.009 | 0.005 to 0.007 | 0.009 to 0.022 | 0.010 to 0.020 |
| Favorita | 0.009 to 0.025 | 0.010 to 0.024 | 0.006 to 0.022 | 0.005 to 0.019 |
| M5 | 0.012 to 0.017 | 0.012 to 0.017 | 0.015 to 0.020 | 0.015 to 0.020 |
| NYC taxi | 0.013 to 0.034 | 0.013 to 0.031 | 0.018 to 0.056 | 0.026 to 0.056 |
| Web traffic | 0.047 to 0.029 | 0.065 to 0.045 | 0.052 to 0.034 | 0.066 to 0.043 |
| Contrast | Amazon | Favorita | M5 | Web traffic |
| Probe point forecast | 6 / 5 | 6 / 6 | 6 / 6 | 6 / 6 |
| Probe raw logistic | 2 / 2 | 5 / 6 | 6 / 6 | 6 / 6 |
| Probe raw LightGBM | 3 / 1 | 5 / 6 | 6 / 6 | 6 / 6 |
| Pretrained randomized | 5 / 3 | 6 / 6 | 6 / 6 | 6 / 6 |
| Unfiltered | (a) Context and target | (b) Context | (c) Context 5% | |
| Windows / series | 41,816 / 1,597 | 1,907 / 210 | 2,406 / 445 | 6,208 / 673 |
| AP skill | ||||
| LightGBM linear probe | 0.115 (6/0/0) | 0.039 (6/0/0) | 0.047 (6/0/0) | 0.041 (6/0/0) |
| LightGBM GBM probe | 0.075 (6/0/0) | 0.049 (6/0/0) | 0.058 (6/0/0) | 0.045 (6/0/0) |
| Linear probe raw logistic | 0.033 (6/0/0) | 0.018 (2/4/0) | 0.023 (4/2/0) | 0.020 (6/0/0) |
| Pretrained randomized | 0.021 (5/0/1) | 0.013 (3/3/0) | 0.052 (5/1/0) | 0.028 (5/1/0) |
| TSFM | Amazon | Favorita | M5 | NYC taxi | Web traffic |
| Median output | |||||
| Chronos-2 | 93.3 (84.4) | 24.0 (18.0) | 49.0 (37.8) | 29.6 (26.3) | 53.0 (36.3) |
| Chronos-Bolt | 90.4 (73.5) | 22.5 (18.2) | 37.5 (29.0) | 27.2 (20.4) | 38.9 (28.7) |
| Moirai-2.0 | 92.1 (81.6) | 28.8 (23.7) | 54.3 (45.2) | 56.0 (48.2) | 56.8 (43.0) |
| TimesFM-2.5 | 76.7 (66.0) | 14.4 (10.7) | 24.6 (17.9) | 25.9 (21.7) | 27.8 (21.2) |
| TimesFM-3 | 92.3 (85.9) | 25.7 (21.5) | 47.8 (40.5) | 36.9 (33.8) | 52.6 (43.2) |
| AP skill | AUC skill | |||||||
| Point | Approx. | [95% CI] | Point | Approx. | [95% CI] | |||
| Amazon | ||||||||
| Chronos-2 | 0.0057 | 0.0007 | 0.0051 | [ 0.0068, 0.0034] | 0.0100 | 0.0012 | 0.0113 | [ 0.0156, 0.0071] |
| Chronos-Bolt | 0.0055 | 0.0033 | 0.0022 | [ 0.0035, 0.0010] | 0.0116 | 0.0047 | 0.0069 | [ 0.0107, 0.0031] |
| Moirai-1.1 | 0.0004 | 0.0003 | 0.0007 | [ 0.0022, 0.0009] | 0.0009 | 0.0006 | 0.0003 | [ 0.0049, 0.0055] |
| Moirai-2.0 | 0.0050 | 0.0002 | 0.0052 | [ 0.0067, 0.0038] | 0.0050 | 0.0047 | 0.0097 | [ 0.0140, 0.0052] |
| AP skill | AUC skill | |||||||||
| Max ref. | Max ref. | |||||||||
| Dataset | Above | Median | Point | Avg. | Probes | Above | Median | Point | Avg. | Probes |
| Amazon | 4/7 | 0.003 | 0.002 | 0.003 | 2/6 | 4/7 | 0.006 | 0.006 | 0.006 | 5/6 |
| Favorita | 7/7 | 0.012 | 0.023 | 0.041 | 5/6 | 6/7 | 0.012 | 0.005 | 0.028 | 5/6 |
| M5 | 7/7 | 0.011 | 0.013 | 0.032 | 1/6 | 5/7 | 0.011 | 0.001 | 0.031 | 2/6 |
| NYC taxi | 7/7 | 0.022 | 0.031 | 0.057 | 0/6 | 7/7 | 0.059 | 0.020 | 0.039 | 2/6 |
| Contrast | Amazon | Favorita | M5 | NYC taxi | Web traffic |
| AP skill | |||||
| Probe readout, point, scaled | 0.006 (6/0/0) | 0.043 (6/0/0) | 0.027 (6/0/0) | 0.009 (4/2/0) | 0.040 (6/0/0) |
| Probe readout, point, log | 0.006 (6/0/0) | 0.021 (6/0/0) | 0.017 (6/0/0) | 0.005 (3/3/0) | 0.037 (6/0/0) |
| Probe readout, point + quantiles, scaled | 0.008 (4/0/0) | 0.035 (4/0/0) | 0.025 (4/0/0) | 0.005 (3/0/1) | 0.034 (4/0/0) |
| Probe readout, point + quantiles, log | 0.007 (4/0/0) | 0.018 (4/0/0) | 0.017 (4/0/0) | 0.003 (2/1/1) | 0.029 (4/0/0) |
| Readout point forecast, scaled | 0.001 (0/4/2) | 0.038 (6/0/0) | 0.004 (4/1/1) | 0.009 (4/2/0) | 0.008 (6/0/0) |
| Pattern difference | |||||||
| Backbone | Eff.-rank ratio | Rate diff. | Amazon | Favorita | M5 | NYC taxi | Web traffic |
| Chronos-2 | 0.41–0.72 | 0.188–0.603 | 0.145 | 0.336 | 0.331 | 0.389 | 0.370 |
| Chronos-Bolt | 2.74–4.02 | 0.489–0.739 | 0.130 | 0.352 | 0.281 | 0.272 | 0.328 |
| Sundial | 0.60–0.70 | 0.301–0.564 | 0.080 | 0.186 | 0.142 | 0.126 | 0.253 |
| TimeMoE | 0.15–0.20 | 0.020–0.245 | 0.132 | 0.203 | 0.234 | 0.307 | 0.186 |
| Timer | 4.29–5.23 | 0.171–0.495 | 0.216 | 0.018 | 0.158 | 0.225 | 0.054 |