Beyond Average Error through Oracle-Informed Stress Tests for Time-Series Forecasting
Organizations: Tsinghua University
Abstract
Average squared error cannot reveal whether forecasting performance degrades because the future becomes less predictable or because forecasts move farther from the conditional mean. We introduce paired, mechanism-controlled stress tests that decompose changes in expected squared error at each lead time into environmental risk and forecast-oracle distance, using an origin-conditioned predictive oracle unavailable to the evaluated models. Three end-to-end controls have known attribution. Specifically, the null, environmental-only, and information-gap controls verify that the pipeline assigns changes to the correct component. We then apply the benchmark to 24 deployable forecasters. Under frequent switching, 14 methods have higher realized MSE but lower oracle distance; under outlier-variance feedback, 19 have higher MSE but lower scale-standardized MSE. Short- and long-lead stress-response rankings have Spearman correlation 0.624, revealing substantial horizon-dependent reordering. We then study multivariate relation shifts. Across six models and three coupling severities, oracle distance accounts for only 0.7-3.9% of the decomposed expected-risk increase, and environmental-risk majority persists in an eight-channel system and a matched-difficulty audit of Ring, Block, and Hub relations. Finally, prespecified contrasts on independent data-generating process (DGP) realizations show that several visually compelling discovery profiles, including trend accumulation and the hypothesized switching reversal, do not replicate. The benchmark thus combines component-wise diagnosis with a held-out stability audit. It complements real-data out-of-distribution evaluation, which measures performance under realistic shifts when exact oracle attribution is unavailable.
Figures & tables
| A. End-to-end controls with known attribution | ||||
| Evidence | Scope | Reference quantity | Observed model term | Audit decision |
| Identical-data null | 400 profiles | all 400 exactly zero | ||
| Innovation scale | 100 seeds | |||
| Delete latest AR input | 100 seeds | ; | 95% CI | |
| B. Coupling attribution across severity and structure | ||||
| Evidence | Scope | Environmental | Added-distance fraction | Upper-bound audit |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Record | Frozen | SHA-256 prefix | Purpose |
|---|---|---|---|
| Stage-2 execution v3 | Sep 10, 20:58 | 7951e4b02fa39a86 | Documents uniform OOM fallback; science unchanged |
| Stage-2.5/3 plan | Sep 10, 23:05 | 27c3500e69930ebb | Locks new seeds, tasks, and All-30 analysis |
| Stage-3 clarification | Sep 10, 23:50 | fb56eb02cf0f58bd | Fixes estimands before model results |
| Stage-4 manifest | Sep 11, 13:32 | 50fadb5ec398f18f | Locks controls, severity, and second system |
| Topology manifest | Sep 11, 15:27 | bb0346f88d410ad5 | Locks matched Ring/Block/Hub audit |
| Hyp. | Alternative | A: effect [95% CI]; | B: effect [95% CI]; | C: effect [95% CI]; |
|---|---|---|---|---|
| H1 | two-sided | [ , ] | [ , ] | [ , ] |
| H2 | [ , ] | [ , ] | [ , ] | |
| H3 | [ , ] | [ , ] | [ , ] | |
| H4 | [ , ] | [ , ] | [ , ] | |
| H5 | two-sided | [ , ] | [ , ] | [ , ] |
| Sign-flip sensitivity: Holm-adjusted (H5 unadjusted as a one-test family) | ||||
| Model | MSE real | [95% CI] | ||||
|---|---|---|---|---|---|---|
| Chronos-2 | 0.01242 | 0.01263 | 0.01235 | 0.00028 | 0.022 [0.013, 0.033] | 0.031 |
| TTM | 0.01258 | 0.01264 | 0.01235 | 0.00029 | 0.023 [0.013, 0.035] | 0.033 |
| DLinear | 0.01250 | 0.01263 | 0.01235 | 0.00028 | 0.022 [0.017, 0.028] | 0.027 |
| PatchTST | 0.01275 | 0.01289 | 0.01235 | 0.00054 | 0.042 [0.034, 0.050] | 0.049 |
| TSMixer | 0.01227 | 0.01251 | 0.01235 | 0.00016 | 0.013 [0.008, 0.018] | 0.017 |
| iTransformer | 0.01223 | 0.01259 | 0.01235 | 0.00024 | 0.019 [0.013, 0.027] | 0.025 |
| Model | Origin 3,712 | Origin 3,904 |
|---|---|---|
| Chronos-2 | 0.006 (0.009) | 0.038 (0.055) |
| TTM | 0.005 (0.008) | 0.040 (0.060) |
| DLinear | 0.009 (0.013) | 0.035 (0.043) |
| PatchTST | 0.005 (0.008) | 0.075 (0.088) |
| TSMixer | 0.010 (0.016) | 0.015 (0.024) |
| iTransformer | 0.009 (0.014) | 0.029 (0.042) |
| Model | 95% CI | ||
|---|---|---|---|
| Chronos-2 | 0.0020 | [-0.0000, 0.0041] | 0.324 |
| TimesFM | 0.0032 | [-0.0004, 0.0067] | 0.378 |
| AR | -0.0002 | [-0.0004, 0.0001] | 0.638 |
| NST | 0.0007 | [-0.0003, 0.0017] | 0.592 |
| SegRNN | 0.0002 | [-0.0001, 0.0004] | 0.638 |
| TSMixer | -0.0161 | [-0.1040, 0.0719] | 0.707 |
| Model | 0.26 | 0.34 | 0.42 |
|---|---|---|---|
| Chronos-2 | 1.05 (1.65) | 1.25 (2.43) | 2.33 (6.52) |
| TTM | 1.46 (2.37) | 1.64 (2.60) | 2.01 (3.41) |
| DLinear | 1.68 (2.11) | 1.74 (2.18) | 1.94 (2.54) |
| PatchTST | 3.17 (4.30) | 3.39 (4.45) | 3.93 (5.15) |
| TSMixer | 0.68 (1.03) | 0.71 (1.04) | 0.85 (1.41) |
| iTransformer | 1.14 (1.77) | 1.20 (1.90) | 1.46 (2.37) |
| Model | (%) | (%) | ||
|---|---|---|---|---|
| DLinear | 0.003110 | 0.000041 | 1.29 | 1.61 |
| PatchTST | 0.003110 | 0.000095 | 2.96 | 3.92 |
| Chronos-2 | 0.003110 | 0.000026 | 0.84 | 2.19 |
| System | Channels | Graph | Coupling Base Shift | Spectral radius Base Shift | Models | Seeds | |
|---|---|---|---|---|---|---|---|
| Ring VAR | 4 | directed cycle | 6 | 20 | |||
| Two-community VAR | 8 | two dense blocks | 3 | 20 |
| Ring | Block | Hub | ||||
|---|---|---|---|---|---|---|
| Model | ||||||
| Chronos-2 | 0.44 (0.72) | 2.04 (4.45) | 0.30 (0.43) | 2.92 (5.04) | 0.84 (1.42) | 0.69 (3.56) |
| TTM | 0.55 (0.81) | 2.71 (4.63) | 0.23 (0.46) | 5.78 (9.84) | 0.86 (1.51) | 3.59 (8.28) |
| DLinear | 0.61 (1.07) | 2.85 (3.72) | 0.44 (0.70) | 4.10 (5.61) | 0.97 (1.67) | 2.80 (3.61) |
| PatchTST | 0.61 (1.04) | 6.02 (8.04) | 0.55 (0.92) | 9.05 (11.97) | 1.11 (2.02) | 6.11 (8.17) |
| TSMixer | 0.61 (0.95) | 0.80 (1.54) | 0.62 (0.92) | 1.79 (3.12) | 1.11 (2.20) | 1.17 (2.21) |
| Lead block | Base | Shift | Shift Base |
|---|---|---|---|
| 1–24 | 0.550 [0.410, 0.690] | 0.261 [0.248, 0.274] | -0.289 [-0.425, -0.153] |
| 169–192 | 0.250 [0.250, 0.250] | 0.250 [0.250, 0.250] | -0.000 [-0.000, 0.000] |
| Benchmark | Real-data breadth | Controlled mechanisms | Strict pairing | Origin oracle | Risk split | profile | Adaptation |
|---|---|---|---|---|---|---|---|
| TFB / GIFT-Eval ( Qiu et al., 2024 ; Aksu et al., 2024 ) | Yes | – | – | – | – | – | – |
| SynTSBench ( Tan et al., 2025 ) | – | Yes | – | Partial | – | Partial | – |
| FinStressTS ( Sun et al., 2026 ) | – | Yes | Partial | Partial | – | Partial | – |
| Ours | Partial | Yes | Yes | Yes | Yes | Yes | Yes |
| Protocol | Base | Shift | What the contrast supports |
|---|---|---|---|
| Statistical anchors | Refit from visible Base context | Refit from visible Shift context | Stress response of a fixed fitting rule |
| Univariate target-trained (15) | Train/validate/test on Base (70/10/20) | Separately train/validate/test on Shift (70/10/20) | Mechanism-controlled response of an environment-specific fitting procedure |
| Multivariate target-trained | Fit on common pre-change prefix | Identical pre-change data/seed; no post-change update | Immediate OOD at the change point; passive contextual response at the later origin |
| Frozen TSFM (5) | Fixed pretrained checkpoint | The identical checkpoint; only context changes | Fixed-model cross-environment stress |
| Matched adaptation (5) | Random/zero/adapt arms on Base support/test | Separately matched arms on Shift support/test | Whether environment-specific limited-data intervention repairs the diagnosed gap |
| Entry | Protocol | Raw MSE | Oracle distance | SMSE |
|---|---|---|---|---|
| DGP oracle mean | Evaluator reference | 0.2558 | 0.0000 | 0.9948 |
| AR | Statistical | 0.2700 | 0.0159 | 1.0148 |
| ETS | Statistical | 0.3339 | 0.0912 | 1.1516 |
| Naive | Statistical | 0.4464 | 0.2185 | 1.8908 |
| Seasonal naive | Statistical | 0.5994 | 0.3509 | 2.4151 |
| SegRNN | Target-trained | 0.2672 | 0.0063 | 1.0213 |
| Term | Definition | Interpretation | Interpretive scope |
|---|---|---|---|
| Raw risk | Operational squared loss | Jointly reflects environmental risk and oracle distance | |
| Environmental risk | Evaluator-conditioned uncertainty | Uses declared pre-origin information only | |
| Forecast–oracle distance / | Information restriction plus model approximation/estimation | Combined distance, not architecture error alone | |
| SMSE | Mean pointwise error relative to local predictive scale | Secondary diagnostic, not MASE or probabilistic calibration | |
| Stress profile | Paired response over mechanism and individual lead | Aggregation requires deployment weights over |
| Stress | Model-resolved evidence | Structure-linked interpretation | Evidential status |
|---|---|---|---|
| Strong trend | AR is numerically unchanged. TSMixer increases by , TTM by , and Time-MoE by ; PatchTST changes by . | AR explicitly fits a linear trend in each environment. Persistence and fixed-window mixing do not enforce slope extrapolation, which can produce a level-dependent direct forecast. PatchTST shows that lacking an explicit trend term is not sufficient for failure. | The TSMixer failure is clear in discovery, but the locked H2 claim about stronger late-lead accumulation does not replicate. |
| Frequent switching | Descriptively, Naive increases by , whereas AR, ETS, TSMixer, and Chronos-2 change by , , , and ; ETS is not in the locked panel. The formal H3/H4 panel is AR, Chronos-2, TimesFM, and TSMixer, whose late effect is negative in all three held-out batches. | A last-value predictor persists the current regime and becomes stale after rapid switches. Other forecasting procedures may approach the common long-run predictive center more closely under faster switching. | The panel-level late reduction is confirmed. Its mechanism and individual-model differences remain descriptive. |
| Outlier-variance feedback | Nineteen of 24 methods combine higher raw MSE with lower SMSE. Seasonal naive and naive increase their oracle distance by and , while AR changes by only . | Copying a recent or seasonal observation can propagate an event-contaminated value. A fitted low-order dynamic can dampen that transient. Much of the raw-loss increase still reflects predictive scale rather than model distance. | The cross-protocol direction is broad, but the proposed propagation route has not been isolated by an architectural ablation. |
| Relation shift | At , every Base and Shift forecast is bitwise identical. At , PatchTST fractions are , , and for Ring, Block, and Hub, compared with , , and for TSMixer. | Because weights are fixed, the later separation reflects how each architecture converts 192 post-shift observations into a forecast. Patch tokenization and temporal mixing are associated with different passive context responses, not different online adaptation. | The fractions replicate on independent topology seeds; which internal operation causes the difference remains open. |
| Control | Prespecified pass rule | Observed attribution | Result |
|---|---|---|---|
| Identical-data null | all paired effects are numerically zero | 400/400 profiles: | pass |
| Innovation | and | , , | pass |
| Delete last AR input | and | , , , CI [0.000262, 0.000372] | pass |
| Pair | Case role(s) | Mechanism | Exact parameters | Formal horizons |
|---|---|---|---|---|
| Trend strength | Base; moderate; strong | linear mean | , | 1,24,96,192 |
| Frequency acceleration | Base; Shift | sinusoidal mean | , , | 1,24,96,192 |
| Innovation shape | Base; ; skew; mixture | standardized | Gaussian; ; shape ; | 1,24,96,192 |
| GARCH persistence | Base; Shift | GARCH(1,1) | 1,24,96,192 | |
| Markov rate | Base; Shift | two-state Gaussian | , , stay | 1,24,96,192 |
| Additive outliers | Base; strong | symmetric events | 1,24,96,192 |
| Implementation | lr / bs | Main dimensions | Other fixed settings |
|---|---|---|---|
| TimeKAN | downsample layers/window , order 0 | ||
| TimeMixer | average downsample layers/window | ||
| PaiFilter | hidden 256 | label length 0 | |
| TexFilter | embed/hidden | dropout 0 | |
| TimesNet | top- , factor 3 | ||
| SegRNN | dropout .5; segment for , else |
| Model | Repository | Revision prefix | Context/horizon and output |
|---|---|---|---|
| Chronos-2 | amazon/chronos-2 | 29ec3766d36d | ; point + 9 quantiles |
| TimesFM 2.5 | google/timesfm-2.5-200m-pytorch | 1d952420fba8 | ; point + 9 quantiles |
| Moirai 2.0 R-small | Salesforce/moirai-2.0-R-small | 30f43ff08c84 | ; point + 9 quantiles |
| TTM-R2 | ibm-granite/granite-timeseries-ttm-r2 | d6a79570cac0 / 25f4a00a25e1 | : native 512–96; : native 512–192; point only |
| Time-MoE-200M | Maple728/TimeMoE-200M | 794591bfeb12 | autoregressive; point only |
| Model | Zero | Adapt | [95% CI] | Repair [95% CI] | |
|---|---|---|---|---|---|
| TimesFM | .279 | .253 | [ ] | 9.3% [3.9,13.5] | .031 |
| Chronos-2 | .271 | .271 | [ ] | 0.04% [ ] | .891 |
| Moirai | .300 | .270 | [ ] | 9.9% [6.0,14.3] | .010 |
| TTM | .344 | .256 | [ ] | 25.5% [21.5,30.1] | .010 |
| Time-MoE | .354 | .255 | [ ] | 27.8% [25.2,31.1] | .010 |
| Donor (DGP seeds) | Frozen model | Reference MSE | Donor MSE | raw MSE | center distance | Paired windows |
|---|---|---|---|---|---|---|
| UCI Air Quality (10) | Chronos-2 | 0.0405 | 0.0369 | -0.0036 | +0.0055 | 80 |
| TimesFM | 0.0423 | 0.0764 | +0.0341 | +0.0389 | 80 | |
| Moirai | 0.0505 | 0.0612 | +0.0108 | +0.0174 | 80 | |
| TTM | 0.1169 | 0.1040 | -0.0129 | -0.0081 | 80 | |
| Time-MoE | 0.1230 | 0.1384 | +0.0154 | +0.0211 | 80 | |
| UCI Electricity (10) | Chronos-2 | 0.0405 | 0.0378 | -0.0027 | +0.0078 | 80 |
| V only | F/M + V | F + M + V | ||||
|---|---|---|---|---|---|---|
| DED | DED | DED | ||||
| 24 | 0.0196 | 0.0073 | 0.3584 | 0.2699 | 0.4709 | 0.5390 |
| 48 | 0.0204 | 0.0078 | 0.3340 | 0.2913 | 0.4587 | 0.5749 |
| 96 | 0.0179 | 0.0069 | 0.3385 | 0.3196 | 0.4523 | 0.6143 |
| 192 | 0.0170 | 0.0065 | 0.3484 | 0.3499 | 0.4494 | 0.6516 |