Time-series forecasting models achieve strong benchmark performance but exhibit severe systematic bias in industrial deployments. This train--deploy gap is conventionally attributed to temporal-structural errors or distribution shifts. We characterize a complementary source that these explanations overlook: canonical losses embed fixed statistical priors, while industrial demand mixes benign and pathological regimes---zero-inflation, skewness, high variability---in which these priors are systematically violated. The induced bias persists even under perfect temporal modeling, remains in a distributional-shape component that normalization cannot remove, and creates an aggregation trade-off invisible to aggregate metrics. We turn these observations into an evaluation toolkit centered on the Regime-wise Relative Bias Vector (RBV): a metric-agnostic, regime-decomposed diagnostic that audits how pooled training allocates systematic mismatch across pathological subpopulations. A controlled attribution analysis decomposes RBV into a model-independent intrinsic floor, set by each loss's estimand, and an excess component attributable to training, tracing observed bias to the loss rather than the model. A large-scale study---13 loss objectives, 3 seeds, 60,000+ series spanning RetailShiftBench and M5, with random-split controls---shows that regime-aware diagnosis separates optimization-type from bias-type failure, and that regime-aware training resolves the pooling-induced bias that capacity scaling cannot, for mean-type losses. A formal structural observation, that risk under evaluation-distribution contamination is affine in the pathology mixture weight, grounds these findings. Our work complements model ranking with mechanism-grounded, regime-oriented evaluation.
Figures & tables
Figure 1: Evolution of time-series forecasting evaluation paradigms.
Loss
Implicit Distributional Priors
Industrial Departure
Mismatch Mechanism
MSE
Symmetric light-tailed residuals, stable finite variance; mean as optimal fitting target
Heavy-tailed skewed demand with unstable variance
Quadratic penalty amplifies burst residuals and averages out tail signals; the mean estimand systematically overshoots sparse regimes.
MAE
Symmetric continuous distribution; median as optimal estimation target
Skewed, burst-heavy demand with frequent zero-inflation
Zeros and a long upper tail pull the median below the mean; the burst mass carrying most expected demand is discarded, widening the gap
Quantile
Stationary distribution; fixed quantile targets valid across sub-populations; non-degenerate tails
Zero-inflated baseline with skewed long-tail demand
Zero-inflation collapses lower-tail quantiles; the estimand degenerates into a step function under sparse bursts
Poisson
Variance = mean; discrete counts; mean estimand under exponential curvature
Over-dispersed burst clustering with variance exceeding the mean
Burst clustering breaks equi-dispersion and mis-sets the fixed exponential curvature; the mass redistribution is invisible to linear-residual metrics
Huber
Stationary residual scale; outliers are genuine rare anomalies
Heavy-tailed baseline with frequent extreme demand bursts
A fixed threshold clips legitimate bursts as outliers; the effective estimand drifts with per-cell residual scale
Tweedie
Fixed dispersion and Tweedie-power parameters shared across all samples
Heterogeneous SKU-level dispersion under varying industrial statistical regimes
Static dispersion parameters fail to fit heterogeneous regime-wise distributions, leaving the estimand mismatched to varying statistical regimes.
Table 1: Statistical Assumptions and Deployment Mismatch Patterns
Semantics
Regime
XGB
TFT
Rule
MSE
MAE
QL 0.60
Huber
Tweedie
Poisson
MSE
MAE
QL 0.60
Huber
Tweedie
Poisson
Bias
All
0.035
-0.088
0.045
-0.035
0.025
-0.01
-0.004
-0.076
0.05
-0.03
-0.039
-0.086
0.033
Benign
0.022
-0.03
0.073
-0.02
0.017
-0.043
0.015
0.016
0.095
0.022
0.038
-0.001
0.031
Skew
0.061
-0.185
-0.003
-0.051
0.044
0.048
-0.039
-0.232
-0.043
-0.116
-0.17
-0.223
0.045
CV
0.262
-0.685
-0.421
-0.042
0.176
0.341
-0.33
-0.894
-0.726
-0.437
-0.838
-0.868
0.057
ZIR
0.116
-0.385
-0.114
-0.110
0.085
0.156
-0.082
-0.53
-0.182
-0.277
-0.443
-0.532
0.038
Table 2: Systematic bias and cross-regime vector misalignment (RBV) under hierarchical model-loss.
XGB
TFT
Rule
Metric
MSE
MAE
Huber
Tweedie
MSE
MAE
Huber
Tweedie
RBV
0.243→0.041 ( −83.1% )
0.675→0.899 ( +33.2% )
0.077→0.262 ( +240.3% )
0.163→0.014 ( −91.4% )
0.335→0.14 ( −58.2% )
0.914→0.863 ( −5.6% )
0.459→0.572 ( +24.6% )
0.873→0.881 ( +0.9% )
0.033
WMAPE
0.4679→0.462 ( −1.3% )
0.4270→0.4255 ( −0.4% )
0.4670→0.4765 ( +2.0% )
0.4593→0.4539 ( −1.2% )
0.4675→0.4730 ( +1.2% )
0.4475→0.4487 ( +0.3% )
0.4603→0.4595 ( −0.2% )
0.4514→0.4549 ( +0.8% )
0.5182
MSE
11.909→11.068 ( −7.1% )
12.057→11.524 ( −4.4% )
13.575→20.176 ( +48.6% )
11.992→11.465 ( −4.4% )
15.68→14.55 ( −7.2% )
15.182→14.40 ( −5.2% )
14.140→14.84 ( +5.0% )
14.714→14.274 ( −3.0% )
16.452
Table 3: Effect of pathology decomposition.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2: Conceptual schematic of the aggregation-breadth trade-off between distributional-statistical bias and estimation variance. The global-pooled, random-cell (R2, random-label control), and pathology-cell markers correspond to measured configurations (Tab. 3 ; § D.2 , Tab. 8 ); remaining markers are positional schematics without measured magnitude. Measured anchors are consistent with, but do not locate, a total-error minimum near pathology-aware subgrouping.
Variant
Base functional ϕr
Units
Question answered
Pathology bias-RBV (default, § 4.2 )
∑i∈Ir∣yi∣∑i∈Ir(y^i−yi)
demand units
Which regime is systematically over-/under-estimated, and by how much?
Quantile RBV
ϕr(τ)=Er(τ)−τ
probability
Where – and at which τ – is calibration error concentrated?
Share-normalized RBV
Br/S^r
dimensionless (relative bias)
How is mismatch severity allocated per unit of business volume?
Appendix
Table 4: The RBV family: one template, three instantiations. All variants share the same centering and ℓ2 aggregation, and differ only in the base functional ϕr . Er(τ) is the empirical exceedance under the mid-convention Pr[q^τ>y]+21Pr[q^τ=y] (Figs. 8 , 11 ); S^r=∑i∈Iry^i is the regime-level predicted total, and Br=(∑i∈Ir(y^i−yi))/(∑i∈Ir∣yi∣) denotes the regime-wise pooled bias of § 4.2 .
Loss
B [95% CI]
RBV [95% CI]
MAE
−0.257[−0.267,−0.248]
0.409[0.391,0.427]
MSE
0.058[0.048,0.068]
0.419[0.387,0.450]
Huber
−0.107[−0.115,−0.098]
0.241[0.215,0.265]
Tweedie
0.054[0.044,0.063]
0.378[0.350,0.407]
Poisson
0.058[0.049,0.068]
0.477[0.445,0.507]
Quantile τ=0.1
−0.859[−0.866,−0.852]
0.229[0.218,0.242]
Appendix
Table 5: Bootstrap uncertainty (SKU-level, B=1000 ) for the pooled bias B and regime bias variation (RBV), single-origin M5 protocol. All RBV intervals exclude zero.
Figure 3: One-at-a-time threshold sensitivity of regime-decomposed RBV on M5. Each panel varies one threshold ( tzir , tskew , tcv ) with the other two fixed at baseline ( 0.20/1.00/1.50 , dotted line); the remaining 9 losses behave similarly and are omitted for clarity. Rankings and the U-shaped RBV– τ structure are invariant across all threshold settings; the dominant sensitivity is along tcv , whose spread tracks the extremity of the CV cell (89–10,746 SKUs across settings) rather than the threshold definition itself. Note: for this sensitivity analysis, regime cells are formed directly from the continuous pathology statistics (ZIR rate, skewness, CV) without applying the § 3 thresholds, so RBV magnitudes are not directly comparable to Table 2 (which uses the thresholded regime labels); only rankings and shapes are comparable.
Figure 4: Synthetic illustration and parameter sensitivity for Prop. B.1 . (a) Example trajectory under perfect temporal modeling Btemp=0 . (b) Quantitative comparison of loss-implied predictors (MSE-opt, MAE-opt, and a fixed-quantile target) under pinball evaluation. (c) Ablation over quantile level τ : the ranking of predictors flips as the target functional varies, illustrating that the optimal estimator is prior-dependent. (d) Ablation over zero-inflation rate p0 . This figure serves as an illustrative synthetic example; formal statements and proofs are given in the main text.
Figure 5: Dose-response ablation over three core pathological dimensions: zero-inflation rate (ZIR), skewness, and within-series scale variability (CV). We report squared and absolute deviation of the fitted predictor from the clean conditional trajectory as pathological intensity increases. Consistently, both deviation metrics rise monotonically with pathological severity, demonstrating that pathological marginals systematically corrupt the recovery of the underlying signal. This synthetic experiment complements the theoretical analysis in Prop. B.1 .
Figure 6: Coupled-pathology performance across varying α . The linear-sum baseline aggregates excess losses from isolated-ZIR and isolated-CV runs. Pink shaded regions highlight regimes where coupled loss strictly exceeds this reference, indicating emergent super-additive degradation.
Figure 7: RetailShiftBench temporal structure. Aggregate daily sales across all 31,213 series over the full evaluation panel. Sales are dominated by a stable weekly periodicity with no trend breaks or regime shifts, indicating that the benchmark’s difficulty stems from cross-sectional distribution heterogeneity rather than temporal non-stationarity.
Table 6: Cross‑domain instantiation of the pathology taxonomy.
Dataset
#Series
Domain & Granularity
Low Skewness
Low CV
Low ZIR
PCS
Benchmark Grade
ETT
28
Energy (transformer temp), hourly/15-min
0.5714
0.8571
0.8571
0.2381
Unqualified; basic temporal verification only
M4-Daily
4,227
Mixed (macro ∼ micro), daily
0.863
0.9970
1.000
0.046
Insufficient pathological spectrum
M4-Hourly
414
Mixed (macro ∼ micro), hourly
0.865
0.993
1.0
0.048
Insufficient pathological spectrum
M4-Monthly
48,000
Mixed (macro ∼ micro), monthly
0.845
0.999
1.0
0.052
Insufficient pathological spectrum
M4-Quarterly
24,000
Mixed (macro ∼ micro), quarterly
0.849
0.999
1.0
0.051
Insufficient pathological spectrum
M4-Weekly
359
Mixed (macro ∼ micro), weekly
0.652
0.997
1.0
0.117
Insufficient pathological spectrum
Appendix
Table 7: Pathology coverage and industrial validity of mainstream forecasting benchmarks. Higher PCS represents stronger industrial pathological spectrum coverage and better deployment-aligned robustness evaluation capability.
Figure 8: Quantile calibration as an attribution test (XGBoost, single- τ pinball objectives, tree depth fixed across τ ). Coverage is computed on positive-demand days with the mid convention P(y^>y)+21P(y^=y) ; the dashed line marks ideal calibration. Left : global pooling glues pathological regimes below the diagonal while benign stays over-covered. Middle : random-label decomposition with cells matched in count and size to the pathology partition (permutation without replacement) reproduces the glued pattern—granularity alone does not release the curves. Right : pathology-informed decomposition releases the bundle toward the diagonal, with the residual gap at τ=0.9 marking the variance wall. Near-zero low- τ coverage on pathological regimes reflects the estimand floor of zero-inflated marginals (§ 4.2 ), not additional miscalibration.
non-linear loss family
Metric
MSE
Tweedie
RBV
0.243→0.26→0.041↓
0.163→0.151→0.014↓
WMAPE
0.4679→0.4681→0.462↓
0.4593→0.4606→0.4539↓
MSE
11.909→12.615→11.068↓
11.992→12.678→11.465↓
linear loss family
Metric
MAE
Huber
Appendix
Table 8: Loss-function trajectory grid under grouping-scheme interventions (XGB).
Figure 9: Capacity scaling (solid) vs. regime decomposition (dashed) under the MSE loss. (a) RBV (log scale): pooled depth-2 RBV improves 7× ( 0.837→0.120 ) with 32× the trees, and the best pooled configuration overall (1600 trees, depth 4) saturates at a 0.116 ceiling—decomposition (dashed) holds RBV ≤0.06 across all but the most extreme configurations (within-cell overfitting at 1600 trees, depth 12). (b) MSE: decomposition improves aggregate accuracy as well. WMAPE varies within 0.46 – 0.50 across all conditions and is omitted.
Figure 10: Same setup as Fig. 9 under the MAE loss. Decomposition increases RBV (dashed curves lie at or above solid ones in (a)): routing removes the benign mass that dilutes MAE’s within-cell estimand mismatch, exposing the mean–median gap rather than repairing it—whereas aggregate error (b) still improves.
non-linear loss family
Metric
MSE
Tweedie
Poisson
RBV
0.422→0.436±0.0071→0.036↓
0.377→0.359±0.0019→0.062↓
0.476→0.487±0.0018→0.044↓
WMAPE
0.7597→0.7612±0.0004→0.7337↓
0.7553→0.7526±0.0005→0.7259↓
0.7612→0.7612±0.0001→0.7326↓
MSE
4.784→4.745±0.020→4.607↓
4.758→4.699±0.020→4.493↓
4.753→4.690±0.025→4.505↓
linear loss family
Metric
MAE
QL60
Huber
Appendix
Table 9: Loss-function trajectory grid under grouping-scheme interventions on M5 (XGB, fixed-origin protocol; mean over three seeds 42, 41, 39).
Figure 11: Quantile-calibration attribution on M5 (mean over three seeds 42, 41, 39; τ∈{0.1,…,0.9} ; τ=0.5 shown via MAE, since pinball(0.5)≡MAE ). Exceedance uses the mid convention Pr[y^>y]+21Pr[y^=y] . Left: global pooling — the aggregate curve hugs the diagonal while regime curves fan out. Center: random decomposition (size-matched control) — reproduces the pooled panel, isolating label informativeness. Right: regime decomposition — high- τ curves are released to the diagonal while the High-CV collapse at mid τ is exposed. Low- τ elevations reflect the zero-inflation estimand floor ( zir/2 ), not miscalibration.
RBV
WMAPE
Loss
All
Random
Regime
All
Random
Regime
MSE
0.473
0.472
0.400 ↓
0.7526
0.7381
0.6764 ↓
MAE
0.428
0.454
0.416 ↓
0.6660
0.6918
0.6625 ↓
QL60
0.420
0.437
0.480 ↑
0.6791
0.7096
0.6910 ↑
Huber
0.397
0.390
0.369 ↓
0.6911
0.6887
0.6659 ↓
Tweedie
0.500
0.469
0.421 ↓
0.7818
0.7236
0.6810 ↓
Appendix
Table 10: Regime decomposition on the M5 deep backbone (NHiTS). Each cell reports All → Random → Regime-wise training. All: global pooled training; Random: random split with cells matched in count and size to the pathology partition (mean over three shuffle seeds); Regime-wise: per-cell training on the skew–CV–ZIR partition. Single seed (42) for All and Regime-wise.
Method
Bbenign
Bzir
Bskew
Bcv
B
RBV
WMAPE
Croston
0.098
0.077
0.086
0.100
0.087
0.020
0.739
SBA
0.043
0.023
0.031
0.045
0.033
0.019
0.725
TSB
0.095
0.062
0.076
0.061
0.075
0.028
0.734
TSB-HB
0.057
0.035
0.036
0.059
0.043
0.024
0.739
Oracle mean
0.068
0.026
0.035
0.023
0.040
0.036
0.738
Oracle median
0.001
−0.456
−0.390
−0.941
−0.353
0.695
0.668
Appendix
Table 11: Intermittent-demand baselines and oracle references on the M5 single-origin protocol. Bias Br follows the paper’s sample-pooled sum-ratio convention, and RBV is the unweighted ℓ2 norm of the centered bias vector (identical to Tab. 8 ). Excess bias is relative to oracle-mean. Croston’s uniformly positive bias is its known over-prediction; SBA’s damping factor largely removes it. Oracle-median achieves the best WMAPE yet the largest regime bias variation: accuracy under MAE does not imply calibration for the pooled-mean estimand.
estimand
Bbenign
Bzir
Bskew
Bcv
Bglobal
RBV
WMAPE
mean (quadrature)
0.114
−0.119
−0.060
−0.328
0.008
0.381
0.703
median ( q0.5 )
0.070
−0.493
−0.368
−0.853
−0.194
0.789
0.672
Appendix
Table 12: Chronos-Bolt zero-shot on the M5 single-origin protocol (App. E.1 ), horizons 1–7 pooled ( n=187,222 samples). Mean vs. median estimands decoded from the same frozen forward pass. Bias =∑(y^−y)/∑∣y∣ per cell; RBV and WMAPE as in Tab. 2 .
Figure 12: Chronos-Bolt zero-shot evaluation on M5. Left: RBV and aggregate WMAPE as the decoding quantile τ is swept on a frozen checkpoint; the dotted vertical line marks τ∗=0.4 , the quantile minimizing aggregate WMAPE. RBV( τ ) is U-shaped and attains its minimum at the quadrature-mean estimand (Tab. 12 ). Right: per-regime exceed rate, Pr[y^>y]+0.5Pr[y^=y] , of zero-shot quantile decoding against the diagonal: quantile heads over-cover throughout the mid range (exceed 0.574 at nominal 0.5 ), the high-CV segment shows the largest over-coverage, and calibration recovers at high τ . Unlike Fig. 8 , the low- τ end shows no zero-inflation floor—the continuous quantile heads emit no atom at zero, so the ZIR curve tracks the diagonal while the floor in Fig. 11 is a construct of discrete zero-mass predictors. Negative predictions ( 65% at q0.1 ) mechanically depress the exceed rate at low τ (App. E.3 ).
East China Normal University, Shanghai, China. · University of Electronic Science and Technology of China, Chengdu, China. · Aalborg University, Aalborg, Denmark.