Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting
Organizations: JD.com, Inc.
Abstract
Time-series forecasting models achieve strong benchmark performance but exhibit severe systematic bias in industrial deployments. This train--deploy gap is conventionally attributed to temporal-structural errors or distribution shifts. We characterize a complementary source that these explanations overlook: canonical losses embed fixed statistical priors, while industrial demand mixes benign and pathological regimes---zero-inflation, skewness, high variability---in which these priors are systematically violated. The induced bias persists even under perfect temporal modeling, remains in a distributional-shape component that normalization cannot remove, and creates an aggregation trade-off invisible to aggregate metrics. We turn these observations into an evaluation toolkit centered on the Regime-wise Relative Bias Vector (RBV): a metric-agnostic, regime-decomposed diagnostic that audits how pooled training allocates systematic mismatch across pathological subpopulations. A controlled attribution analysis decomposes RBV into a model-independent intrinsic floor, set by each loss's estimand, and an excess component attributable to training, tracing observed bias to the loss rather than the model. A large-scale study---13 loss objectives, 3 seeds, 60,000+ series spanning RetailShiftBench and M5, with random-split controls---shows that regime-aware diagnosis separates optimization-type from bias-type failure, and that regime-aware training resolves the pooling-induced bias that capacity scaling cannot, for mean-type losses. A formal structural observation, that risk under evaluation-distribution contamination is affine in the pathology mixture weight, grounds these findings. Our work complements model ranking with mechanism-grounded, regime-oriented evaluation.
Figures & tables
| Loss | Implicit Distributional Priors | Industrial Departure | Mismatch Mechanism |
|---|---|---|---|
| MSE | Symmetric light-tailed residuals, stable finite variance; mean as optimal fitting target | Heavy-tailed skewed demand with unstable variance | Quadratic penalty amplifies burst residuals and averages out tail signals; the mean estimand systematically overshoots sparse regimes. |
| MAE | Symmetric continuous distribution; median as optimal estimation target | Skewed, burst-heavy demand with frequent zero-inflation | Zeros and a long upper tail pull the median below the mean; the burst mass carrying most expected demand is discarded, widening the gap |
| Quantile | Stationary distribution; fixed quantile targets valid across sub-populations; non-degenerate tails | Zero-inflated baseline with skewed long-tail demand | Zero-inflation collapses lower-tail quantiles; the estimand degenerates into a step function under sparse bursts |
| Poisson | Variance mean; discrete counts; mean estimand under exponential curvature | Over-dispersed burst clustering with variance exceeding the mean | Burst clustering breaks equi-dispersion and mis-sets the fixed exponential curvature; the mass redistribution is invisible to linear-residual metrics |
| Huber | Stationary residual scale; outliers are genuine rare anomalies | Heavy-tailed baseline with frequent extreme demand bursts | A fixed threshold clips legitimate bursts as outliers; the effective estimand drifts with per-cell residual scale |
| Tweedie | Fixed dispersion and Tweedie-power parameters shared across all samples | Heterogeneous SKU-level dispersion under varying industrial statistical regimes | Static dispersion parameters fail to fit heterogeneous regime-wise distributions, leaving the estimand mismatched to varying statistical regimes. |
| Semantics | Regime | XGB | TFT | Rule | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MSE | MAE | QL 0.60 | Huber | Tweedie | Poisson | MSE | MAE | QL 0.60 | Huber | Tweedie | Poisson | |||
| Bias | All | 0.035 | -0.088 | 0.045 | -0.035 | 0.025 | -0.01 | -0.004 | -0.076 | 0.05 | -0.03 | -0.039 | -0.086 | 0.033 |
| Benign | 0.022 | -0.03 | 0.073 | -0.02 | 0.017 | -0.043 | 0.015 | 0.016 | 0.095 | 0.022 | 0.038 | -0.001 | 0.031 | |
| Skew | 0.061 | -0.185 | -0.003 | -0.051 | 0.044 | 0.048 | -0.039 | -0.232 | -0.043 | -0.116 | -0.17 | -0.223 | 0.045 | |
| CV | 0.262 | -0.685 | -0.421 | -0.042 | 0.176 | 0.341 | -0.33 | -0.894 | -0.726 | -0.437 | -0.838 | -0.868 | 0.057 | |
| ZIR | 0.116 | -0.385 | -0.114 | -0.110 | 0.085 | 0.156 | -0.082 | -0.53 | -0.182 | -0.277 | -0.443 | -0.532 | 0.038 | |
| XGB | TFT | Rule | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Metric | MSE | MAE | Huber | Tweedie | MSE | MAE | Huber | Tweedie | |
| RBV | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | |
| WMAPE | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | |
| MSE | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | ( ) | |
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
| Variant | Base functional | Units | Question answered |
|---|---|---|---|
| Pathology bias-RBV (default, § 4.2 ) | demand units | Which regime is systematically over-/under-estimated, and by how much? | |
| Quantile RBV | probability | Where – and at which – is calibration error concentrated? | |
| Share-normalized RBV | dimensionless (relative bias) | How is mismatch severity allocated per unit of business volume? |
| Loss | [95% CI] | RBV [95% CI] |
|---|---|---|
| MAE | ||
| MSE | ||
| Huber | ||
| Tweedie | ||
| Poisson | ||
| Quantile |
| Pathology | Retail | Energy / Grid | Manufacturing / IoT | Finance |
|---|---|---|---|---|
| Skew | Promo bursts; seasonal peaks | Extreme weather load spikes | Defect clustering; rework spikes | Jump‑driven return asymmetry; earnings announcement effects |
| ZIR | Stockouts; intermittently stocked SKUs | Turbine downtime; zero irradiance | Sensor dropout; line stoppages | Halted trading; zero‑tick intervals in illiquid assets |
| CV | Long‑tail SKU heterogeneity | Ramping‑driven output dispersion across seasons | Batch‑to‑batch variance | Cross‑sectional dispersion of return scales |
| Coupled Distributional Pathologies | Sporadic luxury bursts; Skew‑CV: high‑variance promo series | Zero‑output + sudden load spikes; weather‑driven price spikes | Stoppage + large‑batch orders; high‑variance defect batches | Illiquidity + price jumps; earnings‑driven clustered volatility |
| Coupling with Temporal‑Structure | Pathology + holiday / lifecycle shifts | Pathology + policy / tariff changes | Pathology + line changeovers | Pathology + macro regime transitions |
| Dataset | #Series | Domain & Granularity | Low Skewness | Low CV | Low ZIR | PCS | Benchmark Grade |
|---|---|---|---|---|---|---|---|
| ETT | 28 | Energy (transformer temp), hourly/15-min | 0.5714 | 0.8571 | 0.8571 | 0.2381 | Unqualified; basic temporal verification only |
| M4-Daily | 4,227 | Mixed (macro micro), daily | 0.863 | 0.9970 | 1.000 | 0.046 | Insufficient pathological spectrum |
| M4-Hourly | 414 | Mixed (macro micro), hourly | 0.865 | 0.993 | 1.0 | 0.048 | Insufficient pathological spectrum |
| M4-Monthly | 48,000 | Mixed (macro micro), monthly | 0.845 | 0.999 | 1.0 | 0.052 | Insufficient pathological spectrum |
| M4-Quarterly | 24,000 | Mixed (macro micro), quarterly | 0.849 | 0.999 | 1.0 | 0.051 | Insufficient pathological spectrum |
| M4-Weekly | 359 | Mixed (macro micro), weekly | 0.652 | 0.997 | 1.0 | 0.117 | Insufficient pathological spectrum |
| non-linear loss family | ||
|---|---|---|
| Metric | MSE | Tweedie |
| RBV | ||
| WMAPE | ||
| MSE | ||
| linear loss family | ||
| Metric | MAE | Huber |
| non-linear loss family | |||
|---|---|---|---|
| Metric | MSE | Tweedie | Poisson |
| RBV | |||
| WMAPE | |||
| MSE | |||
| linear loss family | |||
| Metric | MAE | QL60 | Huber |
| RBV | WMAPE | |||||
|---|---|---|---|---|---|---|
| Loss | All | Random | Regime | All | Random | Regime |
| MSE | 0.473 | 0.472 | 0.400 | 0.7526 | 0.7381 | 0.6764 |
| MAE | 0.428 | 0.454 | 0.416 | 0.6660 | 0.6918 | 0.6625 |
| QL60 | 0.420 | 0.437 | 0.480 | 0.6791 | 0.7096 | 0.6910 |
| Huber | 0.397 | 0.390 | 0.369 | 0.6911 | 0.6887 | 0.6659 |
| Tweedie | 0.500 | 0.469 | 0.421 | 0.7818 | 0.7236 | 0.6810 |
| Method | RBV | WMAPE | |||||
| Croston | 0.020 | 0.739 | |||||
| SBA | 0.019 | 0.725 | |||||
| TSB | 0.028 | 0.734 | |||||
| TSB-HB | 0.024 | 0.739 | |||||
| Oracle mean | 0.036 | 0.738 | |||||
| Oracle median | 0.695 | 0.668 |
| estimand | RBV | WMAPE | |||||
|---|---|---|---|---|---|---|---|
| mean (quadrature) | |||||||
| median ( ) |