Scale-Invariant Training for Time Series Foundation Models
Organizations: Carnegie Mellon University · Amazon · Nixtla
Abstract
Time series foundation models (TSFMs) are trained on large collections of time series datasets that span various morphologies and domains. This setting exposes models to series whose scales -- typical magnitudes of their values -- can differ substantially. Affine scaling methods such as Reversible Instance Normalization (ReVIN) scale model inputs and reverse the transform before computing the loss. We show that this inversion multiplies each series' gradient by relative to loss on scaled targets, where is the scaling denominator (e.g., standard deviation) and is the loss degree. We call this scale-contaminated training (ScaleCon), because the scale of each series consequently becomes an importance weight, causing high-scale series to dominate training. For any scale-equivariant scaler and residual loss that is homogeneous of degree , including MSE, MAE, and Quantile Loss, we prove that computing loss on scaled targets makes every mini-batch gradient and, consequently, the full optimization trajectory invariant to arbitrary independent rescaling of the training series, yielding scale-invariant training (ScaleIn). Notably, existing TSFMs use both objectives, with neither consistent reporting nor a common convention on how to compute training loss. We isolate the convergence disparity induced by ScaleCon and its correction under ScaleIn in controlled studies on synthetic and real data. In pretraining across four TSFM architectures, ScaleIn lowers MASE in all 24 architecture-benchmark comparisons, with average reductions across TSFMs of 18.8% on GIFT-Eval and 21.9% on the M-competitions. The gains extend to supervised neural forecasting, where it lowers MASE in 16 of 20 matched settings. Most existing time series forecasting pipelines can adopt ScaleIn with a one-line code change.
Figures & tables
| Model | Loss Function | Explicit | Inferred | Code | |
|---|---|---|---|---|---|
| TimeGPT-1 | garza2023timegpt1 | – | – | – | – |
| LagLLama | rasul2023lag | Student-t NLL | – | ScaleCon ¶ | ScaleCon ¶ |
| Time-LLM | jin2023time | MSE | – | ScaleCon | ScaleCon |
| Moirai | salesforce2023moirai | Mixture NLL | – | ScaleCon ¶ | ScaleCon ¶ |
| MOMENT | goswami2024moment | MSE | – | ScaleCon | ScaleCon ‡ |
| TimesFM | das2024timesfm | MSE | – | ScaleCon | ScaleCon ∗‡ |
| Codebase | Reference | ScaleCon | Invariant ( ScaleIn ) | Fixed/User-specified |
|---|---|---|---|---|
| GluonTS | alexandrov2020gluonts | NLL, CRPS, EnergyScore | — | Fixed |
| PyTorch Forecasting | beitner2020pytorch_forecasting | MAE, RMSE, MAPE, SMAPE, QuantileLoss, NormalDist, LogNormal, NegBinomial, Beta, MQF2 | — | Fixed |
| TSLib | wu2023tslib | — | MSE, MAPE, MASE, SMAPE | Fixed |
| NeuralForecast | olivares2022neuralforecast | NLL (Normal, StudentT, NegBinomial, Poisson, Tweedie) | MAE, MSE, RMSE, MAPE, SMAPE, MASE, QuantileLoss, MQLoss | Fixed |
| Chronos2 | Moirai-2.0 | TimesFM2.5 | PatchTST | Reference | |||||||||
| Benchmark | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | S.Naive | ||||
| MASE | |||||||||||||
| M1 | 1.879 | 2.388 | -21.3 | 2.013 | 2.123 | -5.2 | 2.731 | 2.863 | -4.6 | 2.091 | 3.859 | -45.8 | 2.117 |
| M3 | 1.608 | 1.963 | -18.1 | 1.621 | 1.800 | -9.9 | 2.233 | 2.573 | -13.2 | 2.041 | 4.840 | -57.8 | 1.764 |
| M4 | 1.845 | 2.294 | -19.6 | 1.879 | 1.986 | -5.4 | 2.661 | 3.208 | -17.0 | 2.235 | 4.555 | -50.9 | 2.057 |
| Tourism | 2.350 | 2.816 | -16.5 | 2.316 | 2.564 | -9.7 | 3.820 | 4.108 | -7.0 | 2.509 | 3.123 | -19.7 | 2.412 |
| Neural Forecast Baselines | Reference | |||||||||||||
| Benchmark | Metric | N-HiTS | N-BEATS | PatchTST | TCN | LSTM | S.Naive | ETS | ARIMA | |||||
| Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | |||||
| M1 | MASE | 1.826 | 1.822 | 1.911 | 1.941 | 1.998 | ||||||||
| WQL | 0.098 | 0.098 | 0.106 | 0.106 | 0.100 | 0.105 | ||||||||
| M3 | MASE | 1.640 | 1.677 | 2.802 | 1.896 | 2.133 | ||||||||
| WQL | 0.090 | 0.089 | 0.116 | 0.099 | 0.107 | |||||||||
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Frequency | Seasonality | Horizon | Series | Min Length | Max Length | % Erratic | |
|---|---|---|---|---|---|---|---|
| M1 | Monthly | 12 | 18 | 617 | 48 | 150 | 0 |
| Quarterly | 4 | 8 | 203 | 18 | 114 | 0 | |
| Yearly | 1 | 6 | 181 | 15 | 58 | 0 | |
| M3 | Other | 4 | 8 | 174 | 71 | 104 | 0 |
| Monthly | 12 | 18 | 1428 | 66 | 144 | 2 | |
| Quarterly | 4 | 8 | 756 | 24 | 72 | 1 |
| Dataset | Source | Domain | Frequency | # Series | Avg | Min | Max | # Obs |
| Jena Weather | Autoformer [ wu2021autoformer ] | Nature | 10T | 1 | 52,704 | 52,704 | 52,704 | 52,704 |
| Jena Weather | Autoformer [ wu2021autoformer ] | Nature | H | 1 | 8,784 | 8,784 | 8,784 | 8,784 |
| Jena Weather | Autoformer [ wu2021autoformer ] | Nature | D | 1 | 366 | 366 | 366 | 366 |
| BizITObs - Application | AutoMixer [ palaskar2024automixer ] | Web/CloudOps | 10S | 1 | 8,834 | 8,834 | 8,834 | 8,834 |
| BizITObs - Service | AutoMixer [ palaskar2024automixer ] | Web/CloudOps | 10S | 21 | 8,835 | 8,835 | 8,835 | 185,535 |
| BizITObs - L2C | AutoMixer [ palaskar2024automixer ] | Web/CloudOps | 5T | 1 | 31,968 | 31,968 | 31,968 | 31,968 |
| Chronos2 | Moirai-2.0 | TimesFM2.5 | PatchTST | Reference | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Freq. | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | S.Naive | ||||
| M1 | Y | 4.124 | 6.492 | -36.5 | 4.894 | 5.175 | -5.4 | 8.067 | 7.765 | +3.9 | 4.707 | 7.600 | -38.1 | 4.894 |
| Q | 1.827 | 1.997 | -8.5 | 1.814 | 1.885 | -3.8 | 1.943 | 2.394 | -18.8 | 1.963 | 3.625 | -45.9 | 2.078 | |
| M | 1.238 | 1.312 | -5.6 | 1.234 | 1.306 | -5.5 | 1.425 | 1.579 | -9.8 | 1.366 | 2.838 | -51.9 | 1.314 | |
| M3 | Y | 3.070 | 4.085 | -24.8 | 3.119 | 3.643 | -14.4 | 4.880 | 5.309 | -8.1 | 3.889 | 9.529 | -59.2 | 3.154 |
| Q | 1.258 | 1.461 | -13.9 | 1.255 | 1.295 | -3.1 | 1.476 | 1.846 | -20.1 | 1.590 | 3.867 | -58.9 | 1.425 | |
| Chronos2 | Moirai-2.0 | TimesFM2.5 | PatchTST | Reference | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Benchmark | Freq. | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | Sc.In | Sc.Con | S.Naive | ||||
| M1 | Y | 0.110 | 0.126 | -12.8 | 0.154 | 0.157 | -2.1 | 0.248 | 0.287 | -13.6 | 0.196 | 0.307 | -36.1 | 0.184 |
| Q | 0.065 | 0.076 | -13.7 | 0.094 | 0.095 | -1.2 | 0.115 | 0.142 | -18.9 | 0.118 | 0.148 | -20.2 | 0.117 | |
| M | 0.162 | 0.131 | +23.5 | 0.207 | 0.188 | +10.0 | 0.227 | 0.261 | -12.9 | 0.272 | 0.322 | -15.7 | 0.150 | |
| M3 | Y | 0.101 | 0.112 | -10.1 | 0.132 | 0.143 | -7.3 | 0.218 | 0.221 | -1.3 | 0.170 | 0.283 | -40.0 | 0.118 |
| Q | 0.063 | 0.066 | -4.8 | 0.079 | 0.080 | -0.5 | 0.104 | 0.126 | -17.7 | 0.110 | 0.202 | -45.4 | 0.082 | |
| Frequency | Horizon | Context | Eligible series |
|---|---|---|---|
| Yearly | 8 | 16 | 7,841 |
| Quarterly | 8 | 16 | 23,239 |
| Monthly | 24 | 48 | 34,183 |
| Weekly | 13 | 26 | 359 |
| Daily | 14 | 28 | 4,227 |
| Hourly | 48 | 96 | 414 |
| Setting | Value |
|---|---|
| Optimizer | Adam |
| Training updates | 20,000 |
| Series batch size | 32 |
| Sampled windows per batch | 512 |
| Learning rate | |
| Learning-rate decays | 3, evenly spaced |
| Architecture | Parameters |
|---|---|
| NHITS | Three identity stacks, one block per stack, two hidden layers of 512 units per block, pooling kernels , frequency downsampling , ReLU activation, and linear interpolation. |
| NBEATS | Identity, trend, and seasonality stacks, one block per stack, two hidden layers of 512 units per block, ReLU activation, and no shared block weights. |
| PatchTST | Three encoder layers, 16 attention heads, hidden width 128, feed-forward width 256, patch length 16, stride 8, dropout 0.2, GELU activation, and internal RevIN disabled. |
| TCN | Kernel size 2, dilations , encoder width 128, context width 10, and a two-layer decoder of width 128. |
| LSTM | Two encoder layers of width 128, no encoder dropout, and a two-layer decoder of width 128. Forecasts use the direct decoder rather than recursive generation. |
| Architecture | Benchmark | Frequency | ScaleIn | ScaleCon | S.Naive | ETS | ARIMA |
|---|---|---|---|---|---|---|---|
| N-HiTS | M1 | Y | 2.368 | 2.461 | 3.046 | 2.476 | 2.632 |
| Q | 1.859 | 1.857 | 2.527 | 1.899 | 1.966 | ||
| M | 1.251 | 1.248 | 1.859 | 1.338 | 1.333 | ||
| M3 | Y | 2.972 | 2.934 | 3.333 | 2.384 | 3.030 | |
| Q | 1.121 | 1.127 | 1.457 | 1.115 | 1.130 | ||
| M | 0.863 | 0.860 | 1.243 | 0.932 | 0.901 |
| Architecture | Benchmark | Frequency | ScaleIn | ScaleCon | S.Naive | ETS | ARIMA |
|---|---|---|---|---|---|---|---|
| N-HiTS | M1 | Y | 0.119 | 0.117 | 0.219 | 0.144 | 0.204 |
| Q | 0.096 | 0.096 | 0.139 | 0.094 | 0.106 | ||
| M | 0.083 | 0.081 | 0.144 | 0.088 | 0.086 | ||
| M3 | Y | 0.105 | 0.107 | 0.105 | 0.093 | 0.100 | |
| Q | 0.081 | 0.082 | 0.092 | 0.079 | 0.081 | ||
| M | 0.083 | 0.083 | 0.111 | 0.089 | 0.088 |