From HL to H+L-1 Parameters: A Hankel-Toeplitz Forecaster for Long-Term Time Series Forecasting
Organizations: Indiana University Bloomington, IN, USA
Abstract
Linear forecasters have shown competitive accuracy against Transformer-based models in long-term time series forecasting. We study how classical stationary prediction theory can guide parameter sharing for more compact linear forecasters. For centered second-order stationary processes with nonsingular history covariance, the minimum-MSE finite-window linear predictor factors into a Hankel cross-covariance matrix and an inverse Toeplitz covariance matrix. Shared lags and scale cancellation specify this predictor using autocorrelations for lookback and horizon . Building on the innovations representation, our Hankel-Toeplitz Forecaster (HTF) learns one impulse response that defines both an inverse filter and a forecast map. We characterize the finite-history correction and, under summability assumptions, bound the excess risk of truncating the true filters. HTF uses trainable coefficients while allowing a full-rank forecasting matrix. Across seven benchmarks at , its horizon-averaged MSE is within 1.2% of Dense Linear on each dataset with 75-229 times fewer trainable parameters.
Figures & tables
| Dataset | Dense Linear | DLinear | FITS | SparseTSF | DiPE-Linear | TimeBase | HTF (Ours) | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MSE | K | MSE | K | MSE | K | MSE | K | MSE | K | MSE | K | MSE | K | |
| ETTh1 | 0.4088 | 112.90 | 0.4139 | 226.46 | 0.4112 | 40.35 | 0.4044 | 0.22 | 0.4053 | 1.87 | 0.4093 | 0.19 | 0.4075 | 0.67 |
| ETTh2 | 0.3429 | 112.90 | 0.3436 | 226.46 | 0.3388 | 40.35 | 0.3475 | 0.22 | 0.3441 | 5.64 | 0.3535 | 0.19 | 0.3390 | 0.67 |
| ETTm1 | 0.3614 | 112.90 | 0.3624 | 226.46 | 0.3606 | 16.99 | 0.3834 | 0.11 | 0.3635 | 1.87 | 0.3724 | 0.06 | 0.3595 | 0.67 |
| ETTm2 | 0.2581 | 112.90 | 0.2581 | 226.46 | 0.2568 | 17.62 | 0.2842 | 0.11 | 0.2582 | 1.87 | 0.2940 | 0.06 | 0.2572 | 0.67 |
| Weather | 0.2481 | 112.90 | 0.2477 | 226.46 | 0.2483 | 7.80 | 0.2769 | 0.15 | 0.2249 | 7.58 | 0.2935 | 0.04 | 0.2491 | 0.67 |
| Model | Coefficients | Difference |
|---|---|---|
| All 28 settings, | ||
| HTF (Ours) | ||
| Ridge (zero bias) | ||
| Ridge-projected, rank 1 | ||
| Ridge-projected, rank 2 | ||
| Ridge-projected, rank 8 | ||