Authors: Hao Wang, Licheng Pan, Zhichao Chen, Xu Chen, Qingyang Dai, Lei Wang, Haoxuan Li, Zhouchen Lin
Organizations: Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Gaoling School of Artificial Intelligence, Renmin University of China · Department of Control Science and Engineering, Zhejiang University · Center for Data Science, Peking University · Institute for Artificial Intelligence, Peking University · Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China
Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https://github.com/Master-PLC/Time-o1.
Figures & tables
Figure 1 : The autocorrelations and volumes in the label sequence Y and latent components Z .
Models
Time-o1
Fredformer
iTransformer
FreTS
TimesNet
MICN
TiDE
DLinear
FEDformer
Autoformer
Transformer
(Ours)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2022)
(2021)
(2017)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
0.380
0.393
0.387
0.398
0.411
0.414
0.414
0.421
0.438
0.430
0.396
0.421
0.413
0.407
0.403
0.407
0.442
0.457
0.526
0.491
0.799
0.648
ETTm2
0.272
0.317
0.280
0.324
0.295
0.336
0.316
0.365
0.302
0.334
0.308
0.364
0.286
0.328
0.342
0.392
0.308
0.354
0.315
0.358
1.662
0.917
ETTh1
0.431
0.429
0.447
0.434
0.452
0.448
0.489
0.474
0.472
0.463
0.533
0.519
0.448
0.435
0.456
0.453
0.447
0.470
0.477
0.483
0.983
0.774
ETTh2
0.359
0.388
0.377
0.402
0.386
0.407
0.524
0.496
0.409
0.420
0.620
0.546
0.378
0.401
0.529
0.499
0.452
0.461
0.448
0.460
2.688
1.291
Table 1 : Long-term forecasting performance.
Figure 2 : The visualization of label and forecast sequences generated by models trained with TMSE versus Time-o1. In both (a) and (b), the left panels display the time-domain sequences ( Y and Y^ ), while the right panels illustrate their corresponding latent components ( Z and Z^ ).
Loss
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DPTA
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Fredformer
ETTm1
0.379
0.393
0.384
0.394
0.389
0.400
0.389
0.400
0.397
0.402
0.396
0.402
0.387
0.398
ETTh1
0.431
0.429
0.438
0.434
0.452
0.443
0.453
0.442
0.460
0.449
0.460
0.449
0.447
0.434
ECL
0.178
0.270
0.179
0.272
0.190
0.282
0.187
0.280
0.206
0.298
0.202
0.294
0.191
0.284
Weather
0.255
0.276
0.256
0.277
0.257
0.279
0.258
0.280
0.261
0.280
0.260
0.280
0.261
0.282
iTransformer
ETTm1
0.395
0.401
0.405
0.405
0.413
0.416
0.407
0.412
0.417
0.415
0.416
0.415
0.411
0.414
Table 2 : Comparable results with other loss functions for time-series forecast.
Model
Decorrelation
Reduction
Data
T=96
T=192
T=336
T=720
Avg
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
DF
✗
✗
ETTm1
0.326
0.361
0.365
0.382
0.396
0.404
0.459
0.444
0.387
0.398
ETTh1
0.377
0.396
0.437
0.425
0.486
0.449
0.488
0.467
0.447
0.434
ECL
0.150
0.242
0.168
0.259
0.182
0.274
0.214
0.304
0.179
0.270
Weather
0.174
0.228
0.213
0.266
0.270
0.316
0.337
0.362
0.249
0.293
Time-o1 †
✗
✓
ETTm1
0.338
0.366
0.369
0.383
0.397
0.403
0.458
0.441
0.391
0.398
Table 3 : Ablation study results.
ECL
Weather
Transformation
MSE
Δ
MAE
Δ
MSE
Δ
MAE
Δ
None
0.179
-
0.270
-
0.249
-
0.293
-
RPCA
0.171
4.31% ↓
0.261
3.16% ↓
0.244
1.78% ↓
0.286
2.38% ↓
SVD
0.175
2.24% ↓
0.264
2.18% ↓
0.248
0.34% ↓
0.290
0.93% ↓
FA
0.175
2.35% ↓
0.265
1.82% ↓
0.245
1.35% ↓
0.287
1.97% ↓
Ours
0.170
4.86% ↓
0.260
3.57% ↓
0.241
2.94% ↓
0.280
4.28% ↓
Table 4 : Varying transformations results.
Figure 3 : Improvement of Time-o1 applied to different forecast models, shown with colored bars for means over forecast lengths (96, 192, 336, 720) and error bars for 50% confidence intervals.
Table 8
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4 : The label autocorrelation in the original label sequence and the extracted components. The datasets are ETTh1, ETTh2, ETTm1, and Weather from left to right. The forecast length is set to 96.
Figure 5 : Running cost for projection matrix calculation (left panel with varying number of samples, right panel with varying prediction length) and sequence transformation (left panel for forward pass, right panel for backward pass, with average and shaded areas for 95% confidence intervals).
Models
Time-o1
Fredformer
iTransformer
FreTS
TimesNet
MICN
TiDE
DLinear
FEDformer
Autoformer
Transformer
(Ours)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2022)
(2021)
(2017)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.321
0.357
0.326
0.361
0.338
0.372
0.342
0.375
0.368
0.394
0.319
0.366
0.353
0.374
0.346
0.373
0.401
0.434
0.485
0.468
0.503
0.482
192
0.360
0.378
0.365
0.382
0.382
0.396
0.385
0.400
0.406
0.409
0.364
0.395
0.391
0.393
0.380
0.390
0.415
0.446
0.504
0.482
0.807
0.664
336
0.389
0.400
0.396
0.404
0.427
0.424
0.416
0.421
0.454
0.444
0.395
0.425
0.423
0.414
0.413
0.414
0.432
0.450
0.520
0.489
0.847
0.678
720
0.447
0.435
0.459
0.444
0.496
0.463
0.513
0.489
0.527
0.474
0.505
0.499
0.486
0.448
0.472
0.450
0.522
0.500
0.594
0.523
1.037
0.771
Appendix
Table 7 : The comprehensive results on the long-term forecasting task.
Models
Time-o1
Fredformer
iTransformer
FreTS
MICN
DLinear
Fedformer
(Ours)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
Metric
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
Yearly
13.485
3.010
0.791
13.509
3.028
0.794
13.797
3.143
0.818
13.576
3.068
0.801
14.594
3.392
0.873
14.307
3.094
0.827
13.648
3.089
0.806
Quarterly
10.105
1.180
0.889
10.140
1.185
0.893
10.503
1.248
0.932
10.361
1.223
0.916
11.417
1.385
1.023
10.500
1.237
0.928
10.612
1.246
0.936
Monthly
12.649
0.930
0.875
12.696
0.931
0.878
13.227
1.013
0.935
13.088
0.990
0.919
13.834
1.080
0.987
13.362
1.007
0.937
14.181
1.105
1.011
Others
4.852
3.274
1.027
4.848
3.230
1.019
5.101
3.419
1.076
5.563
3.71
1.17
6.137
4.201
1.308
5.12
3.649
1.114
4.823
3.243
1.019
Appendix
Table 8 : The comprehensive results on the short-term forecasting task.
Figure 6 : The forecast sequences generated with DF and Time-o1. The forecast length is set to 336 and the experiment is conducted on ETTm2.
Figure 7 : The forecast sequences generated with DF and Time-o1. The forecast length is set to 192 and the experiment is conducted on ECL.
Figure 8 : Performance of different forecast models with and without Time-o1. The forecast errors are averaged over forecast lengths and the error bars represent 50% confidence intervals.
Trans
PCA
RPCA
SVD
FA
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ECL
96
0.1449
0.2348
0.1450
0.2349
0.1450
0.2350
0.1478
0.2385
0.1500
0.2415
192
0.1592
0.2487
0.1594
0.2487
0.1595
0.2490
0.1619
0.2517
0.1681
0.2591
336
0.1731
0.2645
0.1732
0.2646
0.1730
0.2643
0.1789
0.2711
0.1823
0.2744
720
0.2033
0.2920
0.2066
0.2960
0.2214
0.3066
0.2095
0.2975
0.2145
0.3035
Avg
0.1701
0.2600
0.1710
0.2611
0.1747
0.2637
0.1745
0.2647
0.1787
0.2696
Appendix
Table 9 : Varying transformation results.
Figure 9 : Time-o1 improves Fredformer performance given a wide range of transformed loss strength α . These experiments are conducted on ETTh1 (a), ETTh2 (b), ETTm1 (c), ETTm2 (d), ECL (e), Weather (f) datasets. Different columns correspond to different forecast lengths (from left to right: 96, 192, 336, 720, and their average with shaded areas being 15% confidence intervals).
Figure 10 : Time-o1 improves Fredformer performance given a wide range of rank ratio γ . These experiments are conducted on ETTh1 (a), ETTh2 (b), ETTm1 (c), ETTm2 (d), ECL (e), and Weather (f) datasets. Different columns correspond to different forecast lengths (from left to right: 96, 192, 336, 720, and their average with shaded areas being 15% confidence intervals).
Loss
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DPTA
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Forecast model: FredFormer
ETTm1
96
0.321
0.357
0.326
0.355
0.335
0.368
0.337
0.367
0.332
0.363
0.332
0.364
0.326
0.361
192
0.360
0.378
0.363
0.380
0.366
0.384
0.364
0.384
0.370
0.386
0.370
0.386
0.365
0.382
336
0.389
0.400
0.392
0.400
0.399
0.408
0.397
0.406
0.406
0.409
0.409
0.410
0.396
0.404
720
0.447
0.435
0.455
0.440
0.456
0.441
0.457
0.443
0.478
0.450
0.476
0.448
0.459
0.444
Appendix
Table 10 : Comparable results with different loss functions.
Models
Time-o1
iTransformer
Time-o1
PatchTST
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Input sequence length
96
96
0.163
0.202
0.171
0.210
0.175
0.213
0.200
0.244
192
0.214
0.248
0.246
0.278
0.224
0.257
0.229
0.263
336
0.274
0.294
0.296
0.313
0.276
0.296
0.287
0.303
720
0.351
0.344
0.362
0.353
0.353
0.346
0.363
0.353
Avg
0.250
0.272
0.269
0.289
0.257
0.278
0.270
0.291
Appendix
Table 11 : Varying input sequence length results on the Weather dataset.
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation proves biased when the label sequence exhibits autocorrelation. In this paper, we propose DistDF, which achieves alignment by minimizing a distributional discrepancy between the conditional distributions of forecast and label sequences. Since such conditional discrepancies are difficult to estimate from finite time-series observations, we introduce a joint-distribution Wasserstein discrepancy for time-series forecasting, which provably upper bounds the conditional discrepancy of interest. The proposed discrepancy is tractable, differentiable, and readily compatible with gradient-based optimization. Extensive experiments show that DistDF improves diverse forecasting models and achieves leading performance. Code is available at https://anonymous.4open.science/r/DistDF-F66B.
Hao Wang, Licheng Pan, Yuan Lu +7
Xiaohongshu Inc. · College of Control Science and Technology, Zhejiang University · College of Computer Science and Technology, Zhejiang University +4
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting models primarily adhere to the Direct Forecast (DF) paradigm, generating multi-step forecasts independently and disregarding label autocorrelation over time. In this work, we demonstrate that the learning objective of DF is biased in the presence of label autocorrelation. To address this issue, we propose the Frequency-enhanced Direct Forecast (FreDF), which mitigates label autocorrelation by learning to forecast in the frequency domain, thereby reducing estimation bias. Our experiments show that FreDF significantly outperforms existing state-of-the-art methods and is compatible with a variety of forecast models. Code is available at https://github.com/Master-PLC/FreDF.
Hao Wang, Licheng Pan, Zhichao Chen +6
Department of Control Science and Engineering, Zhejiang University · School of Automation, Central South University · Trust and Safety Team, TikTok Sydney, ByteDance Inc. +3
Autocorrelation is a common property of time-series, where each observation is dependent on its predecessors. In deep time-series forecasting, it raises two central challenges: (1) designing backbone architectures to model autocorrelation in history sequences, and (2) devising loss functions to model autocorrelation in label sequences. Recent studies have made strides in tackling these challenges, but a systematic survey examining both aspects remains lacking. To bridge this gap, this paper reviews deep time-series forecasting from an autocorrelation modeling perspective, offering two contributions beyond existing surveys. First, it introduces a taxonomy that jointly covers both backbone architectures and loss functions, whereas prior surveys provide limited coverage of the latter. Second, it analyzes the motivations and insights underlying the surveyed literature from a unified autocorrelation perspective, providing a holistic overview of the field's evolution. Additional resources and details are available at https://github.com/Master-PLC/Awesome-TSF-Papers.