FreDF: Learning to Forecast in the Frequency Domain
Authors: Hao Wang, Licheng Pan, Zhichao Chen, Degui Yang, Sen Zhang, Yifei Yang, Xinggao Liu, Haoxuan Li, +1 more
Organizations: Department of Control Science and Engineering, Zhejiang University · School of Automation, Central South University · Trust and Safety Team, TikTok Sydney, ByteDance Inc. · Department of Computer Science and Engineering, Shanghai Jiao Tong University · Center for Data Science, Peking University · Generative AI Lab, College of Computing and Data Science, Nanyang Technological University
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting models primarily adhere to the Direct Forecast (DF) paradigm, generating multi-step forecasts independently and disregarding label autocorrelation over time. In this work, we demonstrate that the learning objective of DF is biased in the presence of label autocorrelation. To address this issue, we propose the Frequency-enhanced Direct Forecast (FreDF), which mitigates label autocorrelation by learning to forecast in the frequency domain, thereby reducing estimation bias. Our experiments show that FreDF significantly outperforms existing state-of-the-art methods and is compatible with a variety of forecast models. Code is available at https://github.com/Master-PLC/FreDF.
Figures & tables
Figure 1: Visualizing label autocorrelation in time series forecasting. (a) shows the generation process of time series with dependencies depicted as arrows. (b) shows the label correlation in the time domain, where each element ρi,j indicates the partial correlation between Yi and Yj given L . (c-d) shows the label correlation in the frequency domain, where each element ρi,j indicates the partial correlation between Fi and Fj given L , shown with the real (c) and imaginary part (d). Due to the symmetry inherent in FFT, the forecast length in the frequency domain is halved.
Figure 2: The workflow of FreDF. Key operations in the time and frequency domains are highlighted in red and blue, respectively.
Models
FreDF
iTransformer
FreTS
TimesNet
MICN
TiDE
DLinear
FEDformer
Autoformer
Transformer
TCN
(Ours)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2022)
(2021)
(2017)
(2017)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
0.392
0.399
0.415
0.416
0.407
0.415
0.413
0.418
0.399
0.423
0.419
0.419
0.404
0.407
0.440
0.451
0.596
0.517
0.943
0.733
0.891
0.632
ETTm2
0.278
0.319
0.294
0.335
0.335
0.379
0.297
0.332
0.300
0.356
0.358
0.404
0.344
0.396
0.302
0.348
0.326
0.366
1.322
0.814
3.411
1.432
ETTh1
0.437
0.435
0.449
0.447
0.488
0.474
0.478
0.466
0.525
0.515
0.628
0.574
0.462
0.458
0.441
0.457
0.476
0.477
0.993
0.788
0.763
0.636
ETTh2
0.371
0.396
0.390
0.410
0.550
0.515
0.413
0.426
0.624
0.549
0.611
0.550
0.558
0.516
0.430
0.447
0.478
0.483
3.296
1.419
3.325
1.445
Table 1: Long-term forecasting performance.
Figure 3: Visualization of forecast sequence generated with and without FreDF in the time (a-b) and frequency (c-d) domains, using the iTransformer as the backbone model.
Model
L(tmp)
L(feq)
Data
T=96
T=192
T=336
T=720
Avg
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
DF
✓
✗
ETTm1
0.346
0.379
0.391
0.400
0.426
0.422
0.493
0.460
0.414
0.415
ETTh1
0.390
0.409
0.442
0.440
0.479
0.457
0.483
0.479
0.449
0.446
ECL
0.147
0.239
0.166
0.258
0.178
0.271
0.209
0.298
0.175
0.266
Weather
0.201
0.246
0.250
0.282
0.302
0.317
0.370
0.361
0.280
0.302
FreDF †
✗
✓
ETTm1
0.324
0.361
0.374
0.387
0.403
0.405
0.468
0.443
0.392
0.399
Table 2: Ablation study results.
Model
ETTh1
ETTm1
ECL
MSE
Δ
MAE
Δ
MSE
Δ
MAE
Δ
MSE
Δ
MAE
Δ
iTransformer
0.449
-
0.447
-
0.415
-
0.416
-
0.176
-
0.267
-
+ FreDF-T
0.437
↓ 2.63%
0.435
↓ 2.62%
0.392
↓ 5.49%
0.399
↓ 4.01%
0.170
↓ 3.41%
0.259
↓ 2.77%
+ FreDF-D
0.445
↓ 0.92%
0.440
↓ 1.42%
0.395
↓ 4.77%
0.398
↓ 4.33%
0.171
↓ 2.51%
0.260
↓ 2.52%
+ FreDF-2
0.432
↓ 3.94%
0.431
↓ 3.57%
0.392
↓ 5.60%
0.399
↓ 4.05%
0.166
↓ 5.32%
0.256
↓ 4.20%
Table 3: Varying FFT implementation results.
Figure 4: Benefit of incorporating FreDF in varying models, shown with colored bars for means over forecast lengths (96, 192, 336, 720) and error bars for 99.9% confidence intervals.
Figure 5: Varying projection bases results, shown with colored bars for means over forecast lengths (96, 192, 336, 720) and error bars for 99.9% confidence intervals.
Figure 6: Varying strength of frequency loss ( α ) results, shown with colored lines for T=192, 336.
Figure 7: Learning curve on ETTm1 dataset.
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Visualization of partial correlation and DML approach for partial correlation quantification. (a) The correlation graph where the pseudo correlation is caused by the fork structure T←X→Y . (b) The implementation of DML, where β is the identified strength of the partial correlation T→Y . (c) The partial correlation identified by DML.
Figure 9: More comprehensive visualizations of label autocorrelation in different domains and datasets, with columns representing different datasets: Traffic, ETTh1, and ECL, from left to right. Panels (a-c) show the label correlation in the time domain, where each element ρi,j indicates the partial correlation between Yi and Yj given L . Panels (d-i) show the label correlation in the frequency domain, where each element ρi,j indicates the partial correlation between Fi and Fj given L , shown with the real (d-f) and imaginary part (g-i). Due to the symmetry inherent in FFT, the forecast length in the frequency domain is halved.
Figure 10: More comprehensive visualizations of label autocorrelation in different domains and label lengths, with columns representing label lengths H=48, 96, 192, 336 from left to right. Panels (a-d) show the label correlation in the time domain, where each element ρi,j indicates the partial correlation between Yi and Yj given L . Panels (e-l) show the label correlation in the frequency domain, where each element ρi,j indicates the partial correlation between Fi and Fj given L , shown with the real (e-h) and imaginary part (i-l).
Figure 11: The label sequences (black lines) and forecast sequences generated by DF (blue lines) and FreDF (red lines). The forecast model used is iTransformer, with experiments conducted on selected snapshots characterized by periodicity (a) and trend (b).
Dataset
D
Forecast length
Train / validation / test
Frequency
Domain
ETTh1
7
96, 192, 336, 720
8545/2881/2881
Hourly
Health
ETTh2
7
96, 192, 336, 720
8545/2881/2881
Hourly
Health
ETTm1
7
96, 192, 336, 720
34465/11521/11521
15min
Health
ETTm2
7
96, 192, 336, 720
34465/11521/11521
15min
Health
Weather
21
96, 192, 336, 720
36792/5271/10540
10min
Weather
ECL
321
96, 192, 336, 720
18317/2633/5261
Hourly
Electricity
Appendix
Table 4: Dataset description.
Models
FreDF
iTransformer
FreTS
TimesNet
MICN
TiDE
DLinear
FEDformer
Autoformer
Transformer
TCN
(Ours)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2022)
(2021)
(2017)
(2017)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.324
0.362
0.346
0.379
0.339
0.374
0.338
0.379
0.318
0.366
0.364
0.387
0.345
0.372
0.389
0.427
0.468
0.463
0.591
0.549
0.887
0.613
192
0.373
0.385
0.392
0.400
0.382
0.397
0.389
0.400
0.364
0.396
0.398
0.404
0.381
0.390
0.402
0.431
0.573
0.509
0.704
0.629
0.877
0.626
336
0.402
0.404
0.427
0.422
0.421
0.426
0.429
0.428
0.398
0.428
0.428
0.425
0.414
0.414
0.438
0.451
0.596
0.527
1.171
0.861
0.890
0.636
720
0.469
0.444
0.494
0.461
0.485
0.462
0.495
0.464
0.514
0.501
0.487
0.461
0.473
0.451
0.529
0.498
0.749
0.569
1.307
0.893
0.911
0.653
Appendix
Table 5: The comprehensive results on the long-term forecasting task.
Models
FreDF
FreTS
iTransformer
MICN
DLinear
Fedformer
Autoformer
(Ours)
(2023)
(2024)
(2023)
(2023)
(2023)
(2023)
Metric
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
Yearly
13.556
3.046
0.798
13.576
3.068
0.801
13.797
3.143
0.818
14.594
3.392
0.873
14.307
3.094
0.827
13.648
3.089
0.806
18.477
4.26
1.101
Quarterly
10.374
1.229
0.919
10.361
1.223
0.916
10.503
1.248
0.932
11.417
1.385
1.023
10.500
1.237
0.928
10.612
1.246
0.936
14.254
1.829
1.314
Monthly
12.999
0.983
0.913
13.088
0.99
0.919
13.227
1.013
0.935
13.834
1.080
0.987
13.362
1.007
0.937
14.181
1.105
1.011
18.421
1.616
1.398
Others
5.294
3.614
1.127
5.563
3.71
1.17
5.101
3.419
1.076
6.137
4.201
1.308
5.12
3.649
1.114
4.823
3.243
1.019
6.772
4.963
1.495
Appendix
Table 6: The comprehensive results on the short-term forecasting task.
Models
FreDF
iTransformer
FreTS
TimesNet
MICN
TiDE
DLinear
FEDformer
Autoformer
(Ours)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2022)
(2021)
pmiss
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
0.125
0.00153
0.02790
0.00213
0.03307
0.01102
0.07843
0.01152
0.07267
0.00236
0.03371
0.45052
0.45514
0.00148
0.02380
0.68262
0.38111
0.37654
0.35378
0.25
0.00287
0.03801
0.00402
0.04434
0.01089
0.07753
0.01245
0.07946
0.00284
0.03691
0.41777
0.45884
0.00154
0.02351
0.68235
0.38116
0.37059
0.35261
0.375
0.00256
0.03669
0.00458
0.04663
0.01100
0.07812
0.01407
0.08673
0.00323
0.03900
0.62935
0.55570
0.00175
0.02385
0.68191
0.38105
0.37877
0.36093
0.5
0.00152
0.02739
0.00363
0.04359
0.01102
0.07818
0.01676
0.09610
0.00352
0.04028
0.29342
0.39320
0.00192
0.02219
0.68119
0.38085
0.38052
0.36462
Appendix
Table 7: The comprehensive results on the missing data imputation task.
Figure 12: The forecast sequences generated with DF and FreDF. The forecast length is set to 336 and the experiment is conducted on a snapshot of ETTm2.
Figure 13: The spectrum of forecast sequences generated with DF and FreDF. The forecast length is set to 336 and the experiment is conducted on a snapshot of ETTm2. Only the first 24 frequencies of the spectrum are presented.
Figure 14: The forecast sequences generated with DF and FreDF. The forecast length is set to 336 and the experiment is conducted on another snapshot of ETTm2.
Figure 15: The spectrum of forecast sequences generated with DF and FreDF. The forecast length is set to 336 and the experiment is conducted on another snapshot of ETTm2. Only the first 24 frequencies of the spectrum are presented.
Figure 16: Performance of different forecast models with and without FreDF. The forecast errors are averaged over forecast lengths and the error bars represent 95% confidence intervals.
Figure 17: FreDF improves iTransformer performance given a wide range of frequency loss weight α . These experiments are conducted on ETTh1 (a), ETTh2 (b), ETTm1 (c), ETTm2 (d), ECL (e), Traffic (f) and Weather (g) datasets. Different columns correspond to different forecast lengths (from left to right: 96, 192, 336, 720, and their average with shaded areas being 50% confidence intervals).
Figure 18: FreDF improves Autoformer performance given a wide range of frequency loss weight α . These experiments are conducted on ETTh1 (a), ETTh2 (b), ETTm1 (c), ETTm2 (d), ECL (e), Traffic (f) and Weather (g) datasets. Different columns correspond to different forecast lengths (from left to right: 96, 192, 336, 720, and their average with shaded areas being 50% confidence intervals).
Figure 19: FreDF improves DLinear performance given a wide range of frequency loss weight α . These experiments are conducted on ETTh1 (a), ETTh2 (b), ETTm1 (c), ETTm2 (d), ECL (e), Traffic (f) and Weather (g) datasets. Different columns correspond to different forecast lengths (from left to right: 96, 192, 336, 720, and their average with shaded areas being 50% confidence intervals).
Dataset
ETTm1
ETTh1
Models
FreDF
Dilate
DPP
FreDF
Dilate
DPP
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
96
0.324
0.362
0.498
0.443
0.631
0.495
0.382
0.400
0.790
0.567
0.815
0.577
192
0.373
0.385
0.993
0.625
0.975
0.617
0.430
0.427
0.950
0.643
0.916
0.633
336
0.402
0.404
0.946
0.628
0.945
0.626
0.474
0.451
0.978
0.663
0.986
0.660
720
0.469
0.444
0.999
0.652
1.079
0.678
0.463
0.462
0.922
0.654
0.898
0.649
Appendix
Table 8: Comparable results with DTW-based loss.
Models
FreDF
iTransformer
FreDF
PatchTST
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Input sequence length
96
96
0.164
0.202
0.201
0.247
0.174
0.217
0.200
0.244
192
0.220
0.253
0.250
0.283
0.230
0.266
0.234
0.268
336
0.275
0.294
0.302
0.317
0.279
0.301
0.311
0.321
720
0.356
0.347
0.370
0.362
0.355
0.351
0.365
0.353
Avg
0.254
0.274
0.281
0.302
0.259
0.284
0.278
0.297
Appendix
Table 1: Varying input sequence length results on the Weather dataset.
Figure 1: Running time in the forward pass (left panel) and backward pass (right panel), shown with dashed lines for the average and shaded areas for 99.9% confidence intervals.
Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https://github.com/Master-PLC/Time-o1.
Hao Wang, Licheng Pan, Zhichao Chen +5
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Gaoling School of Artificial Intelligence, Renmin University of China +4
The design of learning objectives is central to training time-series forecasting models. Existing learning objectives such as mean squared error mostly treat each future step as an independent, equally weighted task, which leads to the following two challenges: (1) they overlook the label autocorrelation effect among future steps, leading to biased learning objectives; (2) they fail to set heterogeneous task weights for different forecasting tasks corresponding to varying future steps, limiting the forecasting performance. To fill this gap, we propose a novel quadratic-form weighted learning objective, addressing both issues simultaneously. Specifically, the off-diagonal elements of the weighting matrix account for the label autocorrelation effect, whereas the non-uniform diagonals are expected to match the preferred weights of the forecasting tasks with varying future steps. On this basis, we propose a Quadratic Direct Forecast (QDF) learning algorithm, which trains the forecast model using the adaptively updated quadratic-form weighting matrix. Experiments show that our QDF effectively improves the performance of various forecast models, achieving state-of-the-art results. Code is available at https://github.com/Master-PLC/QDF.
Hao Wang, Licheng Pan, Yuan Lu +7
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · College of Engineering, Purdue University +6
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation proves biased when the label sequence exhibits autocorrelation. In this paper, we propose DistDF, which achieves alignment by minimizing a distributional discrepancy between the conditional distributions of forecast and label sequences. Since such conditional discrepancies are difficult to estimate from finite time-series observations, we introduce a joint-distribution Wasserstein discrepancy for time-series forecasting, which provably upper bounds the conditional discrepancy of interest. The proposed discrepancy is tractable, differentiable, and readily compatible with gradient-based optimization. Extensive experiments show that DistDF improves diverse forecasting models and achieves leading performance. Code is available at https://anonymous.4open.science/r/DistDF-F66B.
Hao Wang, Licheng Pan, Yuan Lu +7
Xiaohongshu Inc. · College of Control Science and Technology, Zhejiang University · College of Computer Science and Technology, Zhejiang University +4