Organizations: Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · College of Engineering, Purdue University · School of Computing and Artificial Intelligence, Shanghai University of Finance and Economics · College of Computer Science and Technology, Zhejiang University · Squirrel AI · Center for Data Science, Peking University · Institute for Artificial Intelligence, Peking University · Pazhou Laboratory (Huangpu), Guangzhou, Guangdong, China
The design of learning objectives is central to training time-series forecasting models. Existing learning objectives such as mean squared error mostly treat each future step as an independent, equally weighted task, which leads to the following two challenges: (1) they overlook the label autocorrelation effect among future steps, leading to biased learning objectives; (2) they fail to set heterogeneous task weights for different forecasting tasks corresponding to varying future steps, limiting the forecasting performance. To fill this gap, we propose a novel quadratic-form weighted learning objective, addressing both issues simultaneously. Specifically, the off-diagonal elements of the weighting matrix account for the label autocorrelation effect, whereas the non-uniform diagonals are expected to match the preferred weights of the forecasting tasks with varying future steps. On this basis, we propose a Quadratic Direct Forecast (QDF) learning algorithm, which trains the forecast model using the adaptively updated quadratic-form weighting matrix. Experiments show that our QDF effectively improves the performance of various forecast models, achieving state-of-the-art results. Code is available at https://github.com/Master-PLC/QDF.
Figures & tables
Figure 1: Statistics of label components conditioned on X , with a forecast horizon of T=96 . (a) Partial correlation and conditional variance estimated from the raw label sequence Y , with colors indicating different X . (b) Partial correlation matrices of label components extracted by FreDF and Time-o1 ( Wang et al., 2025e ; Wang et al., 2025d ) . Calculation details are provided in Appendix A .
Figure 2
Models
QDF
TQNet
PDF
Fredformer
iTransformer
FreTS
TimesNet
MICN
TiDE
PatchTST
DLinear
(Ours)
(2025)
(2024)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2023)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
0.371
0.389
0.376
0.391
0.387
0.396
0.387
0.398
0.411
0.414
0.414
0.421
0.438
0.430
0.396
0.421
0.413
0.407
0.389
0.400
0.403
0.407
ETTm2
0.270
0.317
0.277
0.321
0.283
0.331
0.280
0.324
0.295
0.336
0.316
0.365
0.302
0.334
0.308
0.364
0.286
0.328
0.303
0.344
0.342
0.392
ETTh1
0.431
0.431
0.449
0.439
0.452
0.440
0.447
0.434
0.452
0.448
0.489
0.474
0.472
0.463
0.533
0.519
0.448
0.435
0.459
0.451
0.456
0.453
ETTh2
0.368
0.397
0.375
0.400
0.375
0.399
0.377
0.402
0.386
0.407
0.524
0.496
0.409
0.420
0.620
0.546
0.378
0.401
0.390
0.413
0.529
0.499
Table 1: Long-term forecasting performance.
Figure 2: The forecast sequence of DF (in blue) and QDF (in red), with historical length H=96 .
Loss
QDF
Time-o1
FreDF
Koopman
Soft-DTW
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
TQNet
ETTm1
0.371
0.389
0.372
0.390
0.375
0.390
0.595
0.499
0.387
0.394
0.376
0.391
ETTh1
0.431
0.431
0.437
0.432
0.432
0.432
0.451
0.442
0.453
0.438
0.449
0.439
ECL
0.165
0.257
0.167
0.257
0.168
0.257
0.166
0.258
0.623
0.524
0.175
0.265
Weather
0.242
0.268
0.245
0.269
0.244
0.268
0.282
0.306
0.255
0.276
0.246
0.270
PDF
ETTm1
0.381
0.394
0.386
0.399
0.387
0.400
0.587
0.485
0.396
0.404
0.387
0.396
Table 2: Comparable results with other objectives for time-series forecast.
Model
Hetero.
Auto.
Data
T=96
T=192
T=336
T=720
Avg
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
DF
✗
✗
ETTm1
0.310
0.352
0.356
0.377
0.388
0.400
0.450
0.437
0.376
0.391
ETTh1
0.372
0.391
0.430
0.424
0.486
0.454
0.507
0.486
0.449
0.439
ECL
0.143
0.237
0.161
0.252
0.178
0.270
0.218
0.303
0.175
0.265
Weather
0.160
0.203
0.210
0.247
0.267
0.289
0.346
0.342
0.246
0.270
QDF †
✓
✗
ETTm1
0.309
0.351
0.354
0.378
0.387
0.401
0.450
0.439
0.375
0.392
Table 3: Ablation study results.
Figure 3: Improvement of QDF applied to different forecast models, shown with colored bars for means over forecast lengths (96, 192, 336, 720) and error bars for 50% confidence intervals.
Method
T=96
T=192
T=336
T=720
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
DF
0.143
0.237
0.161
0.252
0.178
0.270
0.218
0.303
iMAML
0.135 5.74%↓
0.230 3.26%↓
0.154 4.31%↓
0.246 2.55%↓
0.170 4.48%↓
0.263 2.47%↓
0.205 5.90%↓
0.293 3.36%↓
MAML
0.136 5.54%↓
0.230 3.20%↓
0.154 4.24%↓
0.246 2.47%↓
0.170 4.71%↓
0.263 2.56%↓
0.205 5.65%↓
0.293 3.09%↓
MAML++
0.135 5.76%↓
0.229 3.33%↓
0.154 4.22%↓
0.246 2.49%↓
0.170 4.72%↓
0.263 2.65%↓
0.204 6.41%↓
0.292 3.67%↓
Reptile
0.136 5.06%↓
0.230 2.90%↓
0.155 3.73%↓
0.247 2.14%↓
0.171 3.91%↓
0.264 2.07%↓
0.206 5.36%↓
0.294 2.96%↓
Table 4: Comparison with meta-learning methods on ECL dataset.
Figure 4: Impact of hyperparameters on the performance of QDF.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: The label autocorrelation effect on the original label sequence and the components extracted by FreDF and Time-o1 [ Wang et al., 2025d , Wang et al., 2025e ] . The datasets are ETTh1, ETTh2, ECL, and Weather from left to right. The forecast length is uniformly set to 96.
Dataset
D
Forecast length
Train / validation / test
Frequency
Domain
ETTh1
7
96, 192, 336, 720
8545/2881/2881
Hourly
Health
ETTh2
7
96, 192, 336, 720
8545/2881/2881
Hourly
Health
ETTm1
7
96, 192, 336, 720
34465/11521/11521
15min
Health
ETTm2
7
96, 192, 336, 720
34465/11521/11521
15min
Health
Weather
21
96, 192, 336, 720
36792/5271/10540
10min
Weather
ECL
321
96, 192, 336, 720
18317/2633/5261
Hourly
Electricity
Appendix
Table 5: Dataset description.
Models
QDF
TQNet
PDF
Fredformer
iTransformer
FreTS
TimesNet
MICN
TiDE
PatchTST
DLinear
(Ours)
(2025)
(2024)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2023)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.307
0.349
0.310
0.352
0.326
0.363
0.326
0.361
0.338
0.372
0.342
0.375
0.368
0.394
0.319
0.366
0.353
0.374
0.325
0.364
0.346
0.373
192
0.352
0.376
0.356
0.377
0.365
0.381
0.365
0.382
0.382
0.396
0.385
0.400
0.406
0.409
0.364
0.395
0.391
0.393
0.363
0.383
0.380
0.390
336
0.383
0.398
0.388
0.400
0.397
0.402
0.396
0.404
0.427
0.424
0.416
0.421
0.454
0.444
0.395
0.425
0.423
0.414
0.404
0.413
0.413
0.414
720
0.441
0.434
0.450
0.437
0.458
0.437
0.459
0.444
0.496
0.463
0.513
0.489
0.527
0.474
0.505
0.499
0.486
0.448
0.463
0.442
0.472
0.450
Appendix
Table 6: Full results on the multi-step forecasting task. The length of history window is set to 96 for all baselines. Avg indicates the results averaged over forecasting lengths: T=96, 192, 336 and 720.
Figure 6: The forecast sequences generated with DF and QDF. The forecast length is set to 336 and the experiment is conducted on ETTm2.
Figure 7: The forecast sequences generated with DF and QDF. The forecast length is set to 192 and the experiment is conducted on ECL.
Loss
QDF
Time-o1
FreDF
Koopman
Soft-DTW
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Forecast model:TQNet
ETTm1
96
0.307
0.349
0.309
0.351
0.314
0.355
0.806
0.578
0.315
0.353
0.310
0.352
192
0.352
0.376
0.353
0.375
0.359
0.378
0.619
0.515
0.360
0.377
0.356
0.377
336
0.383
0.398
0.383
0.398
0.382
0.396
0.507
0.468
0.398
0.402
0.388
0.400
720
0.441
0.434
0.444
0.436
0.444
0.432
0.450
0.437
0.476
0.446
0.450
0.437
Appendix
Table 7: Comparable results with different learning objectives.
Figure 8: Performance of different forecast models with and without QDF. The forecast errors are averaged over forecast lengths and the error bars represent 50% confidence intervals.
Models
QDF
TQNet
QDF
PatchTST
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Input sequence length
96
96
0.158
0.201
0.160
0.203
0.180
0.224
0.189
0.230
192
0.207
0.245
0.210
0.247
0.226
0.262
0.228
0.262
336
0.263
0.286
0.267
0.289
0.279
0.300
0.288
0.305
720
0.342
0.339
0.346
0.342
0.354
0.347
0.362
0.354
Avg
0.242
0.268
0.246
0.270
0.260
0.283
0.267
0.288
Appendix
Table 8: Varying input sequence length results on the Weather dataset.
Figure 9: The running time of the QDF algorithm given varying forecast horizons ( T ). In each subfigure, the left panel considers the complexity of each inner-loop update (i.e., step 4 in Algorithm 1 ), the right panel considers the complexity of each outer-loop update (i.e., step 7 in Algorithm 1 ).
Dataset
QDF
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DF
Forecast model:TQNet
ETTm1
2.305
2.315
2.313
2.783
2.338
2.319
2.338
ETTh1
9.619
9.697
9.875
10.488
10.036
10.283
10.290
ECL
2.509
2.540
2.534
2.601
2.554
5.316
2.578
Weather
3.054
3.098
3.086
3.573
3.121
3.276
3.121
Forecast model:PDF
Appendix
Table 10: The comprehensive results of different learning objectives on the MAPE metric.
Loss
QDF
Time-o1
FreDF
Koopman
DF
Metric
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
SMAPE
MASE
OWA
Forecast model:TQNet
Yearly
13.355
3.015
0.788
13.377
3.004
0.787
13.404
3.022
0.790
22.588
5.512
1.385
13.502
3.074
0.800
Quarterly
10.018
1.174
0.883
10.174
1.200
0.899
10.116
1.196
0.895
17.713
2.415
1.685
10.132
1.192
0.895
Monthly
12.756
0.939
0.884
12.776
0.949
0.889
12.786
0.952
0.891
18.655
1.506
1.355
12.777
0.945
0.887
Others
4.909
3.203
1.022
5.039
3.285
1.048
4.908
3.219
1.024
7.478
5.365
1.633
5.048
3.292
1.050
Appendix
Table 11: The comprehensive results on the short-term forecasting task.
Figure 10: The dynamics of loss functions on ETTm1, ETTm2, ETTh1, ETTh2, ECL, and Weather from left to right. The forecast horizon T=336 .
Direction
Loop
T=64
T=96
T=128
T=192
T=256
T=336
T=512
T=720
Forward
Inner
1.196 ±0.007
1.175 ±0.011
1.176 ±0.009
1.409 ±0.012
1.431 ±0.015
1.590 ±0.011
1.523 ±0.012
1.763 ±0.011
Outer
1.161 ±0.010
1.168 ±0.014
1.162 ±0.009
1.350 ±0.019
1.366 ±0.011
1.357 ±0.013
1.361 ±0.012
1.672 ±0.012
Backward
Inner
1.147 ±0.005
1.144 ±0.005
1.137 ±0.006
1.487 ±0.007
1.492 ±0.008
1.647 ±0.007
1.591 ±0.010
1.770 ±0.010
Outer
0.954 ±0.006
0.967 ±0.006
0.950 ±0.008
1.190 ±0.009
1.194 ±0.009
1.343 ±0.008
1.296 ±0.007
1.477 ±0.008
Appendix
Table 12: Running complexity of QDF algorithm under task splits K=3.
Rank
Rank=1
Rank=0.8
Rank=0.6
Rank=0.4
Rank=0.2
DF
Metrics
MSE
MAE
MAPE
MSE
MAE
MAPE
MSE
MAE
MAPE
MSE
MAE
MAPE
MSE
MAE
MAPE
MSE
MAE
MAPE
Forecast model:TQNet
ETTm1
96
0.307
0.349
2.156
0.308
0.350
2.191
0.310
0.351
2.174
0.309
0.351
2.200
0.310
0.352
2.201
0.310
0.352
2.212
192
0.352
0.376
2.281
0.354
0.377
2.288
0.355
0.379
2.308
0.354
0.378
2.298
0.352
0.377
2.295
0.356
0.377
2.288
336
0.383
0.398
2.329
0.387
0.399
2.349
0.387
0.399
2.344
0.387
0.400
2.356
0.387
0.401
2.357
0.388
0.400
2.338
720
0.441
0.434
2.522
0.445
0.437
2.534
0.446
0.437
2.546
0.443
0.437
2.530
0.445
0.437
2.538
0.450
0.437
2.516
Appendix
Table 13: The performance of varying ranks of the inverse covariance matrix Σˉ .
Figure 11: The learned correlation matrix on different datasets: ETTm1, ETTm2 and Weather from left to right.
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting models primarily adhere to the Direct Forecast (DF) paradigm, generating multi-step forecasts independently and disregarding label autocorrelation over time. In this work, we demonstrate that the learning objective of DF is biased in the presence of label autocorrelation. To address this issue, we propose the Frequency-enhanced Direct Forecast (FreDF), which mitigates label autocorrelation by learning to forecast in the frequency domain, thereby reducing estimation bias. Our experiments show that FreDF significantly outperforms existing state-of-the-art methods and is compatible with a variety of forecast models. Code is available at https://github.com/Master-PLC/FreDF.
Hao Wang, Licheng Pan, Zhichao Chen +6
Department of Control Science and Engineering, Zhejiang University · School of Automation, Central South University · Trust and Safety Team, TikTok Sydney, ByteDance Inc. +3
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation proves biased when the label sequence exhibits autocorrelation. In this paper, we propose DistDF, which achieves alignment by minimizing a distributional discrepancy between the conditional distributions of forecast and label sequences. Since such conditional discrepancies are difficult to estimate from finite time-series observations, we introduce a joint-distribution Wasserstein discrepancy for time-series forecasting, which provably upper bounds the conditional discrepancy of interest. The proposed discrepancy is tractable, differentiable, and readily compatible with gradient-based optimization. Extensive experiments show that DistDF improves diverse forecasting models and achieves leading performance. Code is available at https://anonymous.4open.science/r/DistDF-F66B.
Hao Wang, Licheng Pan, Yuan Lu +7
Xiaohongshu Inc. · College of Control Science and Technology, Zhejiang University · College of Computer Science and Technology, Zhejiang University +4
Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https://github.com/Master-PLC/Time-o1.
Hao Wang, Licheng Pan, Zhichao Chen +5
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Gaoling School of Artificial Intelligence, Renmin University of China +4