Organizations: Xiaohongshu Inc. · College of Control Science and Technology, Zhejiang University · College of Computer Science and Technology, Zhejiang University · School of Computing and Artificial Intelligence, Shanghai University of Finance and Economics · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Squirrel AI · Institute for Artificial Intelligence, Peking University
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation proves biased when the label sequence exhibits autocorrelation. In this paper, we propose DistDF, which achieves alignment by minimizing a distributional discrepancy between the conditional distributions of forecast and label sequences. Since such conditional discrepancies are difficult to estimate from finite time-series observations, we introduce a joint-distribution Wasserstein discrepancy for time-series forecasting, which provably upper bounds the conditional discrepancy of interest. The proposed discrepancy is tractable, differentiable, and readily compatible with gradient-based optimization. Extensive experiments show that DistDF improves diverse forecasting models and achieves leading performance. Code is available at https://anonymous.4open.science/r/DistDF-F66B.
Figures & tables
Figure 1: The conditional correlation of label components given x , where the forecast length is set to T=192 . The correlation matrices are computed for the raw labels (a), the frequency components in FreDF (b) ( Wang et al., 2025d ) and the principal components in Time-o1 (c) ( Wang et al., 2025c ) .
Loss
DistDF
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
TimeBridge
ETTm1
0.383
0.397
0.383
0.397
0.386
0.398
0.460
0.438
0.387
0.400
0.395
0.402
0.387
0.400
ETTh1
0.434
0.436
0.439
0.438
0.439
0.436
0.459
0.449
0.464
0.452
0.452
0.445
0.442
0.440
ECL
0.172
0.267
0.175
0.268
0.175
0.267
0.182
0.277
0.176
0.271
0.173
0.268
0.176
0.271
Weather
0.248
0.275
0.250
0.275
0.254
0.276
0.269
0.293
0.252
0.277
0.260
0.280
0.252
0.277
Fredformer
ETTm1
0.378
0.394
0.379
0.393
0.384
0.394
0.389
0.400
0.389
0.400
0.397
0.402
0.387
0.398
Table 1: Comparative results with other objectives for time-series forecasting.
Figure 2: The forecast sequence of DF (in blue) and DistDF (in red), with history length H=96 .
Model
Align μ
Align Σ
Data
T=96
T=192
T=336
T=720
Avg
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
DF
✗
✗
ETTm1
0.326
0.361
0.365
0.382
0.396
0.404
0.459
0.444
0.387
0.398
ETTh1
0.377
0.396
0.437
0.425
0.486
0.449
0.488
0.467
0.447
0.434
ECL
0.142
0.239
0.161
0.257
0.182
0.278
0.217
0.309
0.176
0.271
Weather
0.168
0.211
0.214
0.254
0.273
0.297
0.353
0.347
0.252
0.277
DistDF †
✓
✗
ETTm1
0.318
0.359
0.361
0.382
0.393
0.404
0.453
0.440
0.381
0.396
Table 2: Ablation study results.
Discrepancy
Ours
EMD
MMD@Linear
MMD@RBF
KL
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
TimeBridge
ETTm1
0.383
0.398
0.388
0.400
0.385
0.400
0.387
0.399
0.387
0.400
0.387
0.400
ETTh1
0.433
0.437
0.441
0.439
0.438
0.437
0.441
0.440
0.437
0.438
0.442
0.440
ECL
0.172
0.267
0.177
0.272
0.174
0.269
0.172
0.266
0.176
0.271
0.176
0.271
Weather
0.248
0.275
0.251
0.276
0.253
0.278
0.250
0.276
0.253
0.277
0.252
0.277
Fredformer
ETTm1
0.379
0.395
0.386
0.397
0.380
0.395
0.385
0.397
0.385
0.397
0.387
0.398
Table 3: Comparative results with other discrepancies for aligning the joint distributions.
Figure 3: Improvement of DistDF applied to different forecasting models, shown with colored bars for means over forecast lengths (96, 192, 336, 720) and error bars for 50% confidence intervals.
Table 7
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Models
DistDF
TimeBridge
Fredformer
iTransformer
FreTS
TimesNet
MICN
TiDE
PatchTST
DLinear
(Ours)
(2025)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2023)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.316
0.357
0.323
0.361
0.326
0.361
0.338
0.372
0.342
0.375
0.368
0.394
0.319
0.366
0.353
0.374
0.325
0.364
0.346
0.373
192
0.358
0.380
0.366
0.385
0.365
0.382
0.382
0.396
0.385
0.400
0.406
0.409
0.364
0.395
0.391
0.393
0.363
0.383
0.380
0.390
336
0.392
0.404
0.398
0.408
0.396
0.404
0.427
0.424
0.416
0.421
0.454
0.444
0.395
0.425
0.423
0.414
0.404
0.413
0.413
0.414
720
0.448
0.437
0.461
0.445
0.459
0.444
0.496
0.463
0.513
0.489
0.527
0.474
0.505
0.499
0.486
0.448
0.463
0.442
0.472
0.450
Appendix
Table 6: Comparative results with other established methods for time-series forecasting.
Figure 4: The forecast sequences generated with DF and DistDF. The forecast length is set to 336 and the experiment is conducted on ETTm2.
Figure 5: The forecast sequences generated with DF and DistDF. The forecast length is set to 192 and the experiment is conducted on ECL.
Loss
DistDF
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Forecast model: TimeBridge
ETTm1
96
0.319
0.358
0.318
0.356
0.325
0.361
0.572
0.493
0.321
0.360
0.321
0.359
0.323
0.361
192
0.363
0.383
0.363
0.382
0.373
0.385
0.410
0.407
0.366
0.386
0.368
0.385
0.366
0.385
336
0.394
0.405
0.396
0.407
0.398
0.406
0.397
0.408
0.397
0.409
0.405
0.410
0.398
0.408
720
0.455
0.442
0.456
0.443
0.450
0.438
0.460
0.445
0.462
0.447
0.486
0.453
0.461
0.445
Appendix
Table 7: Comparable results with different learning objectives.
Figure 6: Performance of different forecasting models with and without DistDF. The forecasting errors are averaged over forecast lengths and the error bars represent 50% confidence intervals.
Figure 7: Running time (ms) with varying forecast length.
Models
TQNet
TQNet †
TimeBridge
TimeBridge †
Fredformer
Fredformer †
iTransformer
iTransformer †
FreTS
FreTS †
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
96
0.372
0.391
0.372
0.391
0.373
0.395
0.372
0.392
0.377
0.396
0.373
0.393
0.385
0.405
0.383
0.403
0.398
0.409
0.399
0.409
192
0.430
0.424
0.430
0.422
0.428
0.426
0.424
0.429
0.437
0.425
0.428
0.425
0.440
0.437
0.438
0.434
0.451
0.442
0.457
0.447
336
0.486
0.454
0.472
0.444
0.471
0.451
0.467
0.450
0.486
0.449
0.466
0.445
0.480
0.457
0.476
0.455
0.501
0.472
0.504
0.474
720
0.507
0.486
0.477
0.468
0.495
0.487
0.472
0.471
0.488
0.467
0.453
0.453
0.504
0.492
0.492
0.483
0.608
0.571
0.557
0.537
Avg
0.449
0.439
0.438
0.431
0.442
0.440
0.434
0.436
0.447
0.434
0.430
0.429
0.452
0.448
0.447
0.444
0.489
0.474
0.479
0.467
Appendix
Table 8: The performance comparison of DF and DistDF on different forecasting models.
Figure 8: Evolution of training objectives and validation metrics across four datasets: ETTm1, ETTh1, and ECL (from left to right).
Models
TimeBridge
TimeBridge †
Fredformer
Fredformer †
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.405
0.402
0.395
0.391
0.391
0.396
0.386
0.390
192
0.467
0.438
0.419
0.408
0.494
0.449
0.493
0.446
336
0.518
0.467
0.460
0.437
0.572
0.500
0.579
0.486
720
0.725
0.514
0.527
0.478
1.821
0.837
0.833
0.563
Avg
0.528
0.455
0.450
0.428
0.820
0.546
0.573
0.471
Appendix
Table 9: The performance comparison of DF and DistDF on the autoregressive forecasting setting.
Models
D3U
D3U †
Metrics
MSE
MAE
CRPS
CRPS sum
MSE
MAE
CRPS
CRPS sum
ETTm1
96
0.317
0.357
0.263
0.723
0.316
0.357
0.265
0.720
192
0.361
0.383
0.285
0.749
0.360
0.383
0.282
0.747
336
0.394
0.404
0.299
0.742
0.390
0.402
0.298
0.731
720
0.460
0.437
0.325
0.892
0.453
0.435
0.328
0.849
Avg
0.383
0.395
0.293
0.776
0.380
0.394
0.293
0.762
Appendix
Table 10: The performance comparison of DF and DistDF on the probabilistic forecasting task.
Models
TimeMixer
TimeMixer †
SCINet
SCINet †
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.329
0.369
0.326
0.369
0.325
0.365
0.319
0.359
192
0.371
0.391
0.373
0.392
0.383
0.397
0.367
0.385
336
0.427
0.425
0.412
0.423
0.436
0.424
0.403
0.406
720
0.564
0.506
0.491
0.459
0.528
0.476
0.469
0.444
Avg
0.422
0.423
0.401
0.411
0.418
0.416
0.389
0.399
Appendix
Table 11: The performance comparison of DF and DistDF on the multi-scale architectures.
Models
DistDF
TimeBridge
DistDF
PatchTST
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Historical sequence length
96
96
0.164
0.209
0.168
0.211
0.179
0.220
0.189
0.230
192
0.212
0.252
0.214
0.254
0.222
0.257
0.228
0.262
336
0.270
0.295
0.273
0.297
0.278
0.298
0.288
0.305
720
0.348
0.345
0.353
0.347
0.354
0.348
0.362
0.354
Avg
0.248
0.275
0.252
0.277
0.258
0.281
0.267
0.288
Appendix
Table 12: Varying input sequence length results on the Weather dataset.
Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https://github.com/Master-PLC/Time-o1.
Hao Wang, Licheng Pan, Zhichao Chen +5
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Gaoling School of Artificial Intelligence, Renmin University of China +4
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting models primarily adhere to the Direct Forecast (DF) paradigm, generating multi-step forecasts independently and disregarding label autocorrelation over time. In this work, we demonstrate that the learning objective of DF is biased in the presence of label autocorrelation. To address this issue, we propose the Frequency-enhanced Direct Forecast (FreDF), which mitigates label autocorrelation by learning to forecast in the frequency domain, thereby reducing estimation bias. Our experiments show that FreDF significantly outperforms existing state-of-the-art methods and is compatible with a variety of forecast models. Code is available at https://github.com/Master-PLC/FreDF.
Hao Wang, Licheng Pan, Zhichao Chen +6
Department of Control Science and Engineering, Zhejiang University · School of Automation, Central South University · Trust and Safety Team, TikTok Sydney, ByteDance Inc. +3
The design of learning objectives is central to training time-series forecasting models. Existing learning objectives such as mean squared error mostly treat each future step as an independent, equally weighted task, which leads to the following two challenges: (1) they overlook the label autocorrelation effect among future steps, leading to biased learning objectives; (2) they fail to set heterogeneous task weights for different forecasting tasks corresponding to varying future steps, limiting the forecasting performance. To fill this gap, we propose a novel quadratic-form weighted learning objective, addressing both issues simultaneously. Specifically, the off-diagonal elements of the weighting matrix account for the label autocorrelation effect, whereas the non-uniform diagonals are expected to match the preferred weights of the forecasting tasks with varying future steps. On this basis, we propose a Quadratic Direct Forecast (QDF) learning algorithm, which trains the forecast model using the adaptively updated quadratic-form weighting matrix. Experiments show that our QDF effectively improves the performance of various forecast models, achieving state-of-the-art results. Code is available at https://github.com/Master-PLC/QDF.
Hao Wang, Licheng Pan, Yuan Lu +7
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · College of Engineering, Purdue University +6