Organizations: Xiaohongshu Inc. · College of Control Science and Technology, Zhejiang University · College of Computer Science and Technology, Zhejiang University · School of Computing and Artificial Intelligence, Shanghai University of Finance and Economics · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Squirrel AI · Institute for Artificial Intelligence, Peking University
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approach resorts to minimizing the conditional negative log-likelihood, typically estimated by the mean squared error. However, this estimation proves biased when the label sequence exhibits autocorrelation. In this paper, we propose DistDF, which achieves alignment by minimizing a distributional discrepancy between the conditional distributions of forecast and label sequences. Since such conditional discrepancies are difficult to estimate from finite time-series observations, we introduce a joint-distribution Wasserstein discrepancy for time-series forecasting, which provably upper bounds the conditional discrepancy of interest. The proposed discrepancy is tractable, differentiable, and readily compatible with gradient-based optimization. Extensive experiments show that DistDF improves diverse forecasting models and achieves leading performance. Code is available at https://anonymous.4open.science/r/DistDF-F66B.
Figures & tables
Figure 1: The conditional correlation of label components given x , where the forecast length is set to T=192 . The correlation matrices are computed for the raw labels (a), the frequency components in FreDF (b) ( Wang et al., 2025d ) and the principal components in Time-o1 (c) ( Wang et al., 2025c ) .
Loss
DistDF
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
TimeBridge
ETTm1
0.383
0.397
0.383
0.397
0.386
0.398
0.460
0.438
0.387
0.400
0.395
0.402
0.387
0.400
ETTh1
0.434
0.436
0.439
0.438
0.439
0.436
0.459
0.449
0.464
0.452
0.452
0.445
0.442
0.440
ECL
0.172
0.267
0.175
0.268
0.175
0.267
0.182
0.277
0.176
0.271
0.173
0.268
0.176
0.271
Weather
0.248
0.275
0.250
0.275
0.254
0.276
0.269
0.293
0.252
0.277
0.260
0.280
0.252
0.277
Fredformer
ETTm1
0.378
0.394
0.379
0.393
0.384
0.394
0.389
0.400
0.389
0.400
0.397
0.402
0.387
0.398
Table 1: Comparative results with other objectives for time-series forecasting.
Figure 2: The forecast sequence of DF (in blue) and DistDF (in red), with history length H=96 .
Model
Align μ
Align Σ
Data
T=96
T=192
T=336
T=720
Avg
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
DF
✗
✗
ETTm1
0.326
0.361
0.365
0.382
0.396
0.404
0.459
0.444
0.387
0.398
ETTh1
0.377
0.396
0.437
0.425
0.486
0.449
0.488
0.467
0.447
0.434
ECL
0.142
0.239
0.161
0.257
0.182
0.278
0.217
0.309
0.176
0.271
Weather
0.168
0.211
0.214
0.254
0.273
0.297
0.353
0.347
0.252
0.277
DistDF †
✓
✗
ETTm1
0.318
0.359
0.361
0.382
0.393
0.404
0.453
0.440
0.381
0.396
Table 2: Ablation study results.
Discrepancy
Ours
EMD
MMD@Linear
MMD@RBF
KL
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
TimeBridge
ETTm1
0.383
0.398
0.388
0.400
0.385
0.400
0.387
0.399
0.387
0.400
0.387
0.400
ETTh1
0.433
0.437
0.441
0.439
0.438
0.437
0.441
0.440
0.437
0.438
0.442
0.440
ECL
0.172
0.267
0.177
0.272
0.174
0.269
0.172
0.266
0.176
0.271
0.176
0.271
Weather
0.248
0.275
0.251
0.276
0.253
0.278
0.250
0.276
0.253
0.277
0.252
0.277
Fredformer
ETTm1
0.379
0.395
0.386
0.397
0.380
0.395
0.385
0.397
0.385
0.397
0.387
0.398
Table 3: Comparative results with other discrepancies for aligning the joint distributions.
Figure 3: Improvement of DistDF applied to different forecasting models, shown with colored bars for means over forecast lengths (96, 192, 336, 720) and error bars for 50% confidence intervals.
Table 7
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Models
DistDF
TimeBridge
Fredformer
iTransformer
FreTS
TimesNet
MICN
TiDE
PatchTST
DLinear
(Ours)
(2025)
(2024)
(2024)
(2023)
(2023)
(2023)
(2023)
(2023)
(2023)
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.316
0.357
0.323
0.361
0.326
0.361
0.338
0.372
0.342
0.375
0.368
0.394
0.319
0.366
0.353
0.374
0.325
0.364
0.346
0.373
192
0.358
0.380
0.366
0.385
0.365
0.382
0.382
0.396
0.385
0.400
0.406
0.409
0.364
0.395
0.391
0.393
0.363
0.383
0.380
0.390
336
0.392
0.404
0.398
0.408
0.396
0.404
0.427
0.424
0.416
0.421
0.454
0.444
0.395
0.425
0.423
0.414
0.404
0.413
0.413
0.414
720
0.448
0.437
0.461
0.445
0.459
0.444
0.496
0.463
0.513
0.489
0.527
0.474
0.505
0.499
0.486
0.448
0.463
0.442
0.472
0.450
Appendix
Table 6: Comparative results with other established methods for time-series forecasting.
Figure 4: The forecast sequences generated with DF and DistDF. The forecast length is set to 336 and the experiment is conducted on ETTm2.
Figure 5: The forecast sequences generated with DF and DistDF. The forecast length is set to 192 and the experiment is conducted on ECL.
Loss
DistDF
Time-o1
FreDF
Koopman
Dilate
Soft-DTW
DF
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Forecast model: TimeBridge
ETTm1
96
0.319
0.358
0.318
0.356
0.325
0.361
0.572
0.493
0.321
0.360
0.321
0.359
0.323
0.361
192
0.363
0.383
0.363
0.382
0.373
0.385
0.410
0.407
0.366
0.386
0.368
0.385
0.366
0.385
336
0.394
0.405
0.396
0.407
0.398
0.406
0.397
0.408
0.397
0.409
0.405
0.410
0.398
0.408
720
0.455
0.442
0.456
0.443
0.450
0.438
0.460
0.445
0.462
0.447
0.486
0.453
0.461
0.445
Appendix
Table 7: Comparable results with different learning objectives.
Figure 6: Performance of different forecasting models with and without DistDF. The forecasting errors are averaged over forecast lengths and the error bars represent 50% confidence intervals.
Figure 7: Running time (ms) with varying forecast length.
Models
TQNet
TQNet †
TimeBridge
TimeBridge †
Fredformer
Fredformer †
iTransformer
iTransformer †
FreTS
FreTS †
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTh1
96
0.372
0.391
0.372
0.391
0.373
0.395
0.372
0.392
0.377
0.396
0.373
0.393
0.385
0.405
0.383
0.403
0.398
0.409
0.399
0.409
192
0.430
0.424
0.430
0.422
0.428
0.426
0.424
0.429
0.437
0.425
0.428
0.425
0.440
0.437
0.438
0.434
0.451
0.442
0.457
0.447
336
0.486
0.454
0.472
0.444
0.471
0.451
0.467
0.450
0.486
0.449
0.466
0.445
0.480
0.457
0.476
0.455
0.501
0.472
0.504
0.474
720
0.507
0.486
0.477
0.468
0.495
0.487
0.472
0.471
0.488
0.467
0.453
0.453
0.504
0.492
0.492
0.483
0.608
0.571
0.557
0.537
Avg
0.449
0.439
0.438
0.431
0.442
0.440
0.434
0.436
0.447
0.434
0.430
0.429
0.452
0.448
0.447
0.444
0.489
0.474
0.479
0.467
Appendix
Table 8: The performance comparison of DF and DistDF on different forecasting models.
Figure 8: Evolution of training objectives and validation metrics across four datasets: ETTm1, ETTh1, and ECL (from left to right).
Models
TimeBridge
TimeBridge †
Fredformer
Fredformer †
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.405
0.402
0.395
0.391
0.391
0.396
0.386
0.390
192
0.467
0.438
0.419
0.408
0.494
0.449
0.493
0.446
336
0.518
0.467
0.460
0.437
0.572
0.500
0.579
0.486
720
0.725
0.514
0.527
0.478
1.821
0.837
0.833
0.563
Avg
0.528
0.455
0.450
0.428
0.820
0.546
0.573
0.471
Appendix
Table 9: The performance comparison of DF and DistDF on the autoregressive forecasting setting.
Models
D3U
D3U †
Metrics
MSE
MAE
CRPS
CRPS sum
MSE
MAE
CRPS
CRPS sum
ETTm1
96
0.317
0.357
0.263
0.723
0.316
0.357
0.265
0.720
192
0.361
0.383
0.285
0.749
0.360
0.383
0.282
0.747
336
0.394
0.404
0.299
0.742
0.390
0.402
0.298
0.731
720
0.460
0.437
0.325
0.892
0.453
0.435
0.328
0.849
Avg
0.383
0.395
0.293
0.776
0.380
0.394
0.293
0.762
Appendix
Table 10: The performance comparison of DF and DistDF on the probabilistic forecasting task.
Models
TimeMixer
TimeMixer †
SCINet
SCINet †
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
ETTm1
96
0.329
0.369
0.326
0.369
0.325
0.365
0.319
0.359
192
0.371
0.391
0.373
0.392
0.383
0.397
0.367
0.385
336
0.427
0.425
0.412
0.423
0.436
0.424
0.403
0.406
720
0.564
0.506
0.491
0.459
0.528
0.476
0.469
0.444
Avg
0.422
0.423
0.401
0.411
0.418
0.416
0.389
0.399
Appendix
Table 11: The performance comparison of DF and DistDF on the multi-scale architectures.
Models
DistDF
TimeBridge
DistDF
PatchTST
Metrics
MSE
MAE
MSE
MAE
MSE
MAE
MSE
MAE
Historical sequence length
96
96
0.164
0.209
0.168
0.211
0.179
0.220
0.189
0.230
192
0.212
0.252
0.214
0.254
0.222
0.257
0.228
0.262
336
0.270
0.295
0.273
0.297
0.278
0.298
0.288
0.305
720
0.348
0.345
0.353
0.347
0.354
0.348
0.362
0.354
Avg
0.248
0.275
0.252
0.277
0.258
0.281
0.267
0.288
Appendix
Table 12: Varying input sequence length results on the Weather dataset.
May 23, 2025·Hao Wang, Licheng Pan, Zhichao Chen +5
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Gaoling School of Artificial Intelligence, Renmin University of China +4
Feb 4, 2024·Hao Wang, Licheng Pan, Zhichao Chen +6
Department of Control Science and Engineering, Zhejiang University · School of Automation, Central South University · Trust and Safety Team, TikTok Sydney, ByteDance Inc. +3
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · College of Engineering, Purdue University +6