Autocorrelation is a common property of time-series, where each observation is dependent on its predecessors. In deep time-series forecasting, it raises two central challenges: (1) designing backbone architectures to model autocorrelation in history sequences, and (2) devising loss functions to model autocorrelation in label sequences. Recent studies have made strides in tackling these challenges, but a systematic survey examining both aspects remains lacking. To bridge this gap, this paper reviews deep time-series forecasting from an autocorrelation modeling perspective, offering two contributions beyond existing surveys. First, it introduces a taxonomy that jointly covers both backbone architectures and loss functions, whereas prior surveys provide limited coverage of the latter. Second, it analyzes the motivations and insights underlying the surveyed literature from a unified autocorrelation perspective, providing a holistic overview of the field's evolution. Additional resources and details are available at https://github.com/Master-PLC/Awesome-TSF-Papers.
Figures & tables
Fig. 1 : The ACF at different lags on time-series datasets.
Fig. 2 : The visualization of representative temporal patterns and their ACFs.
Author
Year
Taxonomy
Backbone Architecture
Loss Function
Non-Transformers
Transformers
Plug-ins
LikeEst
ShapeAlign
DistBal
CondGene
Wen et al. [ 41 ]
2023
Backbone architecture
-
✓
-
-
-
-
-
Li et al. [ 63 ]
2024
Backbone architecture
✓
✓
-
-
-
-
-
Wang et al. [ 2 ]
2024
Backbone architecture
✓
✓
✓
-
-
-
-
Casolaro et al. [ 64 ]
2023
Backbone architecture
✓
✓
-
-
-
-
✓
Liang et al. [ 62 ]
2024
Backbone architecture
✓
✓
-
-
-
-
✓
TABLE I : Overview of recent surveys.
Fig. 4 : Overview of representative non-Transformer architectures. Circular nodes represent model inputs and outputs.
Fig. 5 : Overview of representative Transformer architectures. The circlular nodes represent model input and output. The colored blocks indicate the primary modifications to adapt Transformer to time-series forecasting.
Model
Year
Backbone
Tokenizer
Texts
Cross-modal Alignment Strategy
Fine-tuned LLM parameters
Query
Repro
Pretrain
Finetuning
PromptCast [ 110 ]
2023
BERT
Discrete
✓
✓
-
-
-
-
LLMTime [ 111 ]
2023
GPT3
Discrete
-
✓
-
-
-
-
Time-LLM [ 112 ]
2024
LLaMA
Patching
✓
-
✓
-
-
-
Time-FFM [ 113 ]
2024
GPT2
Patching
✓
-
✓
-
-
-
GPT4TS [ 114 ]
2023
GPT2
Patching
-
-
-
-
✓
LN, PE
TABLE II : Overview of LLM-based time-series backbone architectures with standard self-attention token-mixers.
Fig. 6 : A typical layout of the plug-in layers following [ 115 ] , where the colored blocks represent the backbone architecture.
Loss function
Year
Type
Agnostic
AutoDiff
Analytical
Infer. free
Param. free
MSE [ 146 ]
2021
-
✓
✓
✓
✓
✓
AutoMSE [ 174 ]
2021
Likelihood estimation
✓
✓
✓
✓
✓
FreDF [ 14 ]
2025
Likelihood estimation
✓
✓
✓
✓
✓
Time-o1 [ 42 ]
2025
Likelihood estimation
✓
✓
✓
✓
✓
LatentTSF [ 175 ]
2026
Likelihood estimation
✓
✓
✓
-
-
DBLoss [ 176 ]
2025
Likelihood estimation
✓
✓
✓
✓
✓
TABLE III : Overview of loss functions harnessing label autocorrelation.
Fig. 7 : The computation of loss functions based on likelihood estimation, where the dark blocks contain learnable parameters.
Fig. 8 : The computation of loss functions based on shape alignment, where the dark blocks contain learnable parameters.
Fig. 9 : The computation of loss functions based on distribution balancing, where the dark blocks contain learnable parameters.
Fig. 10 : The computation of loss functions based on conditional generation, where the dark blocks contain learnable parameters.
Modern deep-learning models have achieved remarkable success in time-series forecasting. Yet, their performance degrades in long-term prediction due to error accumulation in autoregressive inference, where predictions are recursively used as inputs. While classical error correction mechanisms (ECMs) have long been used in statistical methods, their applicability to deep learning models remains limited or ineffective. In this work, we revisit the error accumulation problem in deep time-series forecasting and investigate the role and necessity of ECMs in this new context. We propose a simple, architecture-agnostic error correction model that can be integrated with any existing forecaster without requiring retraining. By explicitly decomposing predictions into trend and seasonal components and training the corrector to adjust each separately, we introduce the Universal Error Corrector with Seasonal-Trend Decomposition (UEC-STD), which significantly improves correction accuracy and robustness across 4 backbones and 10 datasets. Our findings provide a practical tool for enhancing forecasts while offering new insights into mitigating autoregressive errors in deep time-series models. Code is available at https://github.com/DA2I2-SLM/UEC-STD.
Minh Hoang Nguyen, Dai Do, Huu Hiep Nguyen +3
Deakin Applied AI Initiative, Deakin University, Australia.
Training time-series forecasting models poses unique challenges in loss function design. Most existing approaches adopt temporal mean squared error, but this study reveals two critical limitations: (1) it ignores the presence of label autocorrelation, which biases it from the true label sequence likelihood; (2) it involves excessive number of tasks, which complicates optimization, especially for long-term forecasting. To address these issues, we introduce Time-o1, a transform-enhanced loss function for time-series forecasting. The central idea is to transform the label sequence into decorrelated components with discriminated significance. Models are then trained to align the most significant components, thereby effectively mitigating label autocorrelation and reducing task amount. Experiments demonstrate that Time-o1 achieves state-of-the-art performance and is compatible with various forecast models. Code is available at https://github.com/Master-PLC/Time-o1.
Hao Wang, Licheng Pan, Zhichao Chen +5
Xiaohongshu Inc. · State Key Lab of General AI, School of Intelligence Science and Technology, Peking University · Gaoling School of Artificial Intelligence, Renmin University of China +4
Time series modeling presents unique challenges due to autocorrelation in both historical data and future sequences. While current research predominantly addresses autocorrelation within historical data, the correlations among future labels are often overlooked. Specifically, modern forecasting models primarily adhere to the Direct Forecast (DF) paradigm, generating multi-step forecasts independently and disregarding label autocorrelation over time. In this work, we demonstrate that the learning objective of DF is biased in the presence of label autocorrelation. To address this issue, we propose the Frequency-enhanced Direct Forecast (FreDF), which mitigates label autocorrelation by learning to forecast in the frequency domain, thereby reducing estimation bias. Our experiments show that FreDF significantly outperforms existing state-of-the-art methods and is compatible with a variety of forecast models. Code is available at https://github.com/Master-PLC/FreDF.
Hao Wang, Licheng Pan, Zhichao Chen +6
Department of Control Science and Engineering, Zhejiang University · School of Automation, Central South University · Trust and Safety Team, TikTok Sydney, ByteDance Inc. +3