Organizations: University of Science and Technology Beijing, China · Zhejiang University of Technology, Hangzhou, China · Institution of Artificial Intelligence, University of Science and Technology Beijing, China
Long-term time series forecasting (LTSF) models predominantly employ patch-based encoders terminated by a flatten readout head that maps the entire encoded historical memory to all future steps through a single shared projection. This implicit coupling of future positions obscures position-specific historical-to-future alignment and amplifies sensitivity to corrupted inputs and extreme supervision noise. We present SACQ, a plug-in structured prediction head that replaces flatten readout while keeping the encoder unchanged. SACQ adopts a two-stage decoding pipeline: it first establishes a coarse patch-grid forecast scaffold, then refines each future position through cross-attention over historical memory and merges the attention-derived correction with the coarse scaffold via a learned per-patch gate. To stabilize optimization under long horizons and noisy labels, we further propose a batch-adaptive scaled log-cosh loss that automatically calibrates robustness to the current residual scale, suppressing outlier gradients while preserving MSE-like sensitivity for typical errors. SACQ attains top-tier test MSE/MAE across PatchTST, DLinear, and patch-Mamba backbones with only modest incremental overhead in parameters and latency. Under inference-time input corruption and training-set label-noise stress tests, SACQ substantially outperforms flatten readouts, with ablation studies validating each architectural component.
Figures & tables
Fig. 1: Flatten readout vs. SACQ decoding. Left: flatten maps encoded historical memory to all future steps through one dense projection, sharing readout parameters across the horizon. Right: SACQ initializes coarse future-patch predictions, refines them with cross-self attention over historical memory, and merges a gated residual correction at each future patch. Details are given in Sec. III-C .
Fig. 2: Cross-self attention module structure. Left: the first L−1 layers, each at hidden width D (cross-attention to H , coarse fusion as in Eq. ( 3 ), self-attention, and FFN). Right: the final ( L -th) layer follows the same high-level pipeline; the ellipsis inside the L -th block denotes submodules that are structurally identical to those in the preceding layers. In the last layer only, a linear layer maps from width D to width p before self-attention; the self-attention and FFN sublayers then run in the p -dimensional space. After the final FFN, the stack outputs Δ for Eq. ( 4 ).
Fig. 3: Robustness on ETTh1 under training-set corruption. (a) Tf=96 , input masking: test MSE versus missing rate ρ on historical windows. (b) Tf=96 , target spikes: test MSE versus spike intensity (clean to +7 σy ). (c) – (d) Same sweeps at long horizon Tf=720 . Curves compare Flatten+MSE, Flatten+scaled log-cosh, SAR+MSE, and SAR+scaled log-cosh; evaluation uses the clean test split.
Fig. 4: SAR ablation under training-set tail shift. (a)–(b) ETTh1 at Tf=96 and 720 . (c)–(d) Weather at Tf=96 and 720 . Solid curves: full SACQ (with SAR) and w/o SAR ; dashed curve: relative test-MSE reduction from SAR. Tail-shift rate increases the fraction of spike-corrupted supervision in the forecast tail during training only.
Fig. 5: Horizon-wise mean gate α under clean versus spiked input histories ( +3σ on the lookback window). Dashed: clean test input. Solid: +3σ spiked test input. Colors index the training-set target-spike ratio ( 0%,1%,3%,5% ).
Fig. 6: Lookback sensitivity on ETTh1. Test MSE versus input length N for DLinear, PatchTST (flatten), SACQ+MSE, and SACQ+scaled log-cosh. (a) Short-horizon forecasting ( Tf=96 ): all methods improve when N grows from 96 to 336–512, then plateau. (b) Long-horizon forecasting ( Tf=720 ): SACQ+scaled log-cosh remains stable at N=1440 , whereas flatten baselines spike at the longest contexts.
Fig. 7: Scaling with backbone width dmodel . Left: parameter count (M). Right: forward-pass latency (ms). Blue solid: total model; red dashed: SACQ-only overhead excluding the coarse flatten map.
Long-term time series forecasting benefits from preserving global structure such as trends and seasonality. Recent LLM-based forecasters often improve accuracy through test-time scaling (e.g., iterative refinement), but these methods are computationally expensive and increasingly prone to global-shape mismatch as the prediction horizon extends. We propose SCALER, a coarse-to-fine forecasting framework that first employs a lightweight Transformer tailored to long-term shape modeling to predict a coarse representation of future dynamics. This predicted shape then serves as a compact guide for an LLM to perform test-time scaling via iterative coarse-to-fine residual token refinement, while processing substantially fewer tokens at each step. By guiding refinement with an explicit future-shape prediction, SCALER reduces reliance on long description prompts, and its fixed-step refinement avoids costly reward-model-based selection, further lowering computational overhead. Experimental results demonstrate that SCALER outperforms strong forecasting baselines in long-term, short-term and zero-shot forecasting while significantly reducing the inference cost associated with scaled LLM for time series forecasting. Code: https://github.com/xuanmay2701/SCALER.
Xuan-May Le, Minh-Tuan Tran, Ling Luo +3
The University of Melbourne Melbourne, Victoria, Australia · Monash University Melbourne, Victoria, Australia
Time Series Foundation Models (TSFMs) have borrowed the long context paradigm from natural language processing under the premise that feeding more history into the model improves forecast quality. But in stochastic domains, distant history is often just high-frequency noise, not signal. Hence, the proposed work tests whether this premise actually holds by running continuous context architectures (PatchTST included) through the ETTh1 benchmark. The obtained results contradict the premise: an inverse scaling law shows up clearly, with forecasting error rising as context gets longer. A 3,000-step window causes performance to drop by over 68%, evidence that attention mechanisms are poor at ignoring irrelevant historical volatility. Retrieval-Augmented Forecasting (RAFT) is evaluated as an alternative. RAFT achieves a mean squared error (MSE) of 0.379 with a fixed 720-step window and selective retrieval, outperforming both long-context configurations and zero-shot foundation models (Chronos, Moirai) despite requiring far less computation. In addition, the retrieval step injects only the most relevant historical segments as dynamic exogenous variables, which gives the model a context-informed inductive bias it cannot build on its own from raw sequences. Therefore, foundation models going forward need to shift architecturally toward selective retrieval.
Rishi Ahuja, Kumar Prateek, Simranjit Singh +1
Department of Information Technology, Dr. B.R. Ambedkar National Institute of Technology Jalandhar, Punjab, 144008, India.
Long-term time series forecasting (LTSF) plays a crucial role in fields such as energy management, finance, and traffic prediction. Transformer-based models have adopted patch-based strategies to capture long-range dependencies, but accurately modeling shape similarities across patches and variables remains challenging due to scale differences. To address this, we introduce patch-mean decoupling (PMD), which separates the trend and residual shape information by subtracting the mean of each patch, preserving the original structure and ensuring that the attention mechanism captures true shape similarities. Futhermore, to more effectively model long-range dependencies and capture cross-variable relationships, we propose Trend Restoration Attention (TRA) and Proximal Variable Attention (PVA). The former module reintegrates the decoupled trend from PMD while calculating attention output. And the latter focuses cross-variable attention on the most relevant, recent time segments to avoid overfitting on outdated correlations. Combining these components, we propose PMDformer, a model designed to effectively capture shape similarity in long-term forecasting scenarios. Extensive experiments indicate that PMDformer outperforms existing state-of-the-art methods in stability and accuracy across multiple LTSF benchmarks. The code is available at https://github.com/aohu1105/PMDformer.
Ao Hu, Liangjian Wen, Jiang Duan +7
Southwestern University of Finance and Economics · Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province · Chengdu Everimaging Science and Technology Co., Ltd. +4