Organizations: University of Science and Technology Beijing, China · Zhejiang University of Technology, Hangzhou, China · Institution of Artificial Intelligence, University of Science and Technology Beijing, China
Long-term time series forecasting (LTSF) models predominantly employ patch-based encoders terminated by a flatten readout head that maps the entire encoded historical memory to all future steps through a single shared projection. This implicit coupling of future positions obscures position-specific historical-to-future alignment and amplifies sensitivity to corrupted inputs and extreme supervision noise. We present SACQ, a plug-in structured prediction head that replaces flatten readout while keeping the encoder unchanged. SACQ adopts a two-stage decoding pipeline: it first establishes a coarse patch-grid forecast scaffold, then refines each future position through cross-attention over historical memory and merges the attention-derived correction with the coarse scaffold via a learned per-patch gate. To stabilize optimization under long horizons and noisy labels, we further propose a batch-adaptive scaled log-cosh loss that automatically calibrates robustness to the current residual scale, suppressing outlier gradients while preserving MSE-like sensitivity for typical errors. SACQ attains top-tier test MSE/MAE across PatchTST, DLinear, and patch-Mamba backbones with only modest incremental overhead in parameters and latency. Under inference-time input corruption and training-set label-noise stress tests, SACQ substantially outperforms flatten readouts, with ablation studies validating each architectural component.
Figures & tables
Fig. 1: Flatten readout vs. SACQ decoding. Left: flatten maps encoded historical memory to all future steps through one dense projection, sharing readout parameters across the horizon. Right: SACQ initializes coarse future-patch predictions, refines them with cross-self attention over historical memory, and merges a gated residual correction at each future patch. Details are given in Sec. III-C .
Fig. 2: Cross-self attention module structure. Left: the first L−1 layers, each at hidden width D (cross-attention to H , coarse fusion as in Eq. ( 3 ), self-attention, and FFN). Right: the final ( L -th) layer follows the same high-level pipeline; the ellipsis inside the L -th block denotes submodules that are structurally identical to those in the preceding layers. In the last layer only, a linear layer maps from width D to width p before self-attention; the self-attention and FFN sublayers then run in the p -dimensional space. After the final FFN, the stack outputs Δ for Eq. ( 4 ).
Fig. 3: Robustness on ETTh1 under training-set corruption. (a) Tf=96 , input masking: test MSE versus missing rate ρ on historical windows. (b) Tf=96 , target spikes: test MSE versus spike intensity (clean to +7 σy ). (c) – (d) Same sweeps at long horizon Tf=720 . Curves compare Flatten+MSE, Flatten+scaled log-cosh, SAR+MSE, and SAR+scaled log-cosh; evaluation uses the clean test split.
Fig. 4: SAR ablation under training-set tail shift. (a)–(b) ETTh1 at Tf=96 and 720 . (c)–(d) Weather at Tf=96 and 720 . Solid curves: full SACQ (with SAR) and w/o SAR ; dashed curve: relative test-MSE reduction from SAR. Tail-shift rate increases the fraction of spike-corrupted supervision in the forecast tail during training only.
Fig. 5: Horizon-wise mean gate α under clean versus spiked input histories ( +3σ on the lookback window). Dashed: clean test input. Solid: +3σ spiked test input. Colors index the training-set target-spike ratio ( 0%,1%,3%,5% ).
Fig. 6: Lookback sensitivity on ETTh1. Test MSE versus input length N for DLinear, PatchTST (flatten), SACQ+MSE, and SACQ+scaled log-cosh. (a) Short-horizon forecasting ( Tf=96 ): all methods improve when N grows from 96 to 336–512, then plateau. (b) Long-horizon forecasting ( Tf=720 ): SACQ+scaled log-cosh remains stable at N=1440 , whereas flatten baselines spike at the longest contexts.
Fig. 7: Scaling with backbone width dmodel . Left: parameter count (M). Right: forward-pass latency (ms). Blue solid: total model; red dashed: SACQ-only overhead excluding the coarse flatten map.
Southwestern University of Finance and Economics · Artificial Intelligence and Digital Finance Key Laboratory of Sichuan Province · Chengdu Everimaging Science and Technology Co., Ltd. +4