Test-time adaptation (TTA) is a promising paradigm for handling distribution shift in time-series forecasting (TSF), where models adapt at inference time, often leveraging delayed observed data to refine predictions. In the multivariate setting, distribution shifts often exhibit cross-variate dependencies, yet existing TSF-TTA methods adapt each variate independently and ignore this cross-variate structure. Exploiting such structure motivates cross-variate interaction, but coupling variates through backbone predictions introduces direct pathways for mixing uncorrected errors across variates, a concern under the delayed supervision of TSF-TTA. We identify the \emph{interaction space} as a key design choice, and show that acting on adapter corrections that refine backbone outputs, the \emph{correction space}, rather than on the predictions themselves, avoids directly propagating backbone errors across variates. We build on this to propose \textsc{CoRe} (\textsc{Co}rrection-space Interaction \textsc{Re}finement), realizing correction-space interaction through (i) Shared-anchor Correction Refinement (SCR), which combines each variate's correction with a shared anchor through a parameter-efficient bottleneck, and (ii) input-conditioned spectral gating, which adaptively modulates the refinement from the current input window. Across seven backbones, six datasets, and four prediction horizons, \textsc{CoRe} reduces MSE by 25.82% on average over backbones and 10.57% over the state-of-the-art TSF-TTA method, with stronger gains at medium-to-long horizons and modest computational overhead. Data and code are available at: https://github.com/yyddou/CoReTTA
Figures & tables
Figure 1 : Direct error mixing in prediction-space cross-variate interaction. A conceptual illustration using variables from Weather (Tpot, VPmax, OT, and other variables). Cross-variate interaction applied to backbone predictions can directly mix their uncorrected errors across variates (Step 2). In TSF-TTA, ground-truth observations become available only after the prediction horizon, delaying corrective feedback (Step 3). Proposition 3.2 formalizes the direct error-mixing pathway, while Figure 5 empirically compares prediction-space and correction-space interaction across prediction horizons.
Figure 2
Figure 4 : Framework of CoRe for multivariate test-time adaptation. A frozen backbone produces a base forecast, while a base adapter produces per-variate corrections. SCR performs cross-variate interaction in the correction space . Spectral descriptor of the current input window provides per-variate gates that modulate the SCR refinement. The refined corrections are added to the unchanged backbone prediction to form the final forecast.
Table 4
Method
DLinear
PatchTST
MICN
Avg. MSE
vs. COSA
COSA
0.2981
0.2927
0.3241
0.3050
–
+SCR (no anchor; per-variate bottleneck)
0.2760
0.2721
0.3018
0.2833
−7.1%
+SCR (no anchor/bottleneck; full C×C mixing)
0.2745
0.2706
0.3044
0.2832
−7.2%
+SCR (loss-trend gate)
0.2712
0.2657
0.2963
0.2777
−9.0%
+SCR (fixed gate, g=1 )
0.2703
0.2651
0.2954
0.2769
−9.2%
CoRe
0.2675
0.2646
0.2942
0.2754
−9.7%
Table 4 : Ablation study of CoRe components. Average MSE over three backbones (DLinear, PatchTST, MICN), six datasets, four horizons, and 10 seeds. vs. COSA : relative MSE reduction over COSA. Loss-trend gate : COSA’s adaptive learning-rate signal repurposed as the SCR gate. Fixed gate ( g=1 ): uniform cross-variate refinement without input-dependent modulation. st : Spectral descriptors for gating (Section 3.5 ).
Figure 6
Appendix figures & tables22 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6 : Controlled comparison of correction-space and prediction-space interaction across all seven backbones.
Backbone
Dataset
H=96
H=192
H=336
H=720
Avg. Gain
TAFAS
+CoRe
TAFAS
+ CoRe
TAFAS
+ CoRe
TAFAS
+ CoRe
DLinear
ETTh1
0.4604
0.4604
0.5099
0.5020
0.5622
0.5515
0.6685
0.6548
+ 1.83%
ETTh2
0.2302
0.2039
0.2836
0.2710
0.3193
0.3229
0.3939
0.4140
+ 2.41%
ETTm1
0.3488
0.3342
0.4159
0.3924
0.4787
0.4681
0.5495
0.5539
+ 2.81%
ETTm2
0.1588
0.1505
0.1927
0.1841
0.2337
0.2298
0.3053
0.3008
+ 3.21%
Weather
0.1828
0.1531
0.2219
0.2136
0.2695
0.2454
0.3550
0.3164
+ 9.95%
Appendix
Table 7 : MSE of TAFAS and TAFAS+ CoRe (ours). Bold = better result. Avg. Gain = relative MSE reduction averaged over H∈{96,192,336,720} .
ETTh1
ETTh2
ETTm1
ETTm2
Backbone
w/o SCR
CoRe -MLP
w/o SCR
CoRe -MLP
w/o SCR
CoRe -MLP
w/o SCR
CoRe -MLP
DLinear
0.4842
0.4767
0.2519
0.2378
0.5285
0.5101
0.2096
0.1949
FreTS
0.4722
0.4681
0.2556
0.2439
0.5243
0.4957
0.2056
0.1923
iTransformer
0.4685
0.4673
0.2710
0.2605
0.5337
0.5145
0.2296
0.2131
PatchTST
0.4595
0.4526
0.2534
0.2494
0.5265
0.5038
0.2121
0.1980
OLS
0.4706
0.4661
0.2503
0.2383
0.5270
0.5087
0.2086
0.1966
Appendix
Table 8: MSE of CoRe -MLP without and with SCR, averaged over 4 horizons and 10 seeds. w/o SCR : standalone MLP correction only. CoRe -MLP : MLP correction + SCR. Gain: relative MSE reduction from SCR.
Backbone
CoRe vs. PETSA
DLinear
22.64%
FreTS
21.90%
iTransformer
20.86%
MICN
24.83%
OLS
21.71%
PatchTST
24.03%
Appendix
Table 9 : Comparison with PETSA [ Medeiros et al., 2025 ] . We report the MSE reduction of CoRe relative to PETSA for each backbone, averaged over six datasets and four prediction horizons. Both methods use the same frozen backbone checkpoints. Higher is better.
DLinear
FreTS
OLS
iTransformer
MICN
H
DynaTTA
COSA
CoRe
DynaTTA
COSA
CoRe
DynaTTA
COSA
CoRe
DynaTTA
COSA
CoRe
DynaTTA
COSA
CoRe
ETTh1
96
0.4708
0.4922
0.4916
0.4511
0.4623
0.4625
0.4486
0.4729
0.4723
0.4509
0.4638
0.4545
0.5804
0.5837
0.5863
192
0.5321
0.5371
0.4961
0.5138
0.5084
0.4692
0.5082
0.5173
0.4763
0.5156
0.5024
0.4589
0.6032
0.5802
0.5203
336
0.5792
0.5208
0.4722
0.5838
0.4970
0.4596
0.5626
0.5041
0.4568
0.6052
0.4956
0.4661
0.7037
0.5889
0.5386
720
0.7047
0.5433
0.4891
0.7086
0.5616
0.5052
0.6933
0.5443
0.4886
0.7118
0.5381
0.5055
0.8282
0.6177
0.5591
ETTh2
96
0.2338
0.2552
0.2426
0.2397
0.2560
0.2462
0.2325
0.2443
0.2389
0.2767
0.3050
0.2803
0.2612
0.2538
0.2476
Appendix
Table 10 : Prediction MSE comparison with DynaTTA across five backbones (lower is better). Bold : best among TTA methods. Underline : second-best among TTA methods. DynaTTA results use the COSA re-implementation with prediction-drift signal.
Dataset
H
RevIN
FAN
CoRe
ETTh1
96
0.4591
0.4620
0.4916
192
0.5121
0.5159
0.4961
336
0.5587
0.5427
0.4722
720
0.7063
0.6593
0.4891
ETTh2
96
0.2302
0.2512
0.2426
192
0.2834
0.2973
0.2275
Appendix
Table 11 : Prediction MSE comparison with RevIN and FAN on DLinear (lower is better).
Dataset
H
Base (TimeXer)
TAFAS
COSA
CoRe
PETSA
ETTh1
96
0.4301
0.4304 ± 0.0004
0.4407 ± 0.0001
0.4367 ± 0.0019
0.4349 ± 0.0002
192
0.4895
0.4996 ± 0.0001
0.4867 ± 0.0004
0.4513 ± 0.0049
0.4922 ± 0.0003
336
0.5605
0.5795 ± 0.0113
0.4776 ± 0.0003
0.4329 ± 0.0025
0.5544 ± 0.0005
720
0.6931
0.6810 ± 0.0010
0.5195 ± 0.0001
0.4813 ± 0.0046
0.6698 ± 0.0014
ETTh2
96
0.2367
0.2317 ± 0.0001
0.2492 ± 0.0008
0.2469 ± 0.0034
0.2367 ± 0.0003
192
0.3000
0.3021 ± 0.0001
0.2907 ± 0.0035
0.2420 ± 0.0039
0.2986 ± 0.0002
Appendix
Table 12 : Evaluation on TimeXer [ Wang et al., 2024 ] . MSE is reported for the frozen TimeXer backbone and four TTA methods over six datasets and four horizons. TTA results are mean ± standard deviation over 10 seeds. Lower is better. Bold denotes the best TTA result in each setting.
Family
Backbones
MSE reduction vs. COSA
CI
DLinear, FreTS, OLS, PatchTST
10.87%
CD
iTransformer, Informer, MICN
10.18%
Appendix
Table 13 : Average MSE reduction of CoRe over COSA by backbone family, averaged over six datasets and four horizons.
Figure 15
Dense
SCR r=C
SCR r=16
Avg params
339,024
9,922
16,480
Avg imp. vs. baseline
+23.09%
+25.82%
+20.37%
Win rate vs. baseline
128/168
162/168
147/168
Appendix
Table 14: Adapter-only parameter counts and accuracy for three cross-variate interaction designs, averaged over 168 settings.
Configuration
avg_imp
Mean pooling, r=4
+18.09%
Mean pooling, r=8
+18.84%
Mean pooling, r=16
+19.66%
Mean pooling, r=32
+20.24%
Mean pooling, r=C
+25.45%
Max pooling, r=16
+19.36%
Appendix
Table 16 : Ablation on SCR pooling strategy and bottleneck rank (DLinear backbone, 24 settings). avg_imp : average MSE improvement over backbone.
Model
Dataset
H
Backbone
COSA
CoRe
COSA → CoRe
DLinear
Electricity
96
0.2078
0.2024
0.1992
+1.58%
192
0.2081
0.1922
0.1771
+7.86%
336
0.2228
0.1947
0.1751
+10.07%
720
0.2644
0.2139
0.1926
+9.96%
Traffic
96
0.6710
0.6685
0.6648
+0.55%
192
0.6251
0.6127
0.6031
+1.57%
Appendix
Table 17 : Results on large-scale datasets Electricity ( C=321 ) and Traffic ( C=862 ), averaged over 10 seeds. COSA → CoRe : MSE improvement of CoRe over COSA. All methods use a fixed bottleneck rank r=16 for these datasets to keep the parameter cost independent of C .
Method
DLinear
PatchTST
MICN
Avg. MSE
vs. COSA
COSA
0.2981
0.2927
0.3241
0.3050
–
+SCR (no anchor; per-variate bottleneck)
0.2760
0.2721
0.3018
0.2833
−7.1%
+SCR (no anchor/bottleneck; full C×C mixing)
0.2745
0.2706
0.3044
0.2832
−7.2%
+SCR (loss-trend gate)
0.2712
0.2657
0.2963
0.2777
−9.0%
+SCR (fixed gate, g=1 )
0.2703
0.2651
0.2954
0.2769
−9.2%
+SCR ( st=[SEt] )
0.2676
0.2638
0.2946
0.2753
−9.7%
Appendix
Table 18 : Ablation study of CoRe components. Average MSE over three backbones (DLinear, PatchTST, MICN), six datasets, four horizons, and 10 seeds. vs. COSA : relative MSE reduction over COSA. Loss-trend gate : COSA’s adaptive learning-rate signal repurposed as the SCR gate. Fixed gate ( g=1 ): uniform cross-variate refinement without input-dependent modulation. SEt : spectral entropy; LBR/MBR/HBR t : low/mid/high band energy ratios (Section 3.5 ).
Figure 9 : Example spectral-gate trajectories over adaptation steps on Weather and Exchange Rate.
Dataset
Mean
Std
Near zero ( ∣g∣<0.1 )
Near init ( g<−0.8 )
ETTh1
−0.583
0.325
5.9%
34.5%
ETTh2
−0.631
0.285
9.3%
37.5%
ETTm1
−0.357
0.421
9.9%
15.1%
ETTm2
−0.498
0.354
15.5%
28.4%
Exchange
−0.676
0.194
2.4%
23.9%
Weather
−0.352
0.371
20.3%
9.0%
Appendix
Table 19 : SCR spectral gate statistics across datasets (DLinear backbone, averaged over 4 horizons and 10 seeds). Mean: time-averaged gate value across all variates and test batches; Std: standard deviation across variates and batches; Near zero ( ∣g∣<0.1 ): fraction of gate activations close to zero, indicating suppressed cross-variate refinement; Near init ( g<−0.8 ): fraction of gate activations close to the initialisation bias bg=−1 , indicating gates that have not moved significantly from initialisation. The gate is bounded in (−1,1) by tanh .
bg
−2.0
−1.0
0.0
vs. default ↓
+0.46%
+0.00%
+1.56%
Appendix
Table 20 : Gate bias sensitivity.
TAFAS
COSA
CoRe
Model
Dataset
RT
pneg
RT
pneg
RT
pneg
DLinear
ETTh1
−0.003
0.519
−0.041
0.402
−0.080
0.300
ETTh2
+0.003
0.570
−0.056
0.279
−0.089
0.223
ETTm1
−0.024
0.330
+0.013
0.383
−0.035
0.382
ETTm2
+0.001
0.546
−0.026
0.391
−0.055
0.330
Exchange
−0.013
0.213
−0.262
0.142
−0.274
0.088
Appendix
Table 21 : Online adaptation consistency over 18 backbone–dataset combinations (3 backbones × 6 datasets), each averaged over H∈{96,192,336,720} . RT : mean per-batch regret relative to the frozen backbone (negative = better than backbone on average); pneg : fraction of batches where the adapted model underperforms the backbone. Best per row in bold . Overall row also reports RT+ (average regret on settings where the method is worse than the backbone) and nworse : number of backbone–dataset combinations (out of 18) with positive mean regret.
Dataset
N
Window
Mean ∣r∣
Shifting (std)
CV
Range
Strong
Weak
ETTh1
7
168
0.222
0.266
0.919
0.771
9.5%
81.0%
ETTh2
7
168
0.325
0.237
0.846
0.865
4.8%
57.1%
ETTm1
7
672
0.224
0.264
0.913
0.771
9.5%
81.0%
ETTm2
7
672
0.325
0.234
0.845
0.872
4.8%
57.1%
Weather
21
1008
0.296
0.180
0.736
0.637
21.0%
64.3%
Exchange Rate †
8
30
0.305
0.479
0.957
0.992
3.6%
42.9%
Appendix
Table 22 : Pairwise Pearson correlation statistics across variates. N : number of variates; Window: rolling window size (approx. 1 week in physical time). Mean ∣r∣ : time-averaged absolute pairwise correlation computed on stationary series after ADF test. Shifting (std) : standard deviation of windowed correlations over time, measuring temporal instability of cross-variate dependencies (computed on raw series to preserve trend-switching signals). CV : coefficient of variation (std/mean). Range : peak-to-trough variation of windowed correlations. Strong ( ∣r∣>0.7 ) and Weak ( ∣r∣≤0.3 ): fraction of variate pairs by static ∣r∣ strength. Datasets above the divider are used in main experiments.
Dataset
CV
Cos-sim std
Fixed gate
Spectral gate
Gain
ETTh1
0.217
0.097
0.4889
0.4873
+0.34%
Exchange
0.322
0.223
0.0890
0.0871
+2.19%
ETTm1
0.428
0.117
0.4598
0.4563
+0.76%
ETTm2
0.514
0.137
0.1716
0.1700
+0.95%
ETTh2
0.554
0.119
0.2329
0.2299
+1.29%
Weather
0.696
0.151
0.1797
0.1747
+2.77%
Appendix
Table 23 : Per-variate correction heterogeneity and the benefit of spectral gating over fixed gating (DLinear backbone). CV : coefficient of variation of per-variate correction norms across variates; Cos-sim std : standard deviation of cosine similarity between each variate’s correction and the cross-variate mean anchor; Fixed gate and Spectral gate : average MSE under each gating strategy (10 seeds); Gain : relative improvement from spectral over fixed gating. Datasets ordered by CV.
Figure 10 : Qualitative comparison between COSA and CoRe on Exchange Rate across four prediction horizons and three selected variates.
Figure 11 : Qualitative comparison between COSA and CoRe on Weather across four prediction horizons and three selected variates.
Component
COSA(ms)
CoRe (ms)
Overhead(ms)
Backbone prediction
1.27
1.04
−0.22
Spectral FFT
-
1.14
+1.14
SCR forward
-
1.45
+1.45
Loss computation
0.84
1.00
+0.16
Backward + update
4.42
4.11
−0.31
Per adaptation step
6.28
6.76
+0.48(+7.7%)
Appendix
Table 24 : Detailed per-step timing breakdown averaged over six datasets and four horizons (DLinear). Overhead = CoRe − COSA; "-" indicates a component absent in that method.
Test-Time Adaptation (TTA) aims to improve time series forecasting under distribution shifts by using limited observations revealed during inference. However, forecasting TTA must operate in a source-free online setting, where the adaptation signal is short, temporally correlated, and potentially noisy. Existing methods can therefore suffer from weak identifiability, error accumulation, and unstable long-horizon corrections when the revealed prefix is sparse or contaminated. To address these issues, we propose STEPS, a Smooth Temporal Error Propagation Solver for TTA in time-series forecasting. STEPS reformulates forecasting TTA as a Dirichlet Boundary Value Problem on a temporal manifold, where the revealed prefix error serves as the boundary condition for the unknown future error field. Then, STEPS solves a smooth and bounded correction field in prediction space: a Local Solver propagates prefix errors under temporal smoothness, a Global Solver retrieves stable cross-window error memory and Spatiotemporal Manifold Fusion (SMF) integrates both solutions into the final correction. Across six standard benchmarks and four frozen backbones, STEPS achieves an average relative MSE reduction of 26.82% over the zero-shot backbone, exceeding the strongest compared TTA baseline by 12.77%. Additional sparse prefix and contamination tests confirm the robustness of STEPS under limited and noisy prefixes.
Jiaqi Liu, Yifan Ouyang, Zhifei Song +2
School of Artificial Intelligence and Robotics, Xiamen University Malaysia · S.M.A.R.T. NEXUS Centre of Excellence, Xiamen University Malaysia Sepang, Selangor 43900, Malaysia
Test-time adaptation (TTA) has recently emerged as a promising approach for improving time series forecasting (TSF) under distribution shift. Existing TSF-TTA methods differ in how they utilize revealed targets, yet the resulting adaptation protocols remain heterogeneous and lack a clearly unified formulation. To address this issue, we revisit TSF-TTA from the perspective of protocol cleanliness and propose an adaptation protocol based solely on matured ground truth, yielding a more principled setting for adaptation. Under this protocol, we further diagnose existing adapters in the frequency domain and find that their prediction corrections often exhibit limited and weakly structured spectral modifications. Motivated by this diagnosis, we propose Frequency-Aware Calibration (FAC), a lightweight calibration method that directly parameterizes prediction corrections in the frequency domain. Across diverse datasets, forecasting horizons, and source forecasters, FAC achieves competitive and consistent performance while requiring substantially fewer trainable parameters than the compared TSF-TTA adapters.
Time Series Foundation Models (TSFMs) advance generalization and data efficiency in time series forecasting by unified large-scale pretraining. But TSFMs remain lacking when adapting to specific downstream forecasting tasks for two reasons. First, the non-stationary and uncertain nature of time series data lead to inevitable temporal distribution shifts between historical training and future testing data, while current Supervised FineTuning (SFT)-based methods are prone to overfitting and may degrade generalization. Second, training data availability varies across forecasting tasks, requiring TSFMs to generalize well under diverse data regimes. To address these challenges, we introduce the Time series Reinforcement Finetuning (TimeRFT) paradigm for TSFM downstream adaptation, which consists of two task-specific training recipes: i) A forecasting quality-based temporal reward mechanism that conducts a multi-faceted evaluation of the contribution of each prediction step to overall forecasting accuracy. ii) A forecasting difficulty-based data selection strategy to identify time series samples with generalizable predictive patterns and informative training signals. Extensive experiments demonstrate TimeRFT can consistently outperform SFT-based adaptation methods across various real-world forecasting tasks and training data regimes, enhancing prediction accuracy and generalization against unforeseen distribution shifts.
Siyang Li, Yize Chen, Zijie Zhu +4
HKUST(GZ) Guangzhou, China · University of Alberta Edmonton, Canada · Alibaba Cloud Guangzhou, China +3