Organizations: Aberdeen Institute of Data Science and Artificial Intelligence South China Normal University Foshan, China · School of Artificial Intelligence South China Normal University Foshan, China · Faculty of Humanities and Arts Macau University of Science and Technology Macau, China
Retrieval-augmented time-series forecasting uses the continuations of historical segments similar to the current context as references for a forecaster. Most existing methods build the retrieval memory once from the training segment, leaving observations revealed after deployment unavailable as references, and generally do not calibrate how much the retrieved information should influence a frozen forecaster. We identify two key determinants of retrieval utility for a frozen forecaster: whether the history still reflects the current state, and whether the correction it induces aligns with the forecaster's residual errors, an alignment that can shift between validation and deployment when the memory becomes stale. We propose FreshCast, a plug-in retrieval framework that keeps the forecaster frozen, continuously updates a non-parametric memory with new observations, forms a memory forecast through relational kernel regression, and calibrates its weight in closed form on the validation segment. Under a simplified generative model, we characterize the optimal combination gain through the second-order relation between forecaster error and memory correction, and show that a sufficiently long look-back can make periodic memory information redundant. Across seven benchmarks and ten forecasting architectures, FreshCast reduces average MSE for every evaluated forecaster and input length, by 14.6% and 5.6% at input lengths 96 and 720, and achieves lower MSE than the evaluated retrieval-augmented and online baselines in their comparison settings. Ablations show that freezing the memory at the end of training removes most of the gain, identifying post-training observations as a primary source of improvement. For a frozen forecaster, useful historical references must remain timely and provide information that helps correct its remaining errors.
Figures & tables
Figure 1. (a) MSE reduction brought by the same retrieval scheme (GTR) and by FreshCast on forecasters ranked by validation error; brackets: gap in percentage points. (b) Average MSE reduction of GTR and FreshCast relative to the forecaster as the input length L grows (MLP, DLinear, iTransformer). (c) Existing retrieval-augmented forecasting versus FreshCast. (d) Validation-estimated optimal combination gain versus realized test gain of FreshCast ( L=96 , ten forecasters; Spearman correlation 0.70).
Figure 2. What retrieval returns. (a) On Traffic, joint distribution of phase offset and age for the candidate pool and for segments retrieved by raw cosine similarity, RAFT, and PFRP. (b) Cumulative age distribution of the ten most useful references in a post hoc evaluation (solid) and of the candidate pool (dashed); legend: share of useful references newer than the static memory. (c) Error of same-phase value readouts by age (solid), without the level term (dashed), and the level term (shaded); Traffic, top; ETTh1, bottom. (d) Level and amplitude of Traffic windows in the static memory, after training, and at test queries; contours: 50% (solid) and 90% (dashed) of each group; circles: medians; dashed box: 1–99% range of the static memory.
Figure 3. Overview of FreshCast. Top: (1) the streaming memory supplies seasonal analogs, seasonal profiles, and shape analogs, and relational kernel regression weights their readouts into the memory forecast r ; (2) w^ is estimated on the validation segment and fixed at test time. Bottom: (a) observed history and causal admissibility s+H≤t ; (b) the two candidate families; (c) the value and shape readouts.
MLP
DLinear
PatchTST
iTransformer
CycleNet
Dataset
Base
+FreshCast
Base
+FreshCast
Base
+FreshCast
Base
+FreshCast
Base
+FreshCast
ETTh1
0.449
0.403
0.462
0.438
0.485
0.461
0.472
0.405
0.450
0.403
ETTh2
0.374
0.368
0.560
0.371
0.530
0.372
0.387
0.375
0.378
0.374
ETTm1
0.383
0.330
0.403
0.339
0.434
0.350
0.413
0.336
0.380
0.330
ETTm2
0.279
0.250
0.347
0.251
0.356
0.251
0.290
0.252
0.270
0.248
Weather
0.258
0.224
0.267
0.227
0.255
0.226
0.270
0.227
0.246
0.222
Table 1. Test MSE of the ten forecasters alone (Base) and with FreshCast at L=96 , averaged over the four horizons. The upper block shows the five original forecasters and the lower block the five forecasters proposed in 2025–2026.
Figure 4. Main results. (a) MSE reduction of FreshCast relative to each forecaster at the three input lengths; top row: mean over forecasters. (b) MSE reduction relative to the forecaster of each retrieval-augmented baseline (hollow; diamonds: RAFT with a streaming bank) and of FreshCast (filled); numbers: gap in percentage points. (c) Distribution of per-setting MSE reduction of FreshCast relative to the forecaster, +GTR, RAFT, and +PFRP (bars: interquartile range; circles: medians; dots: settings), with the number of positive settings. (d) MSE reduction relative to the forecaster at L=96 when new observations are written into parameters (online fine-tuning, DynaME) or into the memory (FreshCast); brackets: gap in percentage points.
Method
Forecaster
L=96
L=336
L=720
GTR
MLP
−7.8%
−4.1%
−3.9%
DLinear
−15.2%
−7.6%
−8.7%
PatchTST
−18.3%
−21.9%
–
iTransformer
−12.6%
−8.0%
−8.2%
RAFT (original)
MLP
−12.7%
−5.1%
−4.0%
RAFT, streaming ‡
MLP
−10.0%
−4.9%
−4.3%
Table 2. Average MSE change of FreshCast relative to retrieval-augmented methods in their comparison settings and to DynaME with the forecaster frozen (negative: lower MSE with FreshCast; dashes: not evaluated).
L=96
L=720
Variant
MLP
iTr.
Cyc.
Recent
MLP
iTr.
Cyc.
Recent
Streaming, validation w^ (full)
−11.4
−12.4
−9.6
−14.0
−4.7
−8.6
−4.8
−4.47
Streaming, training wtrain
−11.0
−12.0
−9.2
−14.2
−3.7
−5.4
−3.4
−3.88
Static, validation w^
−2.15
−2.64
−1.16
−4.134
−0.75
−1.63
+0.18
+0.642
Static, training wtrain
+5.413
+5.712
+5.817
+6.476
−0.47
−3.32
−0.79
+4.979
Shape analogs only
−7.0
−8.1
−5.7
−9.81
−3.51
−6.8
−3.51
−3.310
Table 3. Ablation: MSE change (%) relative to the forecaster; Recent averages the five forecasters proposed in 2025–2026 (Appendix E ). A superscript gives the number of settings worse than the forecaster (of 28, or of 140 for Recent; omitted when zero). The upper block crosses the timeliness of the memory with where the combination weight is estimated; the lower block changes the memory content, the readout, or the kernel, with a streaming memory and the validation estimate w^ .
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Channels
Interval
Train / val / test
Daily
Weekly
ETTh1, ETTh2
7
1 hour
8640 / 2880 / 2880
24
168
ETTm1, ETTm2
7
15 min
34560 / 11520 / 11520
96
672
Weather
21
10 min
7:1:2 (52696)
144
1008
Electricity
321
1 hour
7:1:2 (26304)
24
168
Traffic
862
1 hour
7:1:2 (17544)
24
168
Appendix
Table 4. Datasets.
MLP
DLinear
PatchTST
iTransformer
CycleNet
Dataset
L
Base
+FreshCast
Base
+FreshCast
Base
+FreshCast
Base
+FreshCast
Base
+FreshCast
ETTh1
336
0.430
0.402
0.427
0.424
0.425
0.406
0.463
0.410
0.439
0.410
720
0.441
0.402
0.442
0.424
–
–
0.485
0.409
0.440
0.400
ETTh2
336
0.373
0.361
0.476
0.365
0.344
0.344
0.376
0.359
0.374
0.362
720
0.362
0.357
0.552
0.361
–
–
0.386
0.364
0.367
0.358
ETTm1
336
0.362
0.331
0.358
0.338
0.352
0.329
0.370
0.334
0.362
0.330
Appendix
Table 5. Test MSE of the ten forecasters alone (Base) and with FreshCast at L=336 and L=720 , averaged over the four horizons. PatchTST is not evaluated at L=720 .
SparseTSF
DLinear
PatchTST
TimesNet
Dataset
Base
+PFRP
+FreshCast
Base
+PFRP
+FreshCast
Base
+PFRP
+FreshCast
Base
+PFRP
+FreshCast
ETTh1
0.084
0.078
0.076
0.116
0.112
0.106
0.079
0.079
0.077
0.077
0.079
0.076
ETTh2
0.202
0.193
0.195
0.224
0.214
0.209
0.201
0.191
0.187
0.191
0.192
0.187
ETTm1
0.054
0.052
0.051
0.065
0.064
0.052
0.054
0.053
0.051
0.053
0.053
0.051
ETTm2
0.126
0.120
0.120
0.126
0.122
0.119
0.123
0.120
0.117
0.122
0.120
0.118
Weather
0.0015
0.0015
0.0014
0.0062
0.0051
0.0038
0.0019
0.0017
0.0018
0.0017
0.0017
0.0016
Appendix
Table 6. Univariate setting: test MSE of the four forecasters in the PFRP paper alone (Base), with PFRP, and with FreshCast ( L=96 ), averaged over the four horizons; the lowest value in each group is in bold. The target is the channel specified in the original paper; Weather is given to four decimal places.
Figure 5. Shape-readout share of the kernel weights on seasonal analogs, by level gap (a) and segment age (b); numbers in (a): gap between static and streaming memory in percentage points.
Figure 6. MSE change relative to the forecaster in eight equal test intervals with streaming or static memory (median and range over the seven datasets; L=96 ; MLP, iTransformer, CycleNet; four horizons; five kernel seeds).
Retrieval-augmented generation (RAG) complements parametric models with retrieved external evidence. The same idea is attractive for continuous-output regression, but directly reusing retrieved target values is often not robust when samples differ in output level, numerical scale, or local dynamics. Moreover, conventional forecasting pipelines generally use residuals for model optimization and error diagnosis, but do not retain individual historical residual examples as memory that can be accessed at inference time.For multivariate time-series forecasting, we propose RATL, a plug-in residual-retrieval and feedback-correction method. RATL freezes a base forecaster to construct retrieval keys and turns its historical forecast residuals into a train-only memory specific to that base model. At inference time, RATL retrieves residual trajectories from similar historical contexts subject to causal availability constraints, then uses a set-aware router operating over forecast blocks and variables to select and combine these trajectories. Experiments show that historical residuals matched to the current context contain reusable forecasting information and that RATL improves frozen base forecasters in most experimental settings. Ablations further show that learned routing strengthens raw residual feedback, while validation-based correction-strength selection limits residual over-injection.On real-world benchmarks, we use iTransformer as the primary frozen base forecaster, compare against multiple strong forecasting baselines, and test transferability across backbones. The results show that RATL can further improve base-forecaster performance in most settings.Overall, RATL shifts the retrieved object from historical target values to base-model-specific historical forecast errors, providing a plug-in, residual-memory-based paradigm for learned feedback correction in continuous-output forecasting.
Yuchen He, Yueyang Cang, Zhiyuan Ning +2
Department of Automation, Tsinghua University · State Key Laboratory of Hydroscience and Engineering, Tsinghua University
Retrieval-augmented forecasting promises to adapt frozen Time Series Foundation Models (TSFMs) to new domains without fine-tuning, but recent methods typically rely on learned fusion modules, i.e., trained adapters that merge retrieved examples into the backbone's forecast, based on the assumption that frozen backbones cannot dynamically incorporate retrieved context on their own. We show this assumption is unnecessary. We introduce Align-RAG, a training-free method that applies a closed-form per-pair amplitude rescaling and integer-lag phase shift to retrieved past-future windows before they enter a frozen backbone's context. With no learned parameters, Align-RAG outperforms the state-of-the-art trained retrieval adapter on a frozen Chronos-Bolt on all seven datasets of the standard benchmark (avg -3.75% MSE), showing that the gains previously attributed to learned fusion are recoverable without any training. Align-RAG further improves zero-shot MSE on four additional frozen TSFMs with various architectures by 2.5% to 13.7% per backbone with no per-backbone tuning. To probe why alignment helps, we compare the frozen backbone's prediction shift under aligned demonstrations to the closed-form ridge prediction shift on the same pairs. We find that aligned demonstrations induce prediction shifts that track a closed-form ridge predictor on the same pairs, with a future-shuffle control ruling out a futures-averaging account. Together, these results indicate that frozen TSFMs already support dynamic in-context use of retrievals, and that closed-form alignment should be the default baseline for retrieval-augmented forecasting before any fusion module is trained. Code available at: https://github.com/masadi-99/align-rag
Time series forecasting relies on historical patterns, but real-world series often exhibit non-stationarity and regime shifts that challenge fully parametric forecasters. Inspired by Retrieval-Augmented Generation (RAG), recent work augments forecasters by retrieving relevant historical segments and using them as external evidence at inference time. However, due to the intrinsic non-stationarity of real-world time series, a highly similar past segment does not necessarily imply a similar future, rendering similarity-only retrieval brittle and prone to redundancy. We propose Stationarity-Aware Retrieval-Augmented Time Series Forecasting (SARAF), a framework that adaptively balances relevance and diversity in retrieval. SARAF first forms a candidate pool via temporal similarity with time-aligned enhancement, then applies a diversity-aware selection strategy to cover heterogeneous historical regimes, with the diversification strength automatically modulated by dataset-level stationarity. Moreover, SARAF uses stationarity-aware aggregation to fuse the retrieved futures. Extensive experiments on eight real-world datasets show that SARAF achieves competitive forecasting performance and improves average accuracy and robustness over strong baselines, with particularly clear benefits under challenging non-stationary settings. Code: https://github.com/ShiqiaoZhou/SARAF.
Shiqiao Zhou, Holger Schöner, Zipeng Wu +3
University of Birmingham Birmingham, United Kingdom · Siemens AG Munich, Germany