RACE: Residual-Aware Test-Time Adaptation for Neighbor-Rich Time-Series Foundation Model Forecasting
Organizations: School of Artificial Intelligence, Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · Nanjing, China
Abstract
Time-series foundation models (TSFMs) perform strongly across forecasting tasks, but their per-series inference is ill-suited to neighbor-rich forecasting, where each query has access to related but nonidentical historical series. Continuous glucose monitoring (CGM) and Web/cloud workloads exemplify this setting: CGM trajectories share physiological patterns but vary across individuals, devices, and conditions, while Web/cloud workloads combine common operating regimes with non-stationarity, heavy tails, and bursts. These histories share useful structure, yet neighbors are not equally relevant. Existing methods either fine-tune TSFMs for each target domain, incurring additional costs and offering limited transferability across backbones, or append retrieved series without verifying whether they support the current forecast. The key challenges are conflicting residual evidence from neighboring series and residual patterns that vary across TSFMs and forecasting tasks. We formulate test-time neighborhood scaling: using same-domain neighbor evidence without modifying the backbone. We propose RACE (Residual-Aware Correction of Forecasting Errors), a two-stage framework for using historical neighbors. We first retrieve query-compatible neighbors, align their residuals to the query scale, and aggregate coherent evidence into the training-free RACE-TF correction. Full RACE then uses a lightweight, domain-specific Gate to determine when applying the correction is beneficial, with a reusable training workflow across TSFM backbones. Across four TSFMs, RACE improves all three domain-aggregate metrics on both primary domains, with the largest gains on high-error queries. Within each domain, a Gate trained on one TSFM transfers to other backbones without adaptation, and the resulting pipeline improves all 72 cross-backbone metric comparisons over the matched frozen targets.
Figures & tables
| TSFM + correction | Method status | Web: MSE / MAE / CRPS | CGM: MSE / MAE / CRPS | Best Count |
| Chronos-2 | frozen base | 8.676e+05 / 70.35 / 51.27 | 2.925 / 1.149 / 0.9099 | – |
| + RAF | training-free | +2.503 / +2.423 / +0.216 | +6.826 / +4.869 / +4.585 | 2 |
| + RACE-TF (ours) | training-free | +1.159 / +0.099 / -0.186 | +20.752 / +10.531 / +9.849 | 0 |
| + RACE (ours) | fitted (XGB) | +0.982 / +1.091 / +1.182 | +20.814 / +10.677 / +9.939 | 4 |
| Chronos-Bolt Base | frozen base | 9.277e+05 / 74.87 / 59.49 | 3.170 / 1.233 / 0.9708 | – |
| + RAF | training-free | +5.527 / +1.230 / +1.244 | +6.923 / +4.904 / +4.394 | 1 |
| Source Gate | Web/cloud | CGM (ShanghaiT2DM target) | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Chronos-2 Base | Chronos-Bolt Base | TimesFM 2.5 200M | Toto-2.0- 313m | Gate uplift | Chronos-2 Base | Chronos-Bolt Base | TimesFM 2.5 200M | Toto-2.0- 313m | Gate uplift | |
| Chronos-2 Base | +1.09 | +1.94 | +1.98 | +1.04 | +0.74 | +13.81 | +15.46 | +12.99 | +14.68 | +0.10 |
| Chronos-Bolt Base | +0.88 | +1.81 | +1.21 | +0.87 | +0.35 | +13.74 | +15.46 | +12.89 | +14.68 | +0.03 |
| TimesFM 2.5 200M | +1.09 | +1.93 | +1.97 | +1.02 | +0.78 | +13.63 | +15.44 | +12.95 | +14.47 | -0.04 |
| Toto-2.0-313m | +1.06 | +1.92 | +1.98 | +0.99 | +0.67 | +13.72 | +15.44 | +12.95 | +14.54 | +0.06 |
| Variant | Change | Web (%) | CGM (%) |
|---|---|---|---|
| RACE | None | +1.410 / +1.431 / +1.557 | +21.327 / +11.200 / +10.043 |
| Panel A: evidence-construction ablations (Gate retained) | |||
| Global bias + Gate | Without local retrieval | -0.040 / -0.945 / -1.332 | +3.802 / +1.320 / +1.369 |
| History-only retrieval | Without forecast | +1.214 / +1.209 / +1.336 | +16.484 / +8.943 / +8.021 |
| Residual mean + Gate | Without consensus | +2.247 / +0.984 / +0.917 | +21.716 / +10.154 / +9.086 |
| Panel B: Gate ablation and combined control | |||
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Symbols | Meaning |
|---|---|
| Windows and forecasts | is a query window and a historical window in ’s eligible residual pool. is the frozen time-series foundation model (TSFM); are visible contexts; are realized futures; and are the base, neighbor, and deployed forecasts. |
| Legal retrieval | is the eligible residual pool and its final subset of at most query-compatible neighbors. is the initial exact-ranking width; a short query is expanded exactly after legality filtering. A parent is an original full series; a same-parent candidate’s future must end no later than the query future starts: . |
| Retrieval representations | is the normalized visible history–forecast trajectory. For , , , and are feature extractor , its layer-normalization transform, and its fixed weight. concatenates vectors, is the Euclidean norm, is the Euclidean inner product, and are query and neighbor representations; is the HybridMath unit-norm representation. |
| Residual scale alignment | is the realized neighbor residual. are mean absolute observed-context magnitudes, floored at ; clamps to , and is the dimensionless, clipped residual aligned to the query scale. |
| Evidence sets | denote valid neighbors, coordinate-available neighbors, and usable coordinates. and are the initial and coordinate-specific retained sets; indexes a forecast-horizon coordinate and channel. |
| Evidence summaries and screening | are the residual center, coordinate deviation, signed mean deviation, and within-neighbor standard deviation. are robust centers and median-absolute-deviation scales; is the fixed MAD consistency coefficient; is an across-neighbor scalar collection and one entry. are MAD-standardized displacement and one-sided variability-excess scores; is their fixed screening threshold. |
| Frozen TSFM | MSE gain | MAE gain | CRPS gain | Acceptance rate |
|---|---|---|---|---|
| Chronos-2 | +0.168 | +0.052 | -0.089 | 36.15% |
| Chronos-Bolt Base | +3.533 | +1.409 | +0.955 | 62.26% |
| TimesFM-2.5 | +0.010 | +0.086 | +0.122 | 13.79% |
| Toto-2.0-313m | +4.724 | +1.267 | +0.971 | 41.29% |
| Four-TSFM mean | +2.109 | +0.704 | +0.490 | 38.37% |
| Source Gate | Target TSFM | MSE gain (%) | MAE gain (%) | CRPS gain (%) | Mean Gate uplift (pp) |
|---|---|---|---|---|---|
| Chronos-2 | Chronos-Bolt Base | +1.633 | +2.022 | +2.164 | +0.746 |
| Chronos-2 | TimesFM-2.5 | +1.982 | +1.767 | +2.205 | +0.578 |
| Chronos-2 | Toto-2.0-313m | +1.035 | +1.101 | +0.984 | +0.888 |
| Chronos-Bolt Base | Chronos-2 | +1.019 | +0.804 | +0.829 | +0.526 |
| Chronos-Bolt Base | TimesFM-2.5 | +1.189 | +1.091 | +1.339 | -0.200 |
| Chronos-Bolt Base | Toto-2.0-313m | +1.116 | +0.849 | +0.632 | +0.714 |
| Source Gate | Target TSFM | MSE gain (%) | MAE gain (%) | CRPS gain (%) | Mean Gate uplift (pp) |
|---|---|---|---|---|---|
| Chronos-2 | Chronos-Bolt Base | +22.801 | +12.895 | +10.672 | +0.129 |
| Chronos-2 | TimesFM-2.5 | +19.862 | +9.879 | +9.244 | +0.117 |
| Chronos-2 | Toto-2.0-313m | +22.235 | +11.364 | +10.426 | +0.048 |
| Chronos-Bolt Base | Chronos-2 | +20.708 | +10.609 | +9.895 | +0.027 |
| Chronos-Bolt Base | TimesFM-2.5 | +19.684 | +9.818 | +9.177 | +0.015 |
| Chronos-Bolt Base | Toto-2.0-313m | +22.228 | +11.388 | +10.418 | +0.051 |
| Group | Configuration and changed factor | Web (%) | CGM (%) |
|---|---|---|---|
| MSE / MAE / CRPS | MSE / MAE / CRPS | ||
| Deployed route | RACE: raw-z representation, Consensus Mean, , Gate with MSE target | +1.410 / +1.431 / +1.557 | +21.327 / +11.200 / +10.043 |
| Encoder sensitivity | HybridMath representation replaces raw-z | +1.807 / +1.256 / +1.337 | +20.637 / +10.725 / +9.300 |
| Query control | Chronos-T5 final-token representation; history-only query | -0.071 / -0.312 / -0.406 | +6.330 / +2.145 / +1.815 |
| Query control | History only; frozen forecast removed from raw-z representation | +1.214 / +1.209 / +1.336 | +16.484 / +8.943 / +8.021 |
| Neighborhood | replaces | +1.486 / +1.035 / +1.052 | +14.336 / +6.549 / +5.735 |
| Task | Frozen TSFM | RACE-TF (all queries) | CERR (direct learned) | RACE (Gate-selected) |
|---|---|---|---|---|
| MSE / MAE / CRPS | MSE / MAE / CRPS | MSE / MAE / CRPS | ||
| Web | Chronos-2 | +1.159 / +0.099 / -0.186 | +1.445 / -8.179 / -12.516 | +0.982 / +1.091 / +1.182 |
| Web | Chronos-Bolt Base | +1.782 / +0.954 / +0.846 | +4.290 / -3.513 / -4.963 | +1.647 / +1.838 / +1.945 |
| Web | TimesFM-2.5 | +2.342 / +0.845 / +1.033 | +2.826 / -4.014 / -4.782 | +1.986 / +1.748 / +2.187 |
| Web | Toto-2.0-313m | +1.301 / -0.126 / -0.719 | -2.226 / -10.163 / -13.909 | +1.025 / +1.047 / +0.912 |
| Web | Four-TSFM mean | +1.646 / +0.443 / +0.243 | +1.584 / -6.467 / -9.042 | +1.410 / +1.431 / +1.557 |
| Group | Configuration and changed factor | Web (%) |
|---|---|---|
| MSE / MAE / CRPS | ||
| Scope diagnostic | RACE with raw-z representation; same logical-subset eligible residual pool only | -0.277 / +0.378 / +0.288 |
| Scope diagnostic | RACE with HybridMath representation; same logical-subset eligible residual pool only | +0.106 / +0.501 / +0.523 |
| Scope diagnostic | RACE with raw-z representation; hard Top-3 subset-card eligible residual pool | -0.769 / +0.016 / -0.198 |
| Scope diagnostic | RACE with HybridMath representation; hard Top-3 subset-card eligible residual pool | +0.938 / +1.004 / +1.056 |
| Unconditional transfer | Same-subset mean, Web only; no consensus or Gate | -39.280 / -26.983 / -35.841 |
| Task | Frozen TSFM | RACE-TF beneficial (%) | Acceptance rate (%) | Benefit among accepted (%) | AP | Spearman | |
|---|---|---|---|---|---|---|---|
| CGM | Chronos-2 | 0.514 | 68.09 | 95.25 | 69.05 | 0.773 | 0.336 |
| CGM | Chronos-Bolt Base | 0.601 | 69.91 | 95.74 | 71.63 | 0.780 | 0.400 |
| CGM | TimesFM-2.5 | 0.581 | 66.35 | 90.55 | 68.32 | 0.744 | 0.306 |
| CGM | Toto-2.0-313m | 0.644 | 69.42 | 90.67 | 70.96 | 0.787 | 0.368 |
| Web | Chronos-2 | 0.453 | 37.85 | 52.02 | 52.69 | 0.548 | 0.520 |
| Web | Chronos-Bolt Base | 0.564 | 40.23 | 51.75 | 53.48 | 0.529 | 0.421 |
| Frozen TSFM | Error bin | RACE MSE reduction (%) | RACE-TF MSE reduction (%) | Acceptance rate (%) | Benefit among accepted (%) | |
|---|---|---|---|---|---|---|
| Chronos-2 | 1 | -0.02 | -30.99 | 34.48 | 36.34 | 4820 |
| Chronos-2 | 2 | -0.12 | -60.11 | 42.13 | 50.89 | 4814 |
| Chronos-2 | 3 | -0.52 | -51.59 | 49.74 | 53.86 | 4815 |
| Chronos-2 | 4 | +0.07 | -26.84 | 67.30 | 57.13 | 4814 |
| Chronos-2 | 5 | +60.92 | +42.34 | 66.48 | 56.96 | 4812 |
| Chronos-Bolt Base | 1 | -0.50 | -35.68 | 39.98 | 34.61 | 4820 |
| Frozen TSFM | Error bin | RACE MSE reduction (%) | RACE-TF MSE reduction (%) | Acceptance rate (%) | Benefit among accepted (%) | |
|---|---|---|---|---|---|---|
| Chronos-2 | 1 | -1.74 | -1.77 | 96.36 | 47.80 | 165 |
| Chronos-2 | 2 | +1.21 | +1.02 | 94.51 | 65.16 | 164 |
| Chronos-2 | 3 | +7.20 | +6.91 | 92.07 | 69.54 | 164 |
| Chronos-2 | 4 | +22.32 | +21.96 | 98.17 | 80.75 | 164 |
| Chronos-2 | 5 | +75.22 | +75.76 | 95.12 | 82.05 | 164 |
| Chronos-Bolt Base | 1 | -4.08 | -4.16 | 95.76 | 43.04 | 165 |
| Method | MSE W/L | MAE W/L | CRPS W/L |
|---|---|---|---|
| RACE-TF | 21/19 | 10/30 | 11/29 |
| RACE | 26/14 | 17/23 | 17/23 |
| RAF | 13/27 | 4/36 | 8/32 |
| Feature block (dimension) | Exact construction | Decision information |
|---|---|---|
| Evidence availability (2) | Consensus-retained valid-neighbor count and valid-neighbor fraction. | Amount of legal, consensus-retained evidence. |
| Aligned residual evidence (9) | Mean and standard deviation of values and of their absolute values, median, and root-mean-square over valid aligned residual coordinates; mean and standard deviation of within-neighbor residual standard deviations; and . | Residual level, magnitude, variability, and temporal coherence. |
| Retrieval similarity (8) | Mean, standard deviation, minimum, maximum, 10th/50th/90th percentiles, and Gini coefficient of over valid neighbors. | Closeness and concentration of retrieved support. |
| Frozen forecast shape (2) | Standard deviation and mean absolute first difference of the query-scale-normalized frozen forecast, . | Volatility of the backbone forecast. |
| Observed context shape (2) | Standard deviation and mean absolute first difference of observed . | Recent target dynamics in comparable units. |
| Consensus proposal (3) | Mean absolute value, standard deviation, and maximum absolute value of on . | Size and temporal variation of the proposed correction. |
| Method | Measurement boundary | Chronos-2 | Chronos-Bolt Base | TimesFM-2.5 | Toto-2.0-313m |
| Web / CGM (s) | Web / CGM (s) | Web / CGM (s) | Web / CGM (s) | ||
| RAF | Encoding, Top-1 retrieval, augmented-input forecasting, evaluation, and audit output. | 38.7 / 19.6 | 50.3 / 20.0 | 208.0 / 43.2 | 95.6 / 49.1 |
| RACE-TF | Post-cache route replay; residual evidence construction and one final test pass; no fitted Gate. | 91.0 / 8.7 | 93.2 / 8.4 | 92.8 / 8.5 | 93.7 / 8.7 |
| RACE | Post-cache route replay; residual evidence construction, Gate fitting, model-select, calibration, and one final test pass. | 156.3 / 19.5 | 156.7 / 19.5 | 157.2 / 19.6 | 170.8 / 20.7 |
| TS-RAG (PT) | Frozen supplied ARM; Chronos-Bolt Base interface only. | – | 134.2 / 9.2 | – | – |
| TS-RAG (FT) | Fine-tuned ARM; Chronos-Bolt Base interface only. | – | 814.0 / 178.7 | – | – |