Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning
Authors: Hong-In Won
Organizations: Manufacturing AI Research Center, Korea Institute of Industrial Technology, Incheon, Republic of Korea · HYU-KITECH Joint Department, Graduate School, Hanyang University, Seoul, Republic of Korea
Covariate-aware time-series foundation models (TSFMs) promise training-free what-if answers for instrumented plants: the change in output that a different future input would cause. We test this on forced engineering systems with exact counterfactuals, comparing Chronos-2, TimesFM-2.5 and TabPFN-TS with classical system identification fitted to the same context. Through their default covariate interfaces, TimesFM-2.5 and TabPFN-TS are memoryless: the predicted effect of an input change is a same-time function of that change (R2=1.000 for TimesFM-2.5). Chronos-2 identifies dynamics in context but attenuates them. Its predicted effect is 0.33-0.80 of the true effect, its recovered impulse response has the wrong shape, and its error on a one-degree-of-freedom oscillator levels off at 0.57 with 8192 context samples, where ARX fitted to 256 samples reaches 0.02. Context dither at inference lowers the what-if error on all six synthetic classes without training. A 26-minute fine-tune on synthetic forced systems restores the response magnitude (sensitivity 0.83-0.96) and outperforms structure-agnostic identification on Wiener-Hammerstein and a held-out friction class. A specialised in-context identifier trained on the same data comes close, so the forced-system data carry most of the gain. On three of four measured plants classical identification remains clearly better, and the fine-tuned model loses part of its univariate forecasting skill. Paired counterfactual inputs, together with shuffled future inputs on measured records, test two properties: whether the covariate interface can represent dynamics and whether the pretraining prior covers the plant's time scale. Only the counterfactual pairs expose the attenuation.
Figures & tables
Figure 1: One what-if query on a 1-DOF oscillator ( L=1024 ; median Chronos-2 error of 30 instances). Top: factual future input u and alternative u′ (a: step added at t=H/4 ; b: resonant sine). Bottom: true effect Δy and predicted effects Δy^ . Fine-tuned: Chronos-2 after the fine-tune of Section 5 , mean of three seeds.
Figure 2: What-if error and sensitivity of three TSFMs, the published in-context identification transformer (linear checkpoint) and the best structure-agnostic identifier (best of ARX, OE, N4SID, polynomial NARX, SE-NARX and MLP-NARX) on six synthetic classes ( L=1024 , 30 instances). Medians with 95% bootstrap intervals. † Held out from all fine-tuning.
system
TimesFM-2.5
TabPFN-TS
Chronos-2
+dither
ICL-TF
ICL-TF ∗
FT
FT-wide
best agnostic
known struct.
lag+delay
0.99
1.02
0.36
0.29
0.51
0.09
0.10 ± 0.01
0.09
0.00 (NARX)
–
1-DOF
1.00
1.00
0.67
0.36
0.25
0.26
0.17 ± 0.01
0.19
0.01 (ARX)
–
2-DOF
1.00
1.00
0.68
0.55
0.51
0.39
0.24 ± 0.01
0.27
0.10 (N4SID)
–
Duffing
0.99
1.01
0.79
0.59
0.70
0.40
0.38 ± 0.02
0.43
0.14 (SE-NARX)
0.00
Wiener–Ham.
1.00
1.00
0.80
0.49
0.69
0.22
0.20 ± 0.01
0.24
0.50 (SE-NARX)
0.00
Coulomb †
1.00
1.01
0.61
0.45
0.56
0.36
0.30 ± 0.01
0.34
0.35 (OE)
0.00
Table 1: Median what-if error Eint at L=1024 (30 instances, six alternative inputs each). ICL-TF: published in-context identification transformer [ 9 ] , better of its linear and Wiener–Hammerstein checkpoints; ICL-TF ∗ : the same architecture trained on the FT training series. +dither: Chronos-2 with context dither (Section 4 ). FT: Chronos-2 fine-tuned on synthetic forced systems (median over instances of the per-instance mean over seeds, ± half-range of the seed medians); FT-wide: FT with a time-scale-randomised prior (Section 5 ). ICL-TF and ICL-TF ∗ read the last min(L,400) context samples. Best agnostic: best of ARX, OE, N4SID, polynomial NARX, SE-NARX and MLP-NARX on the evaluation instances. Known struct.: grey-box model with the true equations. † Held out from fine-tuning. Table 4 gives L=256 .
Figure 3: Chronos-2 in-context identification (medians over a separate set of 30 instances per class, five alternatives including a pulse; Appendix A ). (a) What-if error versus context length, with ARX on the same contexts. (b) Sensitivity (thick) and impulse-response gain (thin). (c) Impulse response from a what-if pulse, 1-DOF, L=4096 , median-error instance. (d) What-if error versus context output noise ( L=1024 ).
dataset
TimesFM-2.5
TabPFN-TS
Chronos-2
+dither
ICL-TF
FT
FT-wide
ARX
OE
NARX
Wiener–Hammerstein (n=40)
1.081
1.125
0.556
0.301
0.309
0.182
0.189
0.175
0.184
0.125
shuffle gain
+0.00
+0.00
+0.68
+1.07
+1.15
+1.25
+1.24
+1.23
+1.25
+1.30
Cascaded Tanks (n=18)
1.215
1.417
1.197
0.774
1.553
0.428
0.421
0.541
0.495
0.521
shuffle gain
− 0.01
− 0.01
+0.23
+0.66
+0.16
+1.08
+1.14
+1.06
+1.45
+0.85
Silverbox (n=40)
0.962
1.081
0.544
0.506
0.097
0.629
0.443
0.077
0.095
0.039
shuffle gain
+0.01
+0.00
+0.64
+0.66
+1.18
+0.49
+0.88
+1.22
+1.23
+1.29
Table 2: Measured benchmarks: median factual NRMSE over windows ( L=1024 ; 512 for Cascaded Tanks; H=128 ) and median shuffle gain (increase of NRMSE when the future input is taken from another window). NRMSE above 1 is worse than predicting the window mean. EMPS is excluded (periodic closed-loop record; Section 2 ). FT: training seed 1.
system
L
best agnostic
FT − best
95% CI
per cell
Bonferroni
lag+delay
256
NARX
+0.13
[+0.12, +0.15]
worse
worse
1-DOF
256
ARX
+0.17
[+0.15, +0.20]
worse
worse
2-DOF
256
OE
− 0.03
[ − 0.10, +0.09]
tie
tie
Duffing
256
ARX
− 0.11
[ − 0.17, − 0.01]
better
better
Wiener–Ham.
256
ARX
− 0.26
[ − 0.35, − 0.19]
better
better
Coulomb †
256
ARX
− 0.04
[ − 0.08, − 0.01]
better
tie
Table 3: Fine-tuned Chronos-2 (seed-mean per instance) against the best structure-agnostic identifier on each class and context length: paired median difference of Eint over 30 instances with a 95% bootstrap interval. Negative favours the fine-tuned model.
Figure 4: (a) What-if error at L=1024 : Chronos-2, context dither, fine-tuned (FT, three-seed mean), FT-wide (time-scale-randomised prior) and the best structure-agnostic identifier; medians, 95% bootstrap intervals. (b) Factual NRMSE, measured benchmarks. (c) Silverbox as recorded and resampled threefold (context and horizon matched in physical time); classical reference ARX. In (b, c) FT is seed 1; log axes.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
system
TimesFM-2.5
TabPFN-TS
Chronos-2
+dither
ICL-TF
ICL-TF ∗
FT
FT-wide
best agnostic
known struct.
lag+delay
1.00
1.01
0.86
0.39
0.42
0.19
0.14 ± 0.01
0.14
0.01 (NARX)
–
1-DOF
1.01
1.00
0.84
0.45
0.25
0.27
0.21 ± 0.01
0.24
0.03 (ARX)
–
2-DOF
1.00
1.00
0.89
0.62
0.51
0.44
0.35 ± 0.02
0.37
0.37 (OE)
–
Duffing
1.00
1.01
0.89
0.63
0.75
0.52
0.50 ± 0.01
0.54
0.62 (ARX)
0.01
Wiener–Ham.
1.01
1.00
0.93
0.56
0.79
0.31
0.27 ± 0.01
0.29
0.54 (ARX)
0.00
Coulomb †
1.00
1.00
0.77
0.53
0.61
0.39
0.38 ± 0.01
0.42
0.48 (ARX)
0.01
Appendix
Table 4: As Table 1 at L=256 .
model
system
×0.5
×2
zero
step
resonant
fresh
Chronos-2
lag+delay
0.37/0.85
0.52/0.70
0.32/0.87
0.37/0.66
0.38/0.70
0.30/0.86
Chronos-2
1-DOF
0.68/0.57
0.78/0.35
0.66/0.65
0.66/0.48
0.66/0.39
0.61/0.54
Chronos-2
2-DOF
0.71/0.67
0.84/0.41
0.68/0.65
0.53/0.66
0.67/0.43
0.61/0.58
Chronos-2
Duffing
0.82/0.49
0.96/0.33
0.75/0.55
0.69/0.46
0.71/0.53
0.67/0.50
Chronos-2
Wiener–Ham.
0.82/0.42
0.95/0.20
0.75/0.48
0.83/0.24
0.76/0.35
0.76/0.35
Chronos-2
Coulomb †
0.64/0.57
0.80/0.33
0.50/0.75
0.48/0.68
0.64/0.41
0.53/0.70
Appendix
Table 5: What-if error / sensitivity per alternative input ( L=1024 , medians over 30 instances).
Longer histories can improve time-series foundation models (TSFMs), but require substantially higher inference cost. We therefore ask whether contextual information can be provided more efficiently through a compact set of learned token embeddings. We introduce PaCTS, which generates a small set of instance-adaptive latent prompts in the form of continuous embedding tokens conditioned on the visible context. These prompts serve as compact context surrogates for frozen TSFMs. PaCTS constructs them from instance-specific global statistics and further refines them with segment-level temporal information, capturing both global characteristics and local temporal variations. The prompt module is jointly trained and deployed across heterogeneous time series with the frozen backbone. Extensive experiments demonstrate the effectiveness of prompts as context, consistently improving forecasting across context lengths and model architectures. With a shorter input context, PaCTS can outperform the same frozen backbone using double context while requiring substantially less inference computation. Compared with weight-space adaptation methods, PaCTS achieves stronger improvements and better out-of-distribution generalization.
Zehao Xiao, Shifeng Xie, Lei Zan +5
Huawei Noah’s Ark Lab, Paris, France · LIPADE, Universit´e Paris Cit´e, Paris, France · Huawei Noah’s Ark Lab, Shenzhen, China
This work studies a central gap in interpreting time-series foundation models (TSFMs): a dynamical property may be accessible in a hidden state even when the forecast fails to respond correctly as that property changes. We formalize these properties as Dynamical Parameters, including trend slope, oscillation frequency, and autoregressive dependence. We compare their representation accessibility, measured by recovery from hidden states, with their forecast response, measured by agreement with the expected forecast change. Across nine frozen TSFMs and thirteen laws, 42 of 63 model-parameter cells achieve accessibility above 0.95, whereas their median reference-aligned response relative to the conditional reference is only 0.46. To explain this gap, causal geometry compares the hidden-state change required to produce the reference response with the change induced by the parameter intervention. Directly modifying the hidden state recovers the reference response, but the parameter intervention often moves the state in a different direction. These results show that accessible parameter information need not be expressed in forecasts when input changes miss the required hidden-state direction.
Kang Yang, Gaofeng Dong, Liying Han +1
Department of Electrical and Computer Engineering University of California, Los Angeles
Despite the success of Time Series Foundation Models (TSFMs) on broad benchmarks, their ability to internalize basic temporal logic, especially in settings supported by exogenous covariates, remains under-examined. We introduce SimpleTimeBench, a diagnostic univariate and multivariate "unit test" suite for primitives such as monotonic trends, periodic signals and leading indicator covariates, scenarios where near-perfect forecasts should be trivial. Surprisingly, prominent multivariate TSFMs (Chronos-2, Moirai and Toto) frequently produce suboptimal zero-shot forecasts for these inputs. While fine-tuning Chronos-2 improves its behaviour on specific tasks, we show that this adaptation degrades performance on other fundamental patterns rather than enhancing its generalizable foundational capabilities. This reveals a gap between pre-training scale and basic temporal reasoning, suggesting that current TSFMs could potentially lack the inductive biases needed to capture simple predictable functions. We further demonstrate that these failures are not merely synthetic curiosities: they persist in real-world sensor forecasting, where TSFMs consistently underutilize leading indicators available in observed covariates. This inability to capture simple relationships limits the practical utility and reliability of current multivariate models.