Attenuated in-context identification in time-series foundation models: diagnosis under counterfactual inputs and repair by synthetic forced-system fine-tuning
Authors: Hong-In Won
Organizations: Manufacturing AI Research Center, Korea Institute of Industrial Technology, Incheon, Republic of Korea · HYU-KITECH Joint Department, Graduate School, Hanyang University, Seoul, Republic of Korea
Covariate-aware time-series foundation models (TSFMs) promise training-free what-if answers for instrumented plants: the change in output that a different future input would cause. We test this on forced engineering systems with exact counterfactuals, comparing Chronos-2, TimesFM-2.5 and TabPFN-TS with classical system identification fitted to the same context. Through their default covariate interfaces, TimesFM-2.5 and TabPFN-TS are memoryless: the predicted effect of an input change is a same-time function of that change (R2=1.000 for TimesFM-2.5). Chronos-2 identifies dynamics in context but attenuates them. Its predicted effect is 0.33-0.80 of the true effect, its recovered impulse response has the wrong shape, and its error on a one-degree-of-freedom oscillator levels off at 0.57 with 8192 context samples, where ARX fitted to 256 samples reaches 0.02. Context dither at inference lowers the what-if error on all six synthetic classes without training. A 26-minute fine-tune on synthetic forced systems restores the response magnitude (sensitivity 0.83-0.96) and outperforms structure-agnostic identification on Wiener-Hammerstein and a held-out friction class. A specialised in-context identifier trained on the same data comes close, so the forced-system data carry most of the gain. On three of four measured plants classical identification remains clearly better, and the fine-tuned model loses part of its univariate forecasting skill. Paired counterfactual inputs, together with shuffled future inputs on measured records, test two properties: whether the covariate interface can represent dynamics and whether the pretraining prior covers the plant's time scale. Only the counterfactual pairs expose the attenuation.
Figures & tables
Figure 1: One what-if query on a 1-DOF oscillator ( L=1024 ; median Chronos-2 error of 30 instances). Top: factual future input u and alternative u′ (a: step added at t=H/4 ; b: resonant sine). Bottom: true effect Δy and predicted effects Δy^ . Fine-tuned: Chronos-2 after the fine-tune of Section 5 , mean of three seeds.
Figure 2: What-if error and sensitivity of three TSFMs, the published in-context identification transformer (linear checkpoint) and the best structure-agnostic identifier (best of ARX, OE, N4SID, polynomial NARX, SE-NARX and MLP-NARX) on six synthetic classes ( L=1024 , 30 instances). Medians with 95% bootstrap intervals. † Held out from all fine-tuning.
system
TimesFM-2.5
TabPFN-TS
Chronos-2
+dither
ICL-TF
ICL-TF ∗
FT
FT-wide
best agnostic
known struct.
lag+delay
0.99
1.02
0.36
0.29
0.51
0.09
0.10 ± 0.01
0.09
0.00 (NARX)
–
1-DOF
1.00
1.00
0.67
0.36
0.25
0.26
0.17 ± 0.01
0.19
0.01 (ARX)
–
2-DOF
1.00
1.00
0.68
0.55
0.51
0.39
0.24 ± 0.01
0.27
0.10 (N4SID)
–
Duffing
0.99
1.01
0.79
0.59
0.70
0.40
0.38 ± 0.02
0.43
0.14 (SE-NARX)
0.00
Wiener–Ham.
1.00
1.00
0.80
0.49
0.69
0.22
0.20 ± 0.01
0.24
0.50 (SE-NARX)
0.00
Coulomb †
1.00
1.01
0.61
0.45
0.56
0.36
0.30 ± 0.01
0.34
0.35 (OE)
0.00
Table 1: Median what-if error Eint at L=1024 (30 instances, six alternative inputs each). ICL-TF: published in-context identification transformer [ 9 ] , better of its linear and Wiener–Hammerstein checkpoints; ICL-TF ∗ : the same architecture trained on the FT training series. +dither: Chronos-2 with context dither (Section 4 ). FT: Chronos-2 fine-tuned on synthetic forced systems (median over instances of the per-instance mean over seeds, ± half-range of the seed medians); FT-wide: FT with a time-scale-randomised prior (Section 5 ). ICL-TF and ICL-TF ∗ read the last min(L,400) context samples. Best agnostic: best of ARX, OE, N4SID, polynomial NARX, SE-NARX and MLP-NARX on the evaluation instances. Known struct.: grey-box model with the true equations. † Held out from fine-tuning. Table 4 gives L=256 .
Figure 3: Chronos-2 in-context identification (medians over a separate set of 30 instances per class, five alternatives including a pulse; Appendix A ). (a) What-if error versus context length, with ARX on the same contexts. (b) Sensitivity (thick) and impulse-response gain (thin). (c) Impulse response from a what-if pulse, 1-DOF, L=4096 , median-error instance. (d) What-if error versus context output noise ( L=1024 ).
dataset
TimesFM-2.5
TabPFN-TS
Chronos-2
+dither
ICL-TF
FT
FT-wide
ARX
OE
NARX
Wiener–Hammerstein (n=40)
1.081
1.125
0.556
0.301
0.309
0.182
0.189
0.175
0.184
0.125
shuffle gain
+0.00
+0.00
+0.68
+1.07
+1.15
+1.25
+1.24
+1.23
+1.25
+1.30
Cascaded Tanks (n=18)
1.215
1.417
1.197
0.774
1.553
0.428
0.421
0.541
0.495
0.521
shuffle gain
− 0.01
− 0.01
+0.23
+0.66
+0.16
+1.08
+1.14
+1.06
+1.45
+0.85
Silverbox (n=40)
0.962
1.081
0.544
0.506
0.097
0.629
0.443
0.077
0.095
0.039
shuffle gain
+0.01
+0.00
+0.64
+0.66
+1.18
+0.49
+0.88
+1.22
+1.23
+1.29
Table 2: Measured benchmarks: median factual NRMSE over windows ( L=1024 ; 512 for Cascaded Tanks; H=128 ) and median shuffle gain (increase of NRMSE when the future input is taken from another window). NRMSE above 1 is worse than predicting the window mean. EMPS is excluded (periodic closed-loop record; Section 2 ). FT: training seed 1.
system
L
best agnostic
FT − best
95% CI
per cell
Bonferroni
lag+delay
256
NARX
+0.13
[+0.12, +0.15]
worse
worse
1-DOF
256
ARX
+0.17
[+0.15, +0.20]
worse
worse
2-DOF
256
OE
− 0.03
[ − 0.10, +0.09]
tie
tie
Duffing
256
ARX
− 0.11
[ − 0.17, − 0.01]
better
better
Wiener–Ham.
256
ARX
− 0.26
[ − 0.35, − 0.19]
better
better
Coulomb †
256
ARX
− 0.04
[ − 0.08, − 0.01]
better
tie
Table 3: Fine-tuned Chronos-2 (seed-mean per instance) against the best structure-agnostic identifier on each class and context length: paired median difference of Eint over 30 instances with a 95% bootstrap interval. Negative favours the fine-tuned model.
Figure 4: (a) What-if error at L=1024 : Chronos-2, context dither, fine-tuned (FT, three-seed mean), FT-wide (time-scale-randomised prior) and the best structure-agnostic identifier; medians, 95% bootstrap intervals. (b) Factual NRMSE, measured benchmarks. (c) Silverbox as recorded and resampled threefold (context and horizon matched in physical time); classical reference ARX. In (b, c) FT is seed 1; log axes.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
system
TimesFM-2.5
TabPFN-TS
Chronos-2
+dither
ICL-TF
ICL-TF ∗
FT
FT-wide
best agnostic
known struct.
lag+delay
1.00
1.01
0.86
0.39
0.42
0.19
0.14 ± 0.01
0.14
0.01 (NARX)
–
1-DOF
1.01
1.00
0.84
0.45
0.25
0.27
0.21 ± 0.01
0.24
0.03 (ARX)
–
2-DOF
1.00
1.00
0.89
0.62
0.51
0.44
0.35 ± 0.02
0.37
0.37 (OE)
–
Duffing
1.00
1.01
0.89
0.63
0.75
0.52
0.50 ± 0.01
0.54
0.62 (ARX)
0.01
Wiener–Ham.
1.01
1.00
0.93
0.56
0.79
0.31
0.27 ± 0.01
0.29
0.54 (ARX)
0.00
Coulomb †
1.00
1.00
0.77
0.53
0.61
0.39
0.38 ± 0.01
0.42
0.48 (ARX)
0.01
Appendix
Table 4: As Table 1 at L=256 .
model
system
×0.5
×2
zero
step
resonant
fresh
Chronos-2
lag+delay
0.37/0.85
0.52/0.70
0.32/0.87
0.37/0.66
0.38/0.70
0.30/0.86
Chronos-2
1-DOF
0.68/0.57
0.78/0.35
0.66/0.65
0.66/0.48
0.66/0.39
0.61/0.54
Chronos-2
2-DOF
0.71/0.67
0.84/0.41
0.68/0.65
0.53/0.66
0.67/0.43
0.61/0.58
Chronos-2
Duffing
0.82/0.49
0.96/0.33
0.75/0.55
0.69/0.46
0.71/0.53
0.67/0.50
Chronos-2
Wiener–Ham.
0.82/0.42
0.95/0.20
0.75/0.48
0.83/0.24
0.76/0.35
0.76/0.35
Chronos-2
Coulomb †
0.64/0.57
0.80/0.33
0.50/0.75
0.48/0.68
0.64/0.41
0.53/0.70
Appendix
Table 5: What-if error / sensitivity per alternative input ( L=1024 , medians over 30 instances).