Digital twins increasingly support downstream analytical tasks that depend on time-series data, motivating interest in time-series foundation models (TSFMs) as scalable backbones. However, TSFMs are primarily pretrained for temporal continuation and often underperform on unseen tasks such as regression, and systematic empirical comparisons against state-of-the-art dedicated models in digital twin contexts remain limited. This paper makes three contributions. First, we benchmark five well-known TSFMs with frozen backbones on remaining useful life (RUL) prediction using the C-MAPSS dataset, finding that multivariate architectures substantially outperform univariate ones, particularly under varying operating conditions. This raises a deeper question: when cross-channel dependencies can be modeled through pretrained weights, target-task adaptation, and digital twin-derived representations, how much does each contribute, and are they complementary? Second, we propose a topology-informed fusion approach in which topological constraints, derived from the asset structure the digital twin stores among its information models, explicitly shape cross-attention, so that fused representations respect the physical system's local connectivity rather than relying on unconstrained all-to-all interactions. Third, we conduct an ablation study across C-MAPSS subsets of varying operational complexity that isolates the three sources and their interactions. The sources prove complementary rather than redundant, and topology-constrained attention outperforms unconstrained fusion, though by a small margin, enabling a frozen TSFM informed by digital twin representations to remain competitive or in some cases exceed state-of-the-art performance on this regression task.
Figures & tables
Figure 1 : Digital twin framework for standardized development.
Model
Architecture
Tokenization
Variate Handling
Tasks
MOMENT
Encoder-only
Patch-wise
Univariate
F, I, D, C
Moirai
Encoder-only
Patch-wise
Multivariate
F
Lag-Llama
Decoder-only
Point-wise
Univariate
Prob. F
Chronos
Encoder-Decoder
Quantization
Univariate
Prob. F
TimesFM
Decoder-only
Patch-wise
Univariate
F
Table 1 : Comparison of TSFMs: MOMENT, Moirai, Lag-Llama, Chronos, and TimesFM.
Figure 2 : System logic block diagram [ 8 ] and corresponding turbofan cross-sectional schematic for large-scale commercial turbofan engine (90,000 lb thrust class) [ 19 ] .
Subset
FD001
FD002
FD003
FD004
Instances in training set
100
260
100
249
Instances in testing set
100
259
100
248
Fault modes
1
1
2
2
Operating conditions a
1 b
6 c
1 b
6 c
Max/min cycle for train
362/128
378/128
525/145
543/128
Max/min cycle for test
303/31
367/21
475/38
486/19
Table 2 : Description of the C-MAPSS subsets.
Index
Symbol
Description [Unit]
Component
1
T2
Fan Inlet Total Temp. [°R]
/
2
T24
LPC Outlet Total Temp. [°R]
LPC
3
T30
HPC Outlet Total Temp. [°R]
HPC
4
T50
LPT Outlet Total Temp. [°R]
LPT
5
P2
Fan Inlet Pressure [psia]
/
6
P15
Bypass Duct Total Press. [psia]
/
Table 3 : Description of the C-MAPSS sensors and component mapping.
Figure 3 : Overview of the unified evaluation pipeline for benchmarking TSFMs.
Parameter
MOMENT
Moirai
Lag-Llama
Chronos
TimesFM
Pretrained Checkpoint
MOMENT-1-small
moirai-1.1-R-small
lag-llama.ckpt
chronos-t5-small
timesfm-2.5-200m
Embedding Dimension
512
384
144
512
1,280
Context Length
512 (model-defined)
64 (user-defined)
64 (user-defined)
Variable (user-defined)
64 (user-defined)
Patch Size
8 (model-defined)
8 (user-defined)
N/A (point-wise)
N/A (tokenized)
32 (model-defined)
Padding Strategy
Right-pad
Left-pad
Left-pad
None
Left-pad
Frozen Parameters
35.3M
13.8M
2.4M
46.2M
231.3M
Table 4 : Structural and parameter comparison of TSFMs for the benchmark.
Figure 4 : Hierarchical topological graph illustrated for aviation (left), manufacturing (center), and energy (right) domains. The aviation instantiation corresponds to the C-MAPSS turbofan with LPT-associated ( c6 ) edges highlighted; manufacturing and energy columns depict schematic representations of typical component and sensor types from respective fields [ 30 , 34 ] .
Asset structure
FD001 ( w=30 )
FD002 ( w=15 )
FD003 ( w=30 )
FD004 ( w=15 )
Ground truth
RMSE: 13.12±0.49
RMSE: 24.08±1.36
RMSE: 13.37±0.85
RMSE: 24.31±0.23
NRMSE: 0.1112±0.0042
NRMSE: 0.2024±0.0114
NRMSE: 0.1124±0.0071
NRMSE: 0.2043±0.0019
Score: 269±39
Score: 8071±2759
Score: 349±102
Score: 9167±2531
Incorrect component coupling
RMSE: 13.18±0.80
RMSE: 24.57±1.12
RMSE: 13.26±0.88
RMSE: 25.04±0.34
NRMSE: 0.1117±0.0068
NRMSE: 0.2064±0.0094
NRMSE: 0.1114±0.0074
NRMSE: 0.2104±0.0029
Score: 281±47
Score: 8573±2102
Score: 319±71
Score: 9027±2925
Table 6: Performance (mean ± std.) of the GNN-alone configuration under the ground truth sensor-component structure and under the two alternatives, defined in Section 6.1 , across C-MAPSS subsets.
Pretrain
Target Task
Digital Twin Derived
FD001 ( w=30 )
FD002 ( w=15 )
FD003 ( w=30 )
FD004 ( w=15 )
MOMENT
RMSE: 12.69±0.78
RMSE: 23.17±0.96
RMSE: 12.95±1.04
RMSE: 24.39±0.53
✗
✓
✓
NRMSE: 0.1075±0.0066
NRMSE: 0.1947±0.0081
NRMSE: 0.1088±0.0087
NRMSE: 0.2050±0.0045
Score: 257±54
Score: 5980±1585
Score: 385±130
Score: 7541±2238
RMSE: 12.94±0.52
RMSE: 24.74±1.16
RMSE: 13.10±0.74
RMSE: 25.62±0.90
✗
✓
❍
NRMSE: 0.1097±0.0044
NRMSE: 0.2079±0.0098
NRMSE: 0.1101±0.0062
NRMSE: 0.2153±0.0075
Table 7 : Performance (mean ± std.) of MOMENT and Moirai configurations across C-MAPSS subsets.
Figure 5 : Baseline NRMSE comparison between MOMENT and Moirai in univariate encoding mode across C-MAPSS subsets.
Figure 6 : Marginal effect of cross-channel dependencies for Moirai learnt from large-scale pretraining.
Figure 7 : Marginal effect of target-task cross-channel dependencies for (a) MOMENT and (b) Moirai.
Figure 8 : Marginal effect of digital twin fusion for (a) MOMENT and (b) Moirai.
Figure 9 : NRMSE reduction relative to GNN-alone (baseline) under masked digital twin fusion.
Figure 10 : NRMSE reduction relative to GNN-alone (baseline) under masked digital twin fusion for (a) MOMENT and (b) Moirai.
FD001
FD002
FD003
FD004
Method
RMSE
Score
RMSE
Score
RMSE
Score
RMSE
Score
SVR [ 53 ]
20.96
1382
42.00
589900
21.05
1598
45.35
37114
CNN [ 53 ]
18.45
1287
30.29
13570
19.82
1596
29.16
7886
LSTM [ 79 ]
16.14
338
24.49
4450
16.18
852
28.17
5550
AGCNN [ 39 ]
12.42
226
19.43
1492
13.39
227
21.50
3392
MCLSTM [ 67 ]
13.71
315
—
—
—
—
23.81
4826
Table 8: Comparison of proposed method against state-of-the-art RUL prediction baselines. For the proposed method, the value reported is the mean over the ten folds, under the protocol stated in Section 4.1 ; baseline values are cited from existing works. Bold marks the best value in each column; the proposed method is shown in the shaded row.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Pretraining corpus
Domains
Familiarity level
MOMENT [ 24 ]
Time Series Pile: Informer, Monash, UCR/UEA, TSB-UAD
Monash, M-competitions, Kaggle; plus synthetic (TSMixup, KernelSynth)
Energy, transport, climate, healthcare, retail, web
Modality
TimesFM [ 12 ]
Wikipedia, Google Trends, M4, Electricity, Traffic, Weather; plus synthetic
Web, economics, energy, transport, climate
Modality
Appendix
Table 1 : Pretraining corpora of the five evaluated TSFMs, as reported by their source publications.
Component pair
Ground truth
F1
F2
F3
F4
F5
F6
F7
F8
F9
F10
Fan–LPC
✓
✓
Fan–HPC
✓
✓
✓
✓
✓
✓
✓
Fan–Comb.
✓
Fan–HPT
✓
✓
✓
✓
✓
✓
Fan–LPT
✓
✓
✓
✓
✓
Fan–Core
✓
Appendix
Table 1: Component couplings under the ground truth structure and under the ten alternatives, one per fold (F1–F10). A mark indicates the pair is coupled.
Sensor
Ground truth
F1
F2
F3
F4
F5
F6
F7
F8
F9
F10
s2
LPC
HPC
LPT
Fan
HPT
Fan
LPT
Core
Fan
Core
Core
s3
HPC
Comb.
Comb.
LPC
LPC
LPC
LPC
Fan
LPT
Fan
LPT
s4
LPT
Fan
HPT
Comb.
Fan
HPC
Core
Fan
HPC
LPC
HPC
s7
HPC
Fan
Fan
Fan
Core
Fan
Fan
LPC
LPT
Fan
LPT
s8
Fan
HPC
LPC
LPC
Comb.
Core
Core
HPC
HPC
LPC
HPC
s9
Core
LPC
LPT
HPT
LPT
HPC
HPC
Fan
HPT
HPT
LPC
Appendix
Table 2: Sensor and component couplings under the ground truth structure and under the ten alternatives, one per fold (F1–F10).
Remaining Useful Life (RUL) prediction is essential for industrial predictive maintenance, yet many learning-based approaches rely on extensive feature engineering or large labeled datasets to train task-specific sequence models. In this work, we introduce a lightweight learning approach, in which we leverage a frozen pretrained time-series foundation model (TSFM) and combine it with a small regression head for RUL estimation from multivariate sensor streams. More specifically, we use Chronos-2 as a frozen backbone to extract context window features and train a lightweight regression neural network for RUL prediction. Experiments on real-world industrial sensor data from two device types show that Chronos-2 features consistently improve over recurrent, convolutional, Transformer-based, and gradient-boosting baselines under the same preprocessing and evaluation protocol. We further analyze the impact of context length and find that performance improves significantly with longer histories, indicating that TSFM representation offer a practical and data-efficient alternative for RUL estimation in industrial settings.
Amir El-Ghoussani, Michele De Vita, Ronald Naumann +1
Friedrich-Alexander-Universität Erlangen–Nürnberg (FAU), Germany · Nokia Solutions and Networks GmbH & Co. KG, Germany
Engineering Digital Twins and Prognostics and Health Management (PHM) systems rely on robust perception modules to extract actionable information from heterogeneous and non-stationary time-series data. However, most existing approaches remain task-specific, data-hungry, and difficult to integrate into scalable monitoring and decision-making pipelines. Moreover, purely data-driven models often lack robustness and transferability across varying operating conditions. To address these challenges, this paper proposes a modular foundation model for time-series perception based on a collection of pretrained representation encoders. The framework leverages self-supervised learning on heterogeneous datasets to learn transferable and task-agnostic representations, which can be reused across multiple PHM tasks. A gating mechanism is introduced to dynamically select relevant encoders for a given target dataset, enabling conditional computation and adaptive model composition. The selected representations are projected into a shared latent space and aggregated using a Transformer-based self-attention module that explicitly models cross-encoder interactions. The resulting architecture supports multiple downstream tasks, including imputation, long-term forecasting, and few-shot learning, through lightweight task-specific heads, while keeping pretrained encoders frozen during adaptation. Extensive ablation studies demonstrate the complementary roles of self-supervised pretraining, encoder selection, representation alignment, and adaptive aggregation. Experimental results on the ETT benchmark show competitive performance across tasks, while a real-world industrial case study on virtual sensing for hydro-generator rotor temperature highlights the practical relevance of the approach.
Time Series Foundation Models (TSFMs) leverage extensive pretraining to accurately predict unseen time series during inference, without the need for task-specific fine-tuning. Through large-scale evaluations on standard benchmarks, we find that leading transformer-based TSFMs exhibit redundant components in their intermediate layers. We introduce a set of tools for mechanistic interpretability of TSFMs, including ablations of specific components and direct logit attribution on the residual stream. Our findings are consistent across several leading TSFMs with diverse architectures, and across a diverse set of real-world and synthetic time-series datasets. We discover that all models in our study are robust to ablations of entire layers. Furthermore, we develop a theoretical framework framing transformers as kernel regressors, motivating a purely intrinsic strategy for ablating heads based on the stable rank of the per-head projection matrices. Using this approach, we uncover the specific heads responsible for degenerate phenomena widely observed in TSFMs, such as parroting of motifs from the context and seasonality bias. Our study sheds light on the universal properties of this emerging class of architectures for continuous-time sequence modeling.
Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai +1
ECE Department, UT Austin · Department of Physics, UT Austin · Oden Institute, UT Austin