Organizations: The Chinese University of Hong Kong, Hong Kong · Cornell University, USA · Southeast University, China · University College London, London, UK · Dalian Polytechnic University, China · Hainan Normal University, China · Anhui Agricultural University, China · Beijing University of Posts and Telecommunications, China · East China Normal University, China · Karlsruhe Institute of Technology, Germany
Tabular foundation models (TFMs) provide a promising route to time-series classification, but their effectiveness depends on how sequential data are converted into tabular representations. Existing representations face two challenges: global aggregation can lose the order of temporal evolution, while features computed in independently fitted coordinate systems may not have consistent meanings across sequences. We therefore view representation design for TFMs as a problem in its own right: the representation should preserve local temporal transitions while maintaining a shared feature definition across samples. We propose SwitchPFN, which learns a shared projection and regime codebook from the training sequences, making local dynamic operators and transition features directly comparable across samples. Across the evaluated benchmarks, SwitchPFN achieves the highest mean accuracy among the evaluated methods, improving over the strongest baseline by 4.47% relatively. Ablation studies, parameter sensitivity analyses, and reduced-training-data experiments further examine the contributions of the representation, its main design choices, and its behavior when labeled data are limited.
Figures & tables
Method
Cortical control
Motion
Speech
Heart sounds
Spectroscopy
Avg. ± SD
Rank
SCP1
SCP2
HW
UW
JV
SAD
HB
EC
Channels
6
7
3
3
12
13
61
3
–
–
Classical baselines
DTW [ 26 ]
87.70
55.60
39.10
57.50
97.20
98.60
77.60
28.70
67.75 ± 0.44
14
XGBoost [ 27 ]
92.50
55.30
28.00
86.20
97.60
97.70
76.10
28.00
70.18 ± 0.25
11
Neural sequence models
Table 1: Test accuracy (%) on eight UEA datasets grouped by domain. Avg. equally weights datasets; Rank orders Avg. (1 = best). Best in bold; second best underlined.
r0
Accuracy
B
Accuracy
α
Accuracy
6
78.32 ± 0.31
256
78.31 ± 0.49
0
78.84 ± 0.16
12
79.01 ± 0.59
512
79.16 ± 0.34
0.25
79.01 ± 0.59
24
79.62 ± 0.67
1024
79.01 ± 0.59
0.5
79.00 ± 0.39
–
–
1536
78.73 ± 0.35
1
69.34 ± 1.03
Table 2: Training-only sensitivity: mean ± SD (%) across three repeats of the eight-dataset OOF average. Bold settings identify the frozen recipe; its result is shared across sweeps.
We introduce RocketPFN, a training-free pipeline for time series classification that combines random convolutional feature extraction (Rocket) with in-context classification via a pretrained tabular foundation model (TabPFN v2.5). On 92 UCR datasets (30-resample protocol), RocketPFN matches HC2, the strongest published method on the archive, in mean accuracy (both 0.900, Wilcoxon p=0.50), with no training on the target data and a median inference time of 30 seconds per fold. It also significantly outperforms every individual classifier in the HC2 ensemble. On UEA (20 datasets) the difference is likewise not statistically significant. A separate comparison concerns TSC foundation models: when paired with the same downstream classifier, MOMENT, Mantis, and MantisV2 are all significantly outperformed by RocketPFN using fewer extracted features and no learned parameters (p<0.001 in each case). This holds even when the encoders were pretrained on corpora that include the UCR training samples. We propose this two-stage pipeline as a reference point for evaluating zero-shot TSC foundation models.
Franco Martino O'Rourke, Ana Trisovic, Dimitris Bertsimas
Time Series Foundation Models (TSFMs) leverage extensive pretraining to accurately predict unseen time series during inference, without the need for task-specific fine-tuning. Through large-scale evaluations on standard benchmarks, we find that leading transformer-based TSFMs exhibit redundant components in their intermediate layers. We introduce a set of tools for mechanistic interpretability of TSFMs, including ablations of specific components and direct logit attribution on the residual stream. Our findings are consistent across several leading TSFMs with diverse architectures, and across a diverse set of real-world and synthetic time-series datasets. We discover that all models in our study are robust to ablations of entire layers. Furthermore, we develop a theoretical framework framing transformers as kernel regressors, motivating a purely intrinsic strategy for ablating heads based on the stable rank of the per-head projection matrices. Using this approach, we uncover the specific heads responsible for degenerate phenomena widely observed in TSFMs, such as parroting of motifs from the context and seasonality bias. Our study sheds light on the universal properties of this emerging class of architectures for continuous-time sequence modeling.
Anthony Bao, Venkata Hasith Vattikuti, Jeffrey Lai +1
ECE Department, UT Austin · Department of Physics, UT Austin · Oden Institute, UT Austin
Time series classification is central to domains such as medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-channel interactions. Random convolutional transforms capture these diverse patterns by converting time series into rich, fixed-dimensional feature representations that can be processed by standard tabular classifiers. While these representations are traditionally paired with simple linear models, we investigate whether a pretrained tabular foundation model can exploit them more effectively and how its performance depends on the available data and inference budget. We propose MASHT, a pipeline that combines MultiRocket and Hydra features with an in-context tabular foundation model. Our approach uses a pretrained tabular foundation model to bypass task-specific model training, requiring only feature extraction and direct inference. Extensive experiments demonstrate that MASHT matches state-of-the-art time series classification baselines on univariate tasks, achieving a lower average rank than HIVE-COTE 2.0. On multivariate datasets, MASHT remains highly competitive with the strongest reference methods. Controlled resource experiments show that compact feature tables retain most of the accuracy at substantially lower runtime, while TabPFN outperforms a matched linear baseline across the evaluated label budgets on univariate tasks. These results highlight practical trade-offs between predictive performance, labeled data, and inference cost.