A representation family is a distinct way of extracting features from time series. Ensemble algorithms that combine several representation families remain the most accurate approach to time series classification. Current state-of-the-art ensembles, most notably HIVE-COTE2.0, pair a bespoke classification algorithm with each representation family and combine their predictions using a fixed, non-adaptive rule. We present TIGER (Time-series classification with In-context-learning Gated Ensemble of Representations), which instead applies the same small portfolio of three general-purpose classifiers (Ridge, Extra Trees, and Naive Bayes) to four representations from four distinct families, stacking the resulting twelve base learners' predictions into a meta-feature matrix. The final prediction is produced by an adaptive meta-classification rule that chooses, independently for each data set, between a weighted hard majority vote and TabICLv2, a pretrained tabular foundation model used in-context as a meta-classifier, based on the mean number of training samples available per class. On a 142-data-set benchmark drawn from the UCR time series classification archive, TIGER obtains the best mean accuracy, balanced accuracy, and F1-score among six compared algorithms, including HIVE-COTE2.0, and significantly outperforms each of the other five individually. TIGER's adaptive rule also meaningfully outperforms either of its two constituent meta-classification methods used alone, and its single hyperparameter, tuned using only a twenty-data-set development subset, is shown to generalize to the full evaluation benchmark. We further characterize TIGER's design through an extensive set of ablation experiments and report the design alternatives that we investigated and ultimately discarded.
Figures & tables
Figure 1 : Overview of HIVE-COTE 2.0’s architecture.
Figure 2 : Architecture of TIGER.
Figure 3 : Adaptive cross-validation splitting strategy for out-of-fold predictions.
Figure 4 : Critical difference diagram for the six candidate meta-classification methods, the best base learner (picked using cross-validation), and TIGER’s adaptive rule, over the full 142-data-set benchmark, in terms of mean accuracy.
Figure 5 : TabICLv2 minus weighted-hard-majority-vote accuracy gap against training samples per class, with threshold τ=60.83 .
Data set
Mean samples/class
TabICLv2
WHMV
Δ
PigAirwayPressure
2.0
0.602
0.865
-0.264
PigCVP
2.0
0.823
0.928
-0.105
Beef
6.0
0.702
0.778
-0.076
PigArtPressure
2.0
0.892
0.963
-0.071
Rock
5.0
0.839
0.883
-0.043
WordSynonyms
10.7
0.722
0.765
-0.043
Table 4 : The ten data sets on which TabICLv2 loses the most, in accuracy, compared to WHMV.
Figure 6 : Critical-difference diagram for TIGER and the five literature algorithms on accuracy, over the full 142-data-set benchmark.
Algorithm
ACC
BALACC
AUROC
NLL
F1
TIGER
0.884 (1)
0.863 (1)
0.933 (3)
2.775 (3)
0.863 (1)
KG-MTP
0.876 (2)
0.855 (2)
0.906 (4)
4.457 (4)
0.857 (2)
HC2
0.876 (3)
0.848 (3)
0.962 (1)
0.403 (1)
–
RDST
0.863 (4)
0.836 (4)
0.893 (5)
4.943 (5)
0.835 (3)
WEASEL 2.0
0.858 (5)
0.830 (6)
0.888 (6)
5.114 (6)
0.825 (5)
QUANT
0.855 (6)
0.830 (5)
0.956 (2)
0.533 (2)
0.831 (4)
Table 5 : Summary performance measures of TIGER and five literature algorithms on the full 142-data-set benchmark, ordered by accuracy. Ranks are in parentheses, the best value per column is bold, and “–” marks a metric not reported.
Algorithm
ACC
BALACC
AUROC
NLL
F1
TIGER
0.889 (1)
0.867 (1)
0.935 (3)
2.606 (3)
0.867 (1)
KG-MTP
0.884 (2)
0.862 (2)
0.911 (4)
4.186 (4)
0.865 (2)
HC2
0.882 (3)
0.855 (3)
0.963 (1)
0.387 (1)
–
RDST
0.872 (4)
0.847 (4)
0.902 (5)
4.613 (5)
0.849 (3)
WEASEL 2.0
0.871 (5)
0.845 (5)
0.900 (6)
4.654 (6)
0.844 (4)
QUANT
0.865 (6)
0.838 (6)
0.957 (2)
0.509 (2)
0.840 (5)
Table 6 : Summary performance measures on the 122 held-out data sets, ordered by accuracy. Ranks are in parentheses, the best value per column is bold, and “–” marks a metric not reported.
Algorithm
Total (h)
Median (s)
Minimum (s)
Maximum (min)
QUANT
0.26
2.0
0.3
1.1
KG-MTP
2.08
14.4
1.2
10.4
RDST
3.11
19.3
1.0
13.8
WEASEL 2.0
5.94
24.8
1.6
43.2
TIGER
20.65
145.4
19.0
213.2
HC2
149.25
834.7
55.6
1469.9
Table 7 : Combined training-plus-inference wall-clock runtime for TIGER and five literature algorithms, ordered by median runtime (lowest in bold).
Algorithm
Total (h)
Median (s)
Minimum (s)
Maximum (min)
QUANT
0.25
1.6
0.2
1.1
KG-MTP
0.98
5.8
0.2
9.3
RDST
1.52
8.4
0.2
9.9
WEASEL 2.0
2.90
11.6
0.8
17.3
TIGER
10.90
76.3
15.2
114.4
HC2
100.86
418.0
17.6
1331.4
Table 8 : Training wall-clock runtime for TIGER and five literature algorithms, ordered by median runtime (lowest in bold).
Algorithm
Total (h)
Median (s)
Minimum (s)
Maximum (min)
QUANT
0.02
0.1
0.0
0.1
KG-MTP
1.10
7.1
0.4
6.5
RDST
1.58
8.0
0.3
11.9
WEASEL 2.0
3.04
8.7
0.4
39.3
TIGER
9.76
39.5
1.8
135.7
HC2
48.34
248.2
14.7
361.5
Table 9 : Inference wall-clock runtime for TIGER and five literature algorithms, ordered by median runtime (lowest in bold).
Algorithm
Learner
k -fold
Bag(15)
Bag(25)
Native OOB
KG-MTP
Ridge
0.0327
0.0347
0.0334
–
KG-MTP
Extra Trees
0.0357
–
–
0.0390
KG-MTP
Naive Bayes
0.0546
0.0503
0.0494
–
QUANT
Ridge
0.0360
0.0403
0.0383
–
QUANT
Extra Trees
0.0383
–
–
0.0407
QUANT
Naive Bayes
0.0431
0.0535
0.0520
–
Table 10 : Mean absolute gap between each out-of-fold estimator’s accuracy and the held-out test accuracy, per (family, learner) pair. Smaller is more honest, and the best value per row is bold.
Algorithm
Learner
k -fold
Bag(15)
Bag(25)
Native OOB
KG-MTP
Ridge
44.1
171.1
282.5
–
KG-MTP
Extra Trees
11.1
–
–
54.8
KG-MTP
Naive Bayes
25.8
217.8
360.0
–
QUANT
Ridge
4.1
15.9
26.6
–
QUANT
Extra Trees
5.8
–
–
8.0
QUANT
Naive Bayes
1.7
14.7
24.4
–
Table 11 : Mean wall-clock seconds per unit for each out-of-fold estimator, per (family, learner) pair (fastest in bold).
Accuracy
Runtimes (s)
Data set
Classes
Native
ECOC
Native
ECOC
EOGHorizontalSignal
12
0.814
0.818
132
886
EOGVerticalSignal
12
0.771
0.762
131
880
FacesUCR
14
0.957
0.962
361
1557
GestureMidAirD3
26
0.506
0.511
150
972
PigAirwayPressure
52
0.548
0.526
291
1248
Table 12 : TabICLv2’s native many-class handling vs. its ECOC-wrapped variant, on the six development data sets with more than ten classes (runtimes in seconds, best accuracy per row in bold).
Method
Mean accuracy
Gap to six-method oracle
Best single base learner (CV-picked)
0.8778
0.0094
TIGER’s adaptive rule (deployed)
0.8838
0.0034
WHMV / TabICLv2 oracle (hindsight)
0.8869
0.0003
Six-method oracle (hindsight)
0.8872
0.0000
Best single base learner (test hindsight)
0.8922
–
Table 13 : Mean accuracy of TIGER’s deployed rule against four reference points (red rows use test-set hindsight and cannot be deployed).
Method
Data sets won
Mean margin
Max margin
CAWPE
19
0.0013
0.0103
Hard majority vote
17
0.0005
0.0026
LightGBM
10
0.0002
0.0013
Logistic regression
6
0.0018
0.0082
Table 14 : Per-method breakdown of the six-method oracle’s gain over Table 13 ’s two-method oracle, for the four methods TIGER does not deploy. “Data sets won” counts the data sets where that method ties or leads all six candidates. “Margin” is its own accuracy minus the two-method oracle’s, on the data sets it wins.
Algorithm
Mean accuracy
Mean balanced accuracy
Data sets won
KG-MTP
0.8661
0.8437
75
RDST
0.8577
0.8301
30
WEASEL 2.0
0.8432
0.8148
27
QUANT
0.8328
0.8077
21
Table 15 : Each representation family’s standalone predictive strength before stacking, on 139 of 142 data sets (best per column in bold).
Classes
Data sets
TIGER’s adaptive rule
Six-method oracle
Gap
2 (binary)
50
0.9291
0.9307
0.0016
3–5
38
0.8905
0.8930
0.0025
6–10
30
0.8478
0.8532
0.0054
11+
24
0.8241
0.8297
0.0057
Table 16 : Meta-oracle headroom broken down by number of classes.
Training samples
Data sets
TIGER’s adaptive rule
Six-method oracle
Gap
≤ 60
36
0.9447
0.9469
0.0022
61–217
35
0.8347
0.8378
0.0031
218–466
36
0.8273
0.8339
0.0066
≥ 467
35
0.9285
0.9299
0.0014
Table 17 : Meta-oracle headroom broken down by training set size, binned at the quartiles.
Data set
Mean samples/class
TabICLv2 vs. oracle
≥τ
PhoneHeartbeatSound
84.8
+0.072
✓
EthanolLevel
126.0
+0.072
✓
SemgHandMovementCh2
75.0
+0.061
✓
ScreenType
125.0
+0.044
✓
ChlorineConcentration
155.7
+0.042
✓
SemgHandSubjectCh2
90.0
+0.041
✓
Table 18 : The ten data sets on which TabICLv2, used alone as the meta-classifier, gains the most accuracy relative to the WHMV/TabICLv2 oracle (checkmark: at or above τ=60.83 ).
Figure 7 : TabICLv2’s mean accuracy per family as a function of reduced feature count. “default” is each family’s own unconstrained reduction.
Figure 8 : TabICLv2 at its best reduced-feature tag vs. each family’s original algorithm on raw features.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Data set
Training
Test
Classes
Length
Mean samples/class
ACSF1
100
100
10
1460
10.0
AconityMINIPrinterLarge
2403
1184
2
300
1201.5
AconityMINIPrinterSmall
589
292
2
300
294.5
Adiac
390
391
37
176
10.5
AllGestureWiimoteX
300
700
10
500
30.0
AllGestureWiimoteY
300
700
10
500
30.0
Appendix
Table 20 : The 142 data sets of the full evaluation benchmark (mean samples/class computed on the training set).
Data set
HMV
LR
LightGBM
CAWPE
WHMV
TabICLv2
ACSF1
0.866
0.857
0.855
0.862
0.871
0.873
AconityMINIPrinterLarge
0.954
0.963
0.964
0.953
0.956
0.967
AconityMINIPrinterSmall
0.975
0.978
0.975
0.975
0.975
0.978
Adiac
0.811
0.828
0.814
0.823
0.821
0.852
AllGestureWiimoteX
0.745
0.758
0.755
0.778
0.771
0.780
AllGestureWiimoteY
0.754
0.780
0.786
0.798
0.785
0.804
Appendix
Table 21 : Mean accuracy (over 30 resamples) of the six candidate meta-classification methods, for each of the 142 data sets (HMV = hard majority vote, LR = logistic regression, WHMV = weighted hard majority vote).
Algorithm
ACC
BALACC
AUROC
NLL
F1
TIGER
0.895 (1)
0.862 (1)
0.955 (1)
0.265 (1)
0.860 (1)
KG-MTP
0.879 (2)
0.839 (2)
0.889 (4)
4.374 (4)
0.842 (2)
HC2
0.878 (3)
0.825 (4)
0.950 (2)
0.350 (2)
–
QUANT
0.874 (4)
0.831 (3)
0.948 (3)
0.402 (3)
0.832 (3)
RDST
0.863 (5)
0.813 (5)
0.867 (5)
4.952 (5)
0.807 (4)
WEASEL 2.0
0.855 (6)
0.801 (6)
0.858 (6)
5.240 (6)
0.788 (5)
Appendix
Table 22 : Summary performance measures on the 57 data sets where TIGER’s adaptive rule always selects TabICLv2.
Data set
ACC
BALACC
AUROC
NLL
F1
ACSF1
0.871
0.871
0.928
4.662
0.868
AconityMINIPrinterLarge
0.967
0.964
0.988
0.115
0.951
AconityMINIPrinterSmall
0.978
0.966
0.982
0.101
0.960
Adiac
0.821
0.825
0.908
6.468
0.813
AllGestureWiimoteX
0.771
0.771
0.873
8.266
0.768
AllGestureWiimoteY
0.785
0.785
0.881
7.734
0.783
Appendix
Table 23 : TIGER’s mean ACC, BALACC, AUROC, NLL, and F1 (over 30 resamples), for each of the 142 data sets.
Figure 9 : Reliability of the Platt-calibrated Ridge meta-features, per representation family. The dashed diagonal marks perfect calibration.
Algorithm
Expected Calibration Error
WEASEL 2.0
0.0440
KG-MTP
0.0443
RDST
0.0453
QUANT
0.0583
Appendix
Table 24 : Expected calibration error of the Platt-scaled Ridge meta-features, per representation family (lowest in bold).
Time series classification is central to domains such as medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-channel interactions. Random convolutional transforms capture these diverse patterns by converting time series into rich, fixed-dimensional feature representations that can be processed by standard tabular classifiers. While these representations are traditionally paired with simple linear models, we investigate whether a pretrained tabular foundation model can exploit them more effectively and how its performance depends on the available data and inference budget. We propose MASHT, a pipeline that combines MultiRocket and Hydra features with an in-context tabular foundation model. Our approach uses a pretrained tabular foundation model to bypass task-specific model training, requiring only feature extraction and direct inference. Extensive experiments demonstrate that MASHT matches state-of-the-art time series classification baselines on univariate tasks, achieving a lower average rank than HIVE-COTE 2.0. On multivariate datasets, MASHT remains highly competitive with the strongest reference methods. Controlled resource experiments show that compact feature tables retain most of the accuracy at substantially lower runtime, while TabPFN outperforms a matched linear baseline across the evaluated label budgets on univariate tasks. These results highlight practical trade-offs between predictive performance, labeled data, and inference cost.
Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset or via pretraining on large corpora -- and then fit a task-specific classifier on top. While effective, this decoupling optimizes representation learning independently of the classification objective, requires per-dataset training, and prevents the model from exploiting label information during inference. We introduce TimEE, a 4.5M-parameter foundation model for end-to-end TSC via in-context learning. Given a labeled support set and a query time series, TimEE directly outputs a predicted class distribution in a single forward pass with no per-dataset training required. Following the prior-data fitted network (PFN) framework, TimEE is meta-trained exclusively on synthetic TSC tasks, where each task contains time series with distinct class identities arising from structured distributional shifts in the generative process. Despite seeing no real time series during pre-training, TimEE ranks first in ROC AUC (and third on accuracy) on the UCR benchmark among all compared methods, which include both foundation models and supervised deep learning baselines. To our knowledge, TimEE is the first purely synthetic-pretrained model to reach state-of-the-art performance on the UCR benchmark. These results establish end-to-end ICL with synthetic priors as a compelling, largely unexplored direction for TSC, with scaling, prior design, and richer generation mechanisms as natural avenues for improvement. Code is publicly available at http://github.com/automl/timee.
Jaris Küken, Shi Bin Hoo, Martin Mráz +2
University of Freiburg · Zuse School ELIZA Darmstadt · Prior Labs +1
We introduce RocketPFN, a training-free pipeline for time series classification that combines random convolutional feature extraction (Rocket) with in-context classification via a pretrained tabular foundation model (TabPFN v2.5). On 92 UCR datasets (30-resample protocol), RocketPFN matches HC2, the strongest published method on the archive, in mean accuracy (both 0.900, Wilcoxon p=0.50), with no training on the target data and a median inference time of 30 seconds per fold. It also significantly outperforms every individual classifier in the HC2 ensemble. On UEA (20 datasets) the difference is likewise not statistically significant. A separate comparison concerns TSC foundation models: when paired with the same downstream classifier, MOMENT, Mantis, and MantisV2 are all significantly outperformed by RocketPFN using fewer extracted features and no learned parameters (p<0.001 in each case). This holds even when the encoders were pretrained on corpora that include the UCR training samples. We propose this two-stage pipeline as a reference point for evaluating zero-shot TSC foundation models.
Franco Martino O'Rourke, Ana Trisovic, Dimitris Bertsimas