ChorusTIC: Training-Free Multivariate Time Series Classification via Chorus In-Context Learning
Authors: Juntao Fang, Shifeng Xie, Ruichu Cai, Shengji Zheng, Zijian Li, Keli Zhang, Lujia Pan, Themis Palpanas, +1 more
Organizations: Guangdong University of Technology · Huawei Noah’s Ark Lab · Université Paris Cité · Mohamed bin Zayed University of Artificial Intelligence · Shantou University
Time series classification underpins applications in healthcare, sensing, and industrial monitoring. Although time series foundation models support forecasting and transferable representation learning, classification still typically requires fitting a task-specific classifier on each target dataset, while individual channels of multivariate inputs are often encoded independently. We introduce ChorusTIC, a classification-native foundation model for in-context classification across heterogeneous channel configurations without target-task parameter updates. ChorusTIC combines episode-consistent Random Subchannel Slot Concatenation with a shared dual-axis encoder to model temporal and cross-channel interactions and map variable channel configurations into a fixed-width representation independent of the original channel count. It then calibrates feature axes using context-derived distributions and predicts query labels through leakage-protected in-context learning. We pretrain ChorusTIC solely on synthetic labeled episodes comprising context and query sets that share a task background, with classes distinguished by sparse temporal or cross-channel rules. Evaluations on the complete UEA-30 and UCR-128 archives show strong full-context and low-label performance without target-specific classifier fitting.
Figures & tables
Figure 1: Overview of ChorusTIC. Given labeled context and unlabeled queries, the signal-level Chorus uses one RSSC channel-to-slot assignment throughout the episode. Each sampled group is processed by a shared dual-axis encoder that captures temporal structure within slots and cross-channel interactions across slots. A shared readout summarizes each encoded slot, and fixed-order concatenation yields a fixed-width representation for each sample. The task-level Chorus calibrates feature axes from context-only distributions before row-wise interaction. Finally, the leakage-protected ICL Transformer predicts each query from the labeled context while preventing direct information exchange between queries.
Protocol
Method
Target fit
Avg. Acc.
Avg. Rank
Time-series ICL
ChorusTIC
No
72.27%
3.57
Generic ICL
TabICL
No
65.33%
5.18
TabICLv2
No
67.96%
4.37
Frozen TSFM
MOMENT+SVM
Yes
68.17%
5.48
Mantis+RF
Yes
69.34%
5.22
MantisV2+LR
Yes
70.50%
4.02
Table 1: Classification results on the complete UEA-30 archive. “Target fit” indicates whether a dataset-specific classifier is fitted on the target training split. Best and second-best average accuracies and average ranks are shown in bold and underlined , respectively. Per-dataset results are provided in Appendix D.
Protocol
Method
Target fit
Avg. Acc.
Avg. Rank
Frozen TSFM
MOMENT+SVM
Yes
77.98%
6.11
Mantis+RF
Yes
78.67%
6.42
MantisV2+RF
Yes
78.79%
6.51
MantisV2+LR
Yes
80.03%
5.50
UniShape+RF
Yes
78.86%
5.83
NuTime+RF
Yes
69.39%
9.55
Table 2: Classification results on the UCR-128 archive. “Target fit” indicates whether a classifier is fitted on the target training split. Best and second-best results are shown in bold and underlined , respectively. Per-dataset results are provided in Appendix D.
Method
Target fit
5-shot
10-shot
MOMENT+SVM
Yes
59.42%
63.64%
MantisV2+LR
Yes
61.50%
65.80%
MantisV2+RF
Yes
60.10%
63.66%
UniShape+RF
Yes
62.19%
65.35%
NuTime+RF
Yes
54.68%
59.83%
TabICL
No
59.48%
64.06%
Table 3: Fixed-shot classification accuracy on UEA. Results are averaged over five independently sampled support sets and over the 28 and 24 datasets eligible for the 5-shot and 10-shot settings, respectively. Within each budget, all methods use the same datasets and matched support sets. “Target fit” indicates whether a classifier is fitted on the target support set. Best and second-best results are shown in bold and underlined , respectively.
Figure 2: Scaling with labeled data on UEA-30. Each point reports the average accuracy obtained using the indicated fraction of the official training split.
Factor
Variant
Avg. Acc.
Δ Acc.
Complete model
ChorusTIC
72.27%
0.00
Architecture
No channel-axis attention
71.69%
−0.58
No task-conditioned calibration
70.66%
−1.61
Inference
No label-permutation ensemble
71.38%
−0.89
Pretraining prior
No cross-channel discriminative rules
71.07%
−1.20
Table 4: Ablation results on UEA-30. Δ Acc. is measured relative to the complete model in percentage points.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Configuration
Value
Configuration
Value
Input length L0
512
Task-level embedding width
128
Number of temporal patches M
32
Column-attention blocks
3
RSSC groups G
4
Column-attention heads
4
Slots per RSSC group S
4
Column inducing tokens
128
RSSC slot dimension
32
Row-interaction blocks
3
RSSC sampling strategy
Coverage
Row-attention heads
8
Appendix
Table C.1: Final ChorusTIC architecture and pretraining configuration. All values correspond to the checkpoint used for the reported UCR and UEA results.
Parameter
Final value
Checkpoint
step-6000
Evaluation mode
classifier_v2
Context mode
full training split
Label-permutation ensemble size
8
Cyclic label-permutation ensemble
enabled
RSSC inference draws MR
4
Appendix
Table C.2: Final ChorusTIC inference configuration.
Parameter
Values considered
Final
Selection criterion
Label-permutation ensemble size Mπ
{1,2,4,8}
8
Validation accuracy and inference cost
RSSC inference draws MR
{1,2,4,8}
4
Validation accuracy and inference cost
Cyclic label permutation
Enabled, disabled
Enabled
Validation accuracy
Appendix
Table C.3: Development ranges and final inference hyperparameters.
Method
Official
Version or checkpoint
Target protocol
MOMENT
Yes
MOMENT-1-base
Frozen feature + SVM
Mantis
Yes
Mantis-8M
Frozen feature + RF
MantisV2
Yes
MantisV2
Frozen feature + LR/RF
UniShape
Yes
unishape_checkpoint_zeroshot
Frozen feature + RF
NuTime
Yes
checkpoint_bias9
Frozen CLS + RF
TabICL
Yes
tabicl-classifier-v1.1-20250506
In-context inference
Appendix
Table C.4: Implementations and target-task protocols of the comparison methods. “Official” indicates the use of an author-released implementation.
Item
Configuration
GPU model and count
4× NVIDIA Tesla V100 PCIe
GPU memory
32 GiB per GPU (128 GiB total)
GPU driver
570.86.15
CPU model
Intel Xeon Gold 6140 @ 2.30 GHz
System memory
251 GiB
Operating system
Ubuntu 18.04.6 LTS
Appendix
Table C.5: Computing and software environment.
Protocol
Method
Target fit
Avg. Acc. ↑
Avg. Rank ↓
W/T/L
Time-series ICL
ChorusTIC
No
72.27%
3.57
—
Generic ICL
TabICL
No
65.33%
5.18
18/3/9
TabICLv2
No
67.96%
4.37
16/1/13
Frozen TSFM
MOMENT+SVM
Yes
68.17%
5.48
22/2/6
Mantis+RF
Yes
69.34%
5.22
21/1/8
MantisV2+LR
Yes
70.50%
4.02
18/3/9
Appendix
Table D.1: Classification results on the complete UEA-30 archive. “Target fit” indicates whether a dataset-specific classifier is fitted on the target training split. W/T/L counts are reported from the perspective of ChorusTIC. Best and second-best average accuracies and average ranks are shown in bold and underlined , respectively.
Protocol
Method
Target fit
Avg. Acc.
Avg. Rank
Frozen TSFM
MOMENT+SVM
Yes
77.98%
6.11
Mantis+RF
Yes
78.67%
6.42
MantisV2+RF
Yes
78.79%
6.51
MantisV2+LR
Yes
80.03%
5.50
UniShape+RF
Yes
78.86%
5.83
NuTime+RF
Yes
69.39%
9.55
Appendix
Table D.2: Classification results on the complete UCR-128 archive. “Target fit” indicates whether a classifier is fitted on the target training split. Best and second-best results are shown in bold and underlined , respectively.
Dataset
TabICL
TabICLv2
MOMENT
Mantis
MV2-LR
MV2-RF
UniShape
NuTime
ChorusTIC
ArticularyWordRecognition
0.9800
0.9833
0.9600
0.9927
0.9933
0.9933
0.9900
0.7600
0.9800
AtrialFibrillation
0.2000
0.3333
0.1333
0.2800
0.1333
0.0667
0.2000
0.1333
0.2667
BasicMotions
1.0000
0.9750
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
1.0000
CharacterTrajectories
0.9847
0.9937
0.9742
0.9401
0.9742
0.9568
0.9749
0.8621
0.9889
Cricket
0.9444
0.9444
0.9861
1.0000
0.9861
0.9861
0.9722
0.8611
0.9722
DuckDuckGeese
0.2000
0.2000
0.4600
0.3880
0.4800
0.5000
0.4600
0.2600
0.4600
Appendix
Table D.3: Per-dataset classification accuracy on the complete UEA-30 archive. Best and second-best results within each dataset are shown in bold and underlined , respectively, based on the displayed four-decimal accuracies. Avg. Acc. is the macro-average over datasets; Avg. Rank is computed among the 9 displayed methods using full-precision accuracies. Abbreviations: MOMENT=MOMENT+SVM, Mantis=Mantis+RF, MV2-LR=MantisV2+LR, MV2-RF=MantisV2+RF, UniShape=UniShape+RF, and NuTime=NuTime+RF.
Time series classification (TSC) is dominated by a two-stage paradigm: train a feature encoder -- either from scratch on the target dataset or via pretraining on large corpora -- and then fit a task-specific classifier on top. While effective, this decoupling optimizes representation learning independently of the classification objective, requires per-dataset training, and prevents the model from exploiting label information during inference. We introduce TimEE, a 4.5M-parameter foundation model for end-to-end TSC via in-context learning. Given a labeled support set and a query time series, TimEE directly outputs a predicted class distribution in a single forward pass with no per-dataset training required. Following the prior-data fitted network (PFN) framework, TimEE is meta-trained exclusively on synthetic TSC tasks, where each task contains time series with distinct class identities arising from structured distributional shifts in the generative process. Despite seeing no real time series during pre-training, TimEE ranks first in ROC AUC (and third on accuracy) on the UCR benchmark among all compared methods, which include both foundation models and supervised deep learning baselines. To our knowledge, TimEE is the first purely synthetic-pretrained model to reach state-of-the-art performance on the UCR benchmark. These results establish end-to-end ICL with synthetic priors as a compelling, largely unexplored direction for TSC, with scaling, prior design, and richer generation mechanisms as natural avenues for improvement. Code is publicly available at http://github.com/automl/timee.
Jaris Küken, Shi Bin Hoo, Martin Mráz +2
University of Freiburg · Zuse School ELIZA Darmstadt · Prior Labs +1
The zero-shot evaluation of time series foundation models (TSFMs) for classification typically uses a frozen encoder followed by a task-specific classifier. However, this practice violates the training-free premise of zero-shot deployment and introduces evaluation bias due to classifier-dependent training choices. To address this issue, we propose TIC-FM, an in-context learning framework that treats the labeled training set as context and predicts labels for all test instances in a single forward pass, without parameter updates. TIC-FM pairs a time series encoder and a lightweight projection adapter with a split-masked latent memory Transformer. We further provide theoretical justification that in-context inference can subsume trained classifiers and can emulate gradient-based classifier training within a single forward pass. Experiments on 128 UCR datasets show strong accuracy, with consistent gains in the extreme low-label situation, highlighting training-free transfer for time series classification.The source code is publicly available at https://github.com/fangjuntao/TIC-FM.
Juntao Fang, Shifeng Xie, Shengbin Nie +7
Guangdong University of Technology, Guangzhou, China · Huawei Noah’s Ark Lab, Paris, France · Paris Descartes University, Paris, France +2
Time series classification is central to domains such as medical signal analysis, industrial monitoring, and sensor-based activity recognition, where class information manifests as localized shapes, specific frequencies, temporal shifts, or complex cross-channel interactions. Random convolutional transforms capture these diverse patterns by converting time series into rich, fixed-dimensional feature representations that can be processed by standard tabular classifiers. While these representations are traditionally paired with simple linear models, we investigate whether a pretrained tabular foundation model can exploit them more effectively and how its performance depends on the available data and inference budget. We propose MASHT, a pipeline that combines MultiRocket and Hydra features with an in-context tabular foundation model. Our approach uses a pretrained tabular foundation model to bypass task-specific model training, requiring only feature extraction and direct inference. Extensive experiments demonstrate that MASHT matches state-of-the-art time series classification baselines on univariate tasks, achieving a lower average rank than HIVE-COTE 2.0. On multivariate datasets, MASHT remains highly competitive with the strongest reference methods. Controlled resource experiments show that compact feature tables retain most of the accuracy at substantially lower runtime, while TabPFN outperforms a matched linear baseline across the evaluated label budgets on univariate tasks. These results highlight practical trade-offs between predictive performance, labeled data, and inference cost.