Time-series anomaly detection (TSAD) is difficult to generalize across datasets because heterogeneous temporal dynamics imply different notions of normality and favor different detection criteria. While time-series foundation models provide transferable representations, coupling them with a fixed anomaly-scoring mechanism can overlook this variation. This motivates a different perspective on foundation-model-based TSAD: using foundation models to coordinate specialized anomaly criteria rather than directly imposing a universal one. Based on this view, we propose \textbf{TS-Router}, a generalist-representation, specialist-detection framework that estimates the relative competence of heterogeneous anomaly detectors from pretrained temporal representations and selects suitable specialists for each target series. To avoid relying on specialist-performance labels from real tasks, we derive soft competence supervision from specialists' relative performance on labeled simulated tasks. At deployment, routing requires no target anomaly labels, and only the selected specialists are fitted unsupervisedly on the target series. We bound Top-k set-competence regret under representation coverage and conditional competence stability. Across 16 real-world benchmarks and four complementary evaluation metrics, TS-Router achieves the best overall average rank. Controlled ablations with multiple frozen TSFM encoders further support the use of pretrained representations for competence estimation and adaptive specialist selection. The code is available at https://anonymous.4open.science/r/TS-Router-D8FF.
Figures & tables
Figure 1: Domain-wise VUS-PR of representative specialists relative to TimeRCD, normalized by TimeRCD performance.
Figure 2: Overview of the TS-Router framework.
Method
VUS-PR
Aff.-F1
F1 T
Std.-F1
Overall Rank
#Top-1
#Top-2
Score ↑
Rank ↓
Score ↑
Rank ↓
Score ↑
Rank ↓
Score ↑
Rank ↓
Foundation-guided specialist routing
TS-Router
49.59
2.44
86.87
2.44
47.37
3.00
45.34
2.75
2.66
26
39
Direct zero-shot models
TimeRCD
37.00
3.88
83.94
3.69
40.88
3.62
37.03
4.19
3.84
14
26
DADA
30.16
5.62
81.41
4.88
37.29
5.84
32.60
6.31
5.66
5
10
Table 1: Overall performance across 16 real-world TSAD benchmarks. Score denotes the mean performance across datasets, and Rank denotes the average dataset-wise rank among the 12 compared methods, with ties assigned average ranks. Overall Rank averages ranks across all four metrics. #Top-1 and #Top-2 count first-place and top-two finishes across the 64 dataset–metric combinations. Methods follow different deployment paradigms without using target anomaly labels for model fitting. Best results are in bold and second-best results are underlined .
Variant
Detection performance
Routing quality
VUS-PR ↑
Aff.-F1 ↑
F1 T ↑
Std.-F1 ↑
NDCG@ k↑
Hit@ k↑
Selection and fusion strategy
Best Fixed
39.68
84.34
41.01
36.43
0.607
0.207
Full Ensemble
49.54
85.55
46.22
43.48
–
–
Best Fixed Top-3
47.44
86.56
47.56
43.42
0.625
0.532
Router Top-1
49.00
86.80
43.47
41.41
–
–
Table 2: Routing effectiveness and representation ablations on eleven univariate benchmarks. Best Fixed denotes the hindsight-best single specialist shared across targets for each evaluation metric, whereas Oracle denotes hindsight per-target individual-specialist selection; these are single-specialist detection references, not Top-3 fusion oracles. Best Top-3 always fuses the hindsight-best set of three specialists shared across targets. Routing metrics use k=1 for Best Fixed and Oracle and k=3 otherwise; Router Top-1 shares the TS-Router ranking, so its routing metrics are not repeated. Representation variants differ only in the router representation, with the same router architecture, specialist pool, supervision, and Top-3 normalized mean fusion. The dashed line separates conventional feature representations from frozen TSFM encoders (Chronos, MOMENT, TimeRCD); TimeRCD corresponds to the default TS-Router configuration.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Name
Domain
#TS
Avg. Length
AR (%)
UCR
Misc.
228
67818.7
0.6
NAB
Mixed
28
5099.7
10.6
YAHOO
Web
259
1560.2
0.6
IOPS
Operations
17
72792.3
1.3
MGAB
Synthetic
9
97777.8
0.2
SED
Medical
3
23332.3
4.1
Appendix
Table 3: Statistics of the univariate benchmark datasets.
Name
Domain
#TS
Avg. Length
AR (%)
MSL
Space
16
3119.4
5.1
PSM
Sensor
1
217624.0
11.2
SMAP
Space
27
7855.9
2.9
SMD
Server
22
25466.4
3.8
SWaT
ICS
2
207457.5
12.7
Appendix
Table 4: Statistics of the multivariate benchmark datasets.
Specialist
Primary inductive bias
PCA
Low-dimensional subspace reconstruction; anomalies produce large reconstruction residuals.
Sub-PCA
Windowed PCA on subsequences; anomalies produce large reconstruction residuals relative to local temporal structure.
Sub-HBOS
Marginal density modeling through histogram-based statistics; low-density observations receive larger anomaly scores.
POLY
Smooth temporal structure modeled by polynomial fitting; deviations from the fitted temporal trend indicate anomalies.
Sub-KNN
Neighborhood-distance structure; observations far from their nearest neighbors are considered anomalous.
Sub-LOF
Relative local density; anomalies exhibit substantially lower local density than their surrounding neighborhoods.
Appendix
Table 5: Specialists and their primary anomaly-detection inductive biases.
Specialist
Hyperparameters
PCA
Period rank 1 ; all principal components.
Sub-PCA
Period rank 1 ; all principal components.
Sub-HBOS
Period rank 1 ; 10 histogram bins.
POLY
Period rank 1 ; polynomial degree 4 .
Sub-KNN
Period rank 2 ; 50 neighbors.
Sub-LOF
Period rank 2 ; 30 neighbors.
Appendix
Table 6: Specialist hyperparameters used for competence construction and target deployment.
Metric
Model
Univariate datasets
Multivariate datasets
IOPS
MGAB
NAB
NEK
Power
SED
Stock
TODS
UCR
WSD
YAHOO
MSL
PSM
SMAP
SMD
SWaT
VUS-PR
TS-Router
41.49
43.09
51.97
76.92
22.40
65.42
73.18
71.65
39.60
56.91
77.32
20.86
19.55
38.12
52.24
42.75
TimeRCD
20.23
1.05
24.32
27.88
21.25
80.75
77.28
93.46
23.09
21.77
84.41
20.45
18.69
22.68
37.03
17.58
DADA
24.97
0.57
24.73
46.85
10.61
6.42
99.51
64.83
2.94
33.42
70.74
12.74
17.17
20.02
25.98
21.13
Chronos
19.00
0.60
23.76
31.80
10.95
8.65
97.49
70.66
6.56
18.81
83.54
8.25
14.61
5.18
10.22
16.44
MOMENT
37.35
0.56
45.38
67.74
10.50
4.31
76.97
56.45
6.17
55.26
30.81
9.32
16.48
8.97
15.96
14.90
Appendix
Table 7: Per-dataset performance across 16 real-world TSAD benchmarks. Best results are in bold and second-best results are underlined .
Metric
Representation
Univariate datasets
Avg.
IOPS
MGAB
NAB
NEK
Power
SED
Stock
TODS
UCR
WSD
YAHOO
VUS-PR
TimeRCD
41.49
43.09
51.97
76.92
22.40
65.42
73.18
71.65
39.60
56.91
77.32
56.36
Chronos
38.63
33.44
49.34
50.10
15.99
94.18
71.64
78.81
33.17
53.83
66.87
53.27
MOMENT
40.73
33.44
46.83
33.93
20.56
94.18
87.24
76.35
38.09
57.55
63.87
53.89
Catch22
44.58
9.46
42.75
38.31
22.27
94.18
76.57
84.05
34.11
48.86
54.39
49.96
TSFresh
36.98
35.44
35.53
50.10
17.54
37.17
88.43
73.84
33.18
51.50
79.40
49.01
Appendix
Table 8: Representation ablation on eleven univariate TSB-AD benchmarks. All variants use the same MLP, 11-specialist pool, and Top-3 normalized mean fusion; only the router representation changes. TimeRCD remains frozen. Best results are in bold and second-best results are underlined .
Metric
Method
Univariate datasets
Avg.
IOPS
MGAB
NAB
NEK
Power
SED
Stock
TODS
UCR
WSD
YAHOO
VUS-PR
TS-Router
41.49
43.09
51.97
76.92
22.40
65.42
73.18
71.65
39.60
56.91
77.32
56.36
PCA
23.21
0.60
45.77
78.45
10.49
3.77
80.92
54.01
13.40
16.95
21.15
31.70
Sub-PCA
20.35
0.66
49.82
83.68
10.49
4.00
77.81
52.82
13.53
14.09
19.41
31.51
Sub-HBOS
4.97
0.50
33.26
19.43
15.69
63.09
64.60
64.54
15.07
2.21
12.08
26.86
POLY
27.51
0.70
46.91
55.64
9.30
7.88
79.01
57.39
15.80
35.24
27.89
33.02
Appendix
Table 9: Specialist complementarity on eleven univariate TSB-AD benchmarks. TS-Router uses Top-3 normalized mean fusion, while Oracle denotes hindsight per-target selection of the best individual specialist. Oracle is an individual-specialist reference rather than an upper bound on the fused TS-Router prediction. Best and second-best non-oracle results are in bold and underlined , respectively.
Figure 3: Effect of the number of selected specialists k with the 11-specialist pool. Each panel reports the unweighted mean over eleven univariate benchmarks; the panels show VUS-PR, F1 T , Standard-F1, and Affiliation-F1.
Figure 4: Effect of specialist-pool size with Top-3 selection. Each panel reports the unweighted mean over eleven univariate benchmarks; the panels show VUS-PR, F1 T , Standard-F1, and Affiliation-F1. The router is retrained for each pool size.
Routing head
VUS-PR
VUS-ROC
F1 T
Std.-F1
Aff.-F1
VUS-PR
56.36
88.30
49.62
47.73
88.11
Affiliation-F1
56.27
89.24
49.17
46.83
88.49
F1 T
58.24
89.29
53.03
49.96
89.17
Standard-F1
55.28
87.81
50.00
47.77
87.83
Appendix
Table 10: Top-3 routing with different competence-supervision heads. Values are dataset-averaged percentages over eleven univariate benchmarks. The pool size is 11 and the router is trained with the indicated head.
Metric
Selector
Univariate datasets
Avg.
IOPS
MGAB
NAB
NEK
Power
SED
Stock
TODS
UCR
WSD
YAHOO
VUS-PR
TS-Router
41.49
43.09
51.97
76.92
22.40
65.42
73.18
71.65
39.60
56.91
77.32
56.36
SATzilla
41.77
33.44
41.71
53.12
22.40
74.65
70.43
72.73
41.82
58.30
59.20
51.78
ARGOSMART
44.02
9.46
44.73
31.94
22.36
11.94
82.21
70.71
34.24
54.88
53.37
41.81
MetaOD
42.21
33.44
40.23
57.82
22.40
33.96
70.20
72.73
41.82
56.72
57.89
48.13
ISAC
43.02
33.44
41.84
39.78
22.40
74.65
76.03
72.72
40.42
58.08
57.09
50.86
Appendix
Table 11: Algorithm-selection transfer on eleven univariate TSB-AD benchmarks. Classical selectors are trained on Syn-RCD using the common 11-specialist pool. TS-Router uses its fixed routing policy learned from the same synthetic competence supervision. All selectors use Top- 3 fusion of the predicted specialists. Best results are in bold and second-best results are underlined .
Metric
Method
Univariate datasets
Avg.
IOPS
MGAB
NAB
NEK
SED
Stock
TODS
UCR
WSD
YAHOO
VUS-PR
TS-Router
28.72
45.61
57.45
94.24
63.78
52.84
87.54
29.67
45.05
69.09
57.40
SATzilla
20.45
48.52
43.74
76.40
77.57
90.36
73.87
38.32
35.01
45.81
55.00
ARGOSMART
23.07
48.52
42.86
41.35
77.57
94.57
78.03
25.53
38.48
49.53
51.95
MetaOD
19.58
48.52
39.00
76.40
77.57
90.36
77.94
39.62
35.01
55.13
55.91
ISAC
18.96
42.10
42.93
36.70
36.09
52.00
64.30
31.10
33.26
19.34
37.68
Appendix
Table 12: In-distribution algorithm selection on the official TSB-AD Tuning subset. Evaluation uses the thirty genuinely univariate Tuning series from ten collections; Power does not appear in the Tuning split. Classical selectors are adapted from a small labeled support set across ten random trials, whereas TS-Router uses a fixed routing policy without support-set adaptation. Best results are in bold and second-best results are underlined .
Metric
Method
Held-out target dataset
Avg.
IOPS
MGAB
NAB
NEK
Power
SED
Stock
TODS
UCR
WSD
YAHOO
VUS-PR
TS-Router
41.49
43.09
51.97
76.92
22.40
65.42
73.18
71.65
39.60
56.91
77.32
56.36
SATzilla
29.32
0.97
38.66
39.31
24.97
7.88
77.81
62.13
20.21
43.48
29.34
34.01
ARGOSMART
27.68
24.70
32.40
52.31
9.30
24.70
76.93
57.39
20.83
44.15
36.26
36.97
MetaOD
27.79
0.97
40.12
39.31
24.97
7.88
77.81
63.46
20.43
44.46
29.37
34.23
ISAC
31.77
24.70
41.41
41.83
37.05
7.88
78.19
64.33
25.04
40.10
30.47
38.43
Appendix
Table 13: Leave-one-dataset-out transfer across eleven univariate benchmarks. Classical selectors are trained on the other ten real benchmark collections and deploy their top-ranked specialist (Top-1). TS-Router uses its fixed routing policy learned on Syn-RCD and default Top-3 normalized mean fusion, without real-data retraining of the router. All methods are evaluated on the entirely held-out target dataset. Best results are in bold and second-best results are underlined .
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing 100876, China · China Telecom Research Institute Beijing, China · Department of Computer Science, Missouri University of Science and Technology, Rolla, MO 65409 USA