Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised objectives or surrogate abnormal patterns, providing limited supervision for context-dependent normal--anomalous distinctions. We propose Context-Anchored Pair Supervision (CAPS), a supervision-recovery framework for TSAD. CAPS views ideal anomaly supervision as a matched comparison between normal and anomalous outcomes under the same temporal context, and seeks to recover such supervision without target-domain anomaly labels. Using simulated normal--anomalous pairs, CAPS learns structure and anomaly-semantic representations through reconstruction, background consistency, and within-pair counterfactual recombination. The resulting anomaly representations form a continuous semantic space with coarse modes and induce a sampleable multimodal prior. CAPS conditionally realizes sampled semantics as residual-form effects on target reference trajectories. The resulting context-anchored normal--anomalous counterparts provide temporal supervision for discriminative detector learning. Experiments on nine datasets show that CAPS achieves the strongest aggregate performance across all four evaluation metrics among the compared methods, while complementary ablations and transfer analyses support the roles of context anchoring, semantic disentanglement, and conditional realization.
Figures & tables
Figure 1: Overall illustration of CAPS. Panel (a) illustrates the matched-supervision view of temporal anomalies. Panel (b) presents the CAPS framework: black arrows denote learning from simulated correspondences, while purple arrows denote target-domain counterpart generation.
Method
Aff.-F
F1T
Std.-F1
VUS-PR
Avg.R
T1
T2
Score ↑
Rank ↓
Score ↑
Rank ↓
Score ↑
Rank ↓
Score ↑
Rank ↓
Unsupervised TSAD baselines
LOF
72.26
8.94
28.19
7.44
23.11
8.33
23.19
8.06
8.19
0
2
IForest
47.86
10.67
14.15
10.94
15.65
9.72
18.72
8.67
10.00
0
0
OmniAnomaly
73.48
7.78
28.35
5.56
23.82
6.67
25.04
6.28
6.57
1
6
TranAD
73.39
8.67
21.46
7.67
19.00
8.44
22.54
7.28
8.01
0
0
Table 1: Overall performance across nine TSAD benchmarks. Score and Rank denote the unweighted mean score and average dataset-wise rank, respectively. Avg.R averages the ranks across all four metrics. T1 and T2 count first-place and top-two results across the 36 dataset–metric pairs. Best and second-best aggregate results are bold and underlined. Detailed per-dataset results are reported in Appendix C.5 .
Variant
Affiliation-F
F1T
Standard-F1
VUS-PR
Shuffled residuals
79.27
41.21
37.78
41.30
w/o residual parameterization
82.76
40.96
38.35
42.27
Direct source-residual injection
81.58
43.16
39.96
42.58
w/o counterfactual loss
83.65
42.76
39.11
43.11
CAPS
85.33
46.53
42.91
46.82
Table 2: Ablations and controlled comparisons across nine datasets. Each entry is the unweighted average of the corresponding dataset-level scores. Higher is better; the best results are bolded. Detailed analyses are provided in Appendix D .
Figure 2: Learned representations and cross-context semantic recombination. (a) Projections of zs and za . (b) Each column reuses a fixed anomaly-semantic code across different references; colored diagonal cells indicate the original pairings.
Table 3: Main architectural and optimization configurations of CAPS.
Name
Domain
#TS
Avg. Length
AR (%)
UCR
Misc.
228
67818.7
0.6
NAB
Web
28
5099.7
10.6
YAHOO
Web
259
1560.2
0.6
IOPS
Operations
17
72792.3
1.3
MGAB
Sensor
9
97777.8
0.2
SED
Energy
3
23332.3
4.1
Appendix
Table 4: Benchmark datasets used in the main experiments.
Metric
Method
IOPS
MGAB
NAB
Power
SED
UCR
NEK
TODS
YAHOO
Affiliation-F
LOF
81.06
68.44
75.75
66.76
63.85
73.53
84.74
60.58
75.63
IForest
52.81
68.82
39.84
0.00
70.09
50.56
71.15
44.17
33.30
OmniAnomaly
80.32
67.35
92.35
78.16
61.26
73.53
86.30
50.73
71.31
TranAD
83.19
67.28
90.28
71.56
61.03
73.31
85.02
52.76
76.08
USAD
71.08
67.81
91.54
76.48
55.60
76.00
71.13
47.90
53.05
AnomTrans.
70.79
67.65
79.03
71.57
68.21
80.03
75.32
44.57
64.75
Appendix
Table 5: Performance on nine datasets under four metrics (higher is better). CAPS variants are averaged over five runs. Best and second-best results excluding CAPS(Diagnosis) are shown in bold and underline, respectively; a star marks the better result between CAPS and CAPS(Diagnosis) for each dataset–metric pair.
Metric
Method
IOPS
MGAB
NAB
Power
SED
UCR
NEK
TODS
YAHOO
Affiliation-F
CAPS
89.24
69.22
93.25
85.98
74.09
88.41
86.30
86.85
94.63
Shuffled residuals
83.41
68.04
86.84
78.81
70.16
81.99
81.11
75.03
88.06
F1T
CAPS
60.62
7.58
57.90
29.20
24.27
41.53
80.15
46.17
71.35
Shuffled residuals
53.32
1.76
52.75
20.38
22.88
37.18
74.98
38.15
69.50
Standard-F1
CAPS
51.63
4.54
51.78
29.37
24.11
37.70
76.23
41.24
69.58
Shuffled residuals
46.85
1.62
46.31
20.38
22.81
32.29
68.40
34.12
67.26
Appendix
Table 6: Effect of disrupting context–effect matching in CAPS. Higher is better. The better result for each dataset–metric pair is bolded.
Metric
Method
IOPS
MGAB
NAB
Power
SED
UCR
NEK
TODS
YAHOO
Affiliation-F
CAPS
89.24
69.22
93.25
85.98
74.09
88.41
86.30
86.85
94.63
w/o residual parameterization
84.91
68.07
90.61
85.11
77.54
87.26
81.16
77.92
92.27
F1T
CAPS
60.62
7.58
57.90
29.20
24.27
41.53
80.15
46.17
71.35
w/o residual parameterization
44.06
1.26
54.96
24.03
19.96
39.59
76.53
39.91
68.37
Standard-F1
CAPS
51.63
4.54
51.78
29.37
24.11
37.70
76.23
41.24
69.58
w/o residual parameterization
37.98
1.15
48.91
24.02
19.87
35.67
72.43
35.78
69.34
Appendix
Table 7: Effect of removing residual-form counterpart construction in CAPS. Higher is better. The better result for each dataset–metric pair is bolded.
Method
Affiliation-F
F1T
Standard-F1
VUS-PR
Direct source-residual injection
81.58
43.16
39.96
42.58
CAPS
85.33
46.53
42.91
46.82
Appendix
Table 8: Comparison with synthetic supervision constructed from the same simulated anomaly source. Results are averaged over the nine benchmark datasets. Higher is better.
Dataset
Method
Affiliation-F
F1T
Standard-F1
VUS-PR
YAHOO
TCN-Supervised (BCE)
83.11
59.52
51.24
44.28
Weighted BCE
85.10
59.90
56.90
45.70
Balanced Window Sampling
83.10
37.30
34.50
39.30
Focal Loss
83.70
38.50
35.70
40.30
CAPS
94.63
71.35
69.58
83.70
NEK
TCN-Supervised (BCE)
85.24
72.93
62.58
67.52
Appendix
Table 9: Comparison with label-supervised TCN training under different class-imbalance treatments. All methods use the same TCN detector backbone. Higher is better.
Contamination
Affiliation-F
F1T
Standard-F1
VUS-PR
0%
85.74
34.17
37.54
34.05
1%
84.29
33.08
34.98
31.19
2%
83.94
33.05
35.48
30.77
5%
84.48
31.06
32.69
29.44
10%
84.53
30.83
31.65
29.46
20%
84.20
30.61
29.22
26.52
Appendix
Table 10: Robustness of CAPS to controlled target-reference contamination on IOPS using 100 reference windows per file. Results are averaged over 17 time-series files and five runs. Higher is better.
Variant
Affiliation-F
F1T
Standard-F1
VUS-PR
w/o Lcf
83.65
42.76
39.11
43.11
CAPS
85.33
46.53
42.91
46.82
Appendix
Table 11: Effect of within-pair counterfactual recomposition. Results are averaged over the nine benchmark datasets. Higher is better.
Figure 3: Learned representation spaces. Left : structure codes zs colored by normal/anomalous status. Right : anomaly-semantic codes za colored by simulated anomaly family.
Figure 4: Cross-context semantic recombination. Rows use different references, and each column reuses a fixed anomaly-semantic code za . Colored diagonal cells mark the original pairings. Dashed curves denote references, solid curves denote generated counterparts, and shaded regions indicate anomaly supports.
Method
Affiliation-F
F1T
Standard-F1
VUS-PR
CAPS w/o highest-associated mode
79.31
38.42
37.81
39.87
CAPS
85.33
46.53
42.91
46.82
Appendix
Table 12: Semantic-coverage stress test. The source-side mode with the highest target association is excluded during counterpart generation while the total number of counterparts is preserved. Results are averaged over the nine benchmark datasets. Higher is better.
Organization
K
Affiliation-F
F1T
Standard-F1
VUS-PR
Data-driven
1
82.34
39.92
37.23
41.36
Data-driven
3
84.98
43.75
40.12
45.76
Data-driven
5
83.10
44.04
40.43
44.76
Data-driven
7
81.61
42.38
38.78
43.38
Knowledge-based
3
85.33
46.53
42.91
46.82
Appendix
Table 13: Comparison of semantic-mode organization and granularity. Results are averaged over the nine benchmark datasets. Higher is better.
Figure 5: Router associations on simulated samples. Rows denote source-side modes and columns denote router-associated modes. Values are row-normalized percentages, with sample counts in parentheses.
Strategy
Affiliation-F
F1T
Standard-F1
VUS-PR
Direct source-residual injection
81.58
43.16
39.96
42.58
VAE residual generator
83.28
42.32
38.82
42.63
CAPS (diffusion)
85.33
46.53
42.91
46.82
Appendix
Table 14: Comparison of anomaly-effect construction strategies. Results are averaged over the nine benchmark datasets. Higher is better.
Positive source
Score ↓
Feature ↓
Rule-based injection
0.4531
0.6218
Direct source-residual injection
0.4783
0.6712
CAPS
0.3990
0.6123
Appendix
Table 15: Detector-oriented diagnostics against real target anomalies, averaged over the nine benchmark datasets. Score measures Wasserstein distance between anomaly-score distributions; Feature measures average generated-to-real 5 -NN distance in the penultimate-layer feature space. Lower is better.
Metric
CAPS (TCN)
CAPS (Transformer)
Affiliation-F
85.33
81.10
F1T
46.53
42.12
Standard-F1
42.91
40.37
VUS-PR
46.82
43.28
Appendix
Table 16: Comparison of CAPS under different detector architectures. Results are averaged over the nine benchmark datasets. Higher is better.
Prior
Affiliation-F
F1T
Standard-F1
VUS-PR
Isotropic Gaussian
84.97
45.88
42.36
46.21
Diagonal Gaussian
85.33
46.53
42.91
46.82
Full-covariance Gaussian
85.40
47.11
42.74
46.70
Student- t
85.21
46.47
42.68
46.95
Appendix
Table 17: Sensitivity to the parametric form of the mode-wise anomaly-semantic prior. Results are averaged over the nine benchmark datasets. Higher is better.
Metric
5,000
10,000
20,000
48,000
Affiliation-F
77.66
78.93
82.15
85.33
F1T
44.36
47.24
46.21
46.53
Standard-F1
41.05
43.52
42.78
42.91
VUS-PR
39.68
42.40
44.79
46.82
Appendix
Table 18: Performance under different amounts of simulated training data, averaged over the nine benchmark datasets. Higher is better.
Number of simulated pairs
Simulated-domain training
Pair generation
Test-time detection
5,000
23.6 min
50 ms / pair
2 ms / sequence
10,000
48.5 min
50 ms / pair
2 ms / sequence
20,000
95.1 min
50 ms / pair
2 ms / sequence
48,000
180.5 min
50 ms / pair
2 ms / sequence
Appendix
Table 19: Computational cost under different amounts of simulated training data. Pair generation is performed offline; test-time detection uses only the trained TCN.
Name
Domain
#TS
#Dim
Avg. Length
AR (%)
MSL
Space
16
55
3119.4
5.1
PSM
Sensor
1
25
217624.0
11.2
SMAP
Space
27
25
7855.9
2.9
SMD
Server
22
38
25466.4
3.8
Appendix
Table 20: Multivariate datasets used in the supplementary experiments. #TS denotes the number of evaluated multivariate sequences, #Dim the number of input variables per sequence, and AR the anomaly ratio.
Metric
Model
MSL
PSM
SMAP
SMD
Affiliation-F
LOF
84.35
61.98
63.32
64.13
IForest
63.36
63.78
59.96
69.71
OmniAnomaly
83.15
58.17
91.38
85.82
TranAD
79.91
73.83
87.39
92.20
USAD
81.86
57.86
87.25
85.09
AnomTrans.
74.38
66.04
74.82
73.44
Appendix
Table 21: Results on multivariate anomaly detection benchmarks. Higher is better. Best and second-best results within each dataset–metric pair are shown in bold and underline , respectively.
German Research Center for Artificial Intelligence (DFKI), Kaiserslautern, Germany · Rheinland-Pfälzische Technische Universität Kaiserslautern-Landau, Germany · Paul Wurth S.A, Luxembourg
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing 100876, China · China Telecom Research Institute Beijing, China · Department of Computer Science, Missouri University of Science and Technology, Rolla, MO 65409 USA