Reconstruction errors in multivariate time-series anomaly detection may not reliably distinguish abnormal behavior from benign deviations. Language-derived semantics offer complementary context, but existing multimodal approaches may rely on time-associated paired textual information that is difficult to obtain consistently and is not provided by standard multivariate time-series anomaly detection benchmarks. This setting poses two challenges: (1) conditioning masked reconstruction on window-specific semantics without exposing exact numerical targets or anomaly-specific cues, and (2) using a window-independent concept of normality as a complementary semantic reference rather than an independent anomaly detector. We propose LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Time Series (LEARN-TS), which uses a frozen language model to construct two role-separated semantic representations without requiring temporally paired external text. Window-specific observation semantics encode temporal and cross-variable context without exact numerical values to guide channel-shared patch-masked reconstruction. A fixed, dataset-agnostic normality prompt provides a window-independent semantic reference for aligning normal representations and estimating normality discrepancy. At inference, masking each temporal patch once yields timestamp-level reconstruction evidence, conditionally modulated by discrepancy from a separate unmasked view. Across four benchmarks, LEARN-TS achieves the highest mean performance in 13 of 16 dataset-metric comparisons. Controlled ablations examine observation conditioning, joint normality alignment and scoring, and reference content, showing dataset-dependent ranking benefits and modest average gains from semantic over random references.
Figures & tables
Figure 1: (a) Time-associated paired external textual context may vary across datasets and monitored intervals, providing context-dependent semantic information. (b) LEARN-TS instead derives a shared, window-independent normality reference from a fixed, dataset-agnostic prompt applied unchanged across all datasets and windows.
Figure 2: Overview of the three-phase LEARN-TS framework.
Dataset
Metric
Omni
A.T.
Times
Patch
iTrans
G4TS
TLLM
LMixer
CALF
DADA
LEARN-TS
SWaT
A-PR
0.1492
0.6829
0.1341
0.0890
0.0931
0.0894
0.0846
0.0866
0.0819
0.5385
0.7578
V-PR
0.1371
0.4508
0.1656
0.0982
0.1008
0.0991
0.0930
0.0951
0.0907
0.4543
0.4859
R-F1
0.1718
0.1602
0.1593
0.1097
0.1277
0.1344
0.1190
0.1201
0.1137
0.0990
0.3102
Aff-F1
0.7137
0.7413
0.7940
0.6376
0.6621
0.6826
0.6658
0.6737
0.6570
0.7119
0.7277
SMD
A-PR
0.3713
0.2575
0.4640
0.4893
0.4726
0.4947
0.4514
0.4307
0.4578
0.4754
0.4930
V-PR
0.3752
0.3214
0.6184
0.6376
0.6250
0.6455
0.5839
0.5662
0.5910
0.5912
0.5398
Table 1: Performance comparison on four real-world multivariate time-series anomaly detection datasets. Results are averaged over five runs; standard deviations are reported in Table 7 . For each dataset–metric pair, the best and second-best results are shown in bold and underlined , respectively.
Figure 3: Effect of observation–window correspondence: (a)–(b) matched vs. shuffled observations on SWaT and MSL; (c) A-PR and V-PR gains across all four datasets.
Figure 4: Event-level predictions on three SMD entities: Machine 1-3 with intermediate event density, Machine 2-3 with sparse events, and Machine 2-9 with dense events.
Figure 5: Sensitivity to (a) λfull , (b) λnorm , (c) λgate , and (d) the number of masked training patches. Results are four-dataset macro averages using a fixed seed.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Domain
Entities
Variables
Official train
Official test
Anomaly (%)
SWaT
Industrial control
1
51
496,800
449,919
11.98
SMD
Server monitoring
28
38
708,405
708,420
4.16
PSM
Application server
1
25
132,481
87,841
27.75
MSL
Spacecraft telemetry
27
55
58,317
73,729
10.72
Appendix
Table 3: Dataset statistics before the internal training–validation split. For SMD and MSL, sequence lengths are summed over all entities.
Hyperparameter
Value
Input window length L
128
Patch length ℓ
16
Number of patches P
8
Hidden dimension d
768
Time-series encoder layers
3
Cross-modal fusion layers
2
Appendix
Table 4: Common model and optimization settings used across all datasets.
Dataset
λnorm
λgate
SWaT
0.02
1.00
SMD
0.01
0.05
PSM
0.01
0.1
MSL
0.01
0.05
Appendix
Table 5: Dataset-specific loss and scoring configurations.
Dataset
Channel aggregation GD
Window overlap
SWaT
Mean over all variables
Mean
SMD
Top- k mean
Maximum
PSM
Top- k mean
Maximum
MSL
Target telemetry channel
Mean
Appendix
Table 6: Dataset-specific channel error and overlapping-window aggregation rules.
Method
A-PR
V-PR
R-F1
Aff-F1
(a) SWaT
OmniAnomaly
0.1492 ± 0.0559
0.1371 ± 0.0321
0.1718 ± 0.1140
0.7137 ± 0.0041
Anomaly Transformer
0.6829 ± 0.0193
0.4508 ± 0.0077
0.1602 ± 0.0187
0.7413 ± 0.0149
TimesNet
0.1341 ± 0.0018
0.1656 ± 0.0021
0.1593 ± 0.0057
0.7940 ± 0.0091
PatchTST
0.0890 ± 0.0002
0.0982 ± 0.0002
0.1097 ± 0.0062
0.6376 ± 0.0026
iTransformer
0.0931 ± 0.0010
0.1008 ± 0.0005
0.1277 ± 0.0029
0.6621 ± 0.0025
Appendix
Table 7: Performance of all methods on the four datasets over five independent runs. Results are reported as mean ± sample standard deviation. For each dataset–metric pair, the best and second-best mean results are shown in bold and underlined , respectively.
Dataset
Variant
A-PR
V-PR
R-F1
Aff-F1
SWaT
w/o Obs.
0.6552 ± 0.0652
0.3727 ± 0.0193
0.2446 ± 0.0110
0.7047 ± 0.0012
Concat fusion
0.7001 ± 0.1014
0.4170 ± 0.0508
0.2911 ± 0.0189
0.7270 ± 0.0090
LEARN-TS (Full)
0.7578 ± 0.0117
0.4859 ± 0.0346
0.3102 ± 0.0436
0.7277 ± 0.0069
SMD
w/o Obs.
0.4650 ± 0.0074
0.5266 ± 0.0047
0.3222 ± 0.0070
0.8137 ± 0.0102
Concat fusion
0.4829 ± 0.0122
0.5380 ± 0.0098
0.3262 ± 0.0057
0.8174 ± 0.0084
LEARN-TS (Full)
0.4930 ± 0.0056
0.5398 ± 0.0037
0.3203 ± 0.0033
0.8121 ± 0.0056
Appendix
Table 8: Dataset-wise cross-modal fusion ablation. Results are reported as mean ± sample standard deviation over five random seeds. The best result in each dataset–metric column is shown in bold .
Variant
Align.
Score
A-PR
V-PR
R-F1
Aff-F1
(a) SWaT
No normality
✗
✗
0.2088 ± 0.0076
0.1748 ± 0.0010
0.2971 ± 0.0611
0.7407 ± 0.0157
Scoring only
✗
✓
0.1460 ± 0.0531
0.1530 ± 0.0264
0.2711 ± 0.0635
0.7103 ± 0.0271
Alignment only
✓
✗
0.2004 ± 0.0094
0.1766 ± 0.0008
0.2646 ± 0.0211
0.7299 ± 0.0097
LEARN-TS (Full)
✓
✓
0.7578 ± 0.0117
0.4859 ± 0.0346
0.3102 ± 0.0436
0.7277 ± 0.0069
(b) SMD
Appendix
Table 9: Dataset-wise ablation of normality alignment and discrepancy-guided scoring. Results are reported as mean ± sample standard deviation over five random seeds. The best result in each dataset–metric column is shown in bold .
Normality reference
A-PR
V-PR
R-F1
Aff-F1
(a) SWaT
Random semantic-free
0.7431 ± 0.0426
0.4758 ± 0.0364
0.2933 ± 0.0182
0.7276 ± 0.0017
Dataset-agnostic semantic
0.7578 ± 0.0117
0.4859 ± 0.0346
0.3102 ± 0.0436
0.7277 ± 0.0069
(b) SMD
Random semantic-free
0.4919 ± 0.0052
0.5388 ± 0.0052
0.3153 ± 0.0058
0.8022 ± 0.0143
Dataset-agnostic semantic
0.4930 ± 0.0056
0.5398 ± 0.0037
0.3203 ± 0.0033
0.8121 ± 0.0056
Appendix
Table 10: Dataset-wise comparison between random semantic-free and dataset-agnostic semantic normality references. Results are reported as mean ± sample standard deviation over five random seeds. The best result in each dataset–metric column is shown in bold ; values tied at the displayed precision are both highlighted.
(a) Full-reconstruction weight λfull
Dataset
Value
A-PR
V-PR
R-F1
Aff-F1
SWaT
0.0
0.1766
0.1738
0.2920
0.7381
0.1
0.7654
0.5004
0.2793
0.7121
0.2
0.7662
0.5122
0.3293
0.7332
0.5
0.7684
0.5139
0.3784
0.7343
1.0
0.7665
0.5023
0.2655
0.7276
Appendix
Table 11: Dataset-wise sensitivity to (a) the full-reconstruction weight λfull and (b) the normality-alignment weight λnorm . Results are obtained using a fixed random seed. Bold rows denote the configurations used in the main experiments.
(a) Discrepancy-gating coefficient λgate
Dataset
Value
A-PR
V-PR
R-F1
Aff-F1
SWaT
0.00
0.2132
0.1763
0.2607
0.7387
0.01
0.7622
0.5137
0.2786
0.7330
0.05
0.7640
0.5138
0.2861
0.7338
0.10
0.7670
0.5153
0.3121
0.7337
0.50
0.7686
0.5138
0.3299
0.7394
Appendix
Table 12: Dataset-wise sensitivity to (a) the discrepancy-gating coefficient λgate and (b) the number of masked training patches. Results are obtained using a fixed random seed. Bold rows denote the configurations used in the main experiments.
Dataset
Backbone
Trainable
Peak Train GPU
Embed.
Scoring
Params ↓
(GiB) ↓
(prompts/s) ↑
(windows/s) ↑
SWaT
GPT-2
41.16M
0.795
80.90
1111.46
Qwen2.5-1.5B
41.75M
0.811
38.81
932.19
Llama-3.2-1B
42.14M
0.822
53.18
958.34
PSM
GPT-2
40.79M
0.792
80.15
1194.91
Qwen2.5-1.5B
41.38M
0.803
40.17
1152.70
Appendix
Table 13: Backbone efficiency and detection performance on four datasets. Detection results are mean ± sample standard deviation over five seeds. Embedding throughput measures prompt encoding; scoring throughput excludes text encoding and uses cached embeddings. Bold indicates the best result within each dataset and column.
Figure 6: Dataset-agnostic normality prompt used in LEARN-TS. The prompt is fixed before training and shared unchanged across all datasets and input windows.
Figure 7: Abridged SWaT observation prompt template. Blue placeholders denote window-dependent descriptors and patch-index lists; dataset-level descriptions and rules remain fixed. Exact measurements and explicit anomaly labels are excluded. Repeated entries are omitted for readability.
Multivariate time series anomaly detection has become increasingly important in real-world applications, where labeled data are often scarce. Many existing approaches rely on unsupervised learning to model normal patterns, but they often treat all channels equally. This design can dilute anomaly-relevant signals, since not all channels contribute equally to anomaly detection. In this paper, we propose CALAD, a channel-aware contrastive learning framework for multivariate time series anomaly detection. CALAD governs the construction of contrastive samples using estimated channel relevance, allowing the learning process to reflect anomaly semantics rather than generic similarity. Channel relevance is estimated from reconstruction errors of a transformer-based autoencoder and is used to distinguish channels that are more influential to anomalous behaviors. Using this information, we design a channel-wise augmentation strategy in which positive and negative samples are constructed based on whether anomaly-relevant channels are preserved or perturbed. This encourages invariance to changes in irrelevant channels while being sensitive to changes in anomaly-relevant channels. Furthermore, CALAD combines contrastive learning and an auxiliary reconstruction head, allowing the model to learn discriminative representations while retaining normal structures. Experiments on multiple real-world datasets shows that CALAD consistently outperforms existing methods, particularly under distribution shift scenarios. We provide the code for reproducibility at https://github.com/hirundo1218/CALAD
Jaehyeop Hong, Youngbum Hur
Department of Industrial Engineering, Inha University, Incheon, Republic of Korea
Time series anomaly detection (TSAD) remains challenging not only because anomaly labels are scarce, but also because temporal anomalies are highly context-dependent. Existing methods often rely on unsupervised objectives or surrogate abnormal patterns, providing limited supervision for context-dependent normal--anomalous distinctions. We propose Context-Anchored Pair Supervision (CAPS), a supervision-recovery framework for TSAD. CAPS views ideal anomaly supervision as a matched comparison between normal and anomalous outcomes under the same temporal context, and seeks to recover such supervision without target-domain anomaly labels. Using simulated normal--anomalous pairs, CAPS learns structure and anomaly-semantic representations through reconstruction, background consistency, and within-pair counterfactual recombination. The resulting anomaly representations form a continuous semantic space with coarse modes and induce a sampleable multimodal prior. CAPS conditionally realizes sampled semantics as residual-form effects on target reference trajectories. The resulting context-anchored normal--anomalous counterparts provide temporal supervision for discriminative detector learning. Experiments on nine datasets show that CAPS achieves the strongest aggregate performance across all four evaluation metrics among the compared methods, while complementary ablations and transfer analyses support the roles of context anchoring, semantic disentanglement, and conditional realization.
Yifei Gao, Tian Lan, Yimeng Lu +5
Department of Industrial Engineering, Tsinghua University · Huawei
The core challenge in unsupervised anomaly detection is identifying abnormal patterns without prior knowledge of their characteristics. While existing methods have addressed aspects of this problem, they often struggle to learn a robust representation of the normal data distribution that is distinct from anomalous patterns. In this paper, we present a novel framework, Unified Unsupervised Anomaly Detection (U2AD), that comprehensively addresses anomaly detection in multivariate time series. Our approach learns the underlying data distribution of normal samples by utilizing score-based generative modeling. We introduce a novel time-dependent score network and a unified training objective that together delineate the manifold of normal data while considering both local and global temporal contexts. Reconstruction is then performed via a deterministic sampling process using an ordinary differential equation solver. Our extensive experimental evaluations demonstrate that U2AD not only outperforms current state-of-the-art methods in detection accuracy but also identifies anomalies at significantly earlier stages of their occurrence.
Prithul Sarker, Sushmita Sarker, Nicholas G. Murray +1