Reconstruction errors in multivariate time-series anomaly detection may not reliably distinguish abnormal behavior from benign deviations. Language-derived semantics offer complementary context, but existing multimodal approaches may rely on time-associated paired textual information that is difficult to obtain consistently and is not provided by standard multivariate time-series anomaly detection benchmarks. This setting poses two challenges: (1) conditioning masked reconstruction on window-specific semantics without exposing exact numerical targets or anomaly-specific cues, and (2) using a window-independent concept of normality as a complementary semantic reference rather than an independent anomaly detector. We propose LLM-Enhanced Alignment and Reconstruction with Normality Guidance for Time Series (LEARN-TS), which uses a frozen language model to construct two role-separated semantic representations without requiring temporally paired external text. Window-specific observation semantics encode temporal and cross-variable context without exact numerical values to guide channel-shared patch-masked reconstruction. A fixed, dataset-agnostic normality prompt provides a window-independent semantic reference for aligning normal representations and estimating normality discrepancy. At inference, masking each temporal patch once yields timestamp-level reconstruction evidence, conditionally modulated by discrepancy from a separate unmasked view. Across four benchmarks, LEARN-TS achieves the highest mean performance in 13 of 16 dataset-metric comparisons. Controlled ablations examine observation conditioning, joint normality alignment and scoring, and reference content, showing dataset-dependent ranking benefits and modest average gains from semantic over random references.
Figures & tables
Figure 1: (a) Time-associated paired external textual context may vary across datasets and monitored intervals, providing context-dependent semantic information. (b) LEARN-TS instead derives a shared, window-independent normality reference from a fixed, dataset-agnostic prompt applied unchanged across all datasets and windows.
Figure 2: Overview of the three-phase LEARN-TS framework.
Dataset
Metric
Omni
A.T.
Times
Patch
iTrans
G4TS
TLLM
LMixer
CALF
DADA
LEARN-TS
SWaT
A-PR
0.1492
0.6829
0.1341
0.0890
0.0931
0.0894
0.0846
0.0866
0.0819
0.5385
0.7578
V-PR
0.1371
0.4508
0.1656
0.0982
0.1008
0.0991
0.0930
0.0951
0.0907
0.4543
0.4859
R-F1
0.1718
0.1602
0.1593
0.1097
0.1277
0.1344
0.1190
0.1201
0.1137
0.0990
0.3102
Aff-F1
0.7137
0.7413
0.7940
0.6376
0.6621
0.6826
0.6658
0.6737
0.6570
0.7119
0.7277
SMD
A-PR
0.3713
0.2575
0.4640
0.4893
0.4726
0.4947
0.4514
0.4307
0.4578
0.4754
0.4930
V-PR
0.3752
0.3214
0.6184
0.6376
0.6250
0.6455
0.5839
0.5662
0.5910
0.5912
0.5398
Table 1: Performance comparison on four real-world multivariate time-series anomaly detection datasets. Results are averaged over five runs; standard deviations are reported in Table 7 . For each dataset–metric pair, the best and second-best results are shown in bold and underlined , respectively.
Figure 3: Effect of observation–window correspondence: (a)–(b) matched vs. shuffled observations on SWaT and MSL; (c) A-PR and V-PR gains across all four datasets.
Figure 4: Event-level predictions on three SMD entities: Machine 1-3 with intermediate event density, Machine 2-3 with sparse events, and Machine 2-9 with dense events.
Figure 5: Sensitivity to (a) λfull , (b) λnorm , (c) λgate , and (d) the number of masked training patches. Results are four-dataset macro averages using a fixed seed.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Domain
Entities
Variables
Official train
Official test
Anomaly (%)
SWaT
Industrial control
1
51
496,800
449,919
11.98
SMD
Server monitoring
28
38
708,405
708,420
4.16
PSM
Application server
1
25
132,481
87,841
27.75
MSL
Spacecraft telemetry
27
55
58,317
73,729
10.72
Appendix
Table 3: Dataset statistics before the internal training–validation split. For SMD and MSL, sequence lengths are summed over all entities.
Hyperparameter
Value
Input window length L
128
Patch length ℓ
16
Number of patches P
8
Hidden dimension d
768
Time-series encoder layers
3
Cross-modal fusion layers
2
Appendix
Table 4: Common model and optimization settings used across all datasets.
Dataset
λnorm
λgate
SWaT
0.02
1.00
SMD
0.01
0.05
PSM
0.01
0.1
MSL
0.01
0.05
Appendix
Table 5: Dataset-specific loss and scoring configurations.
Dataset
Channel aggregation GD
Window overlap
SWaT
Mean over all variables
Mean
SMD
Top- k mean
Maximum
PSM
Top- k mean
Maximum
MSL
Target telemetry channel
Mean
Appendix
Table 6: Dataset-specific channel error and overlapping-window aggregation rules.
Method
A-PR
V-PR
R-F1
Aff-F1
(a) SWaT
OmniAnomaly
0.1492 ± 0.0559
0.1371 ± 0.0321
0.1718 ± 0.1140
0.7137 ± 0.0041
Anomaly Transformer
0.6829 ± 0.0193
0.4508 ± 0.0077
0.1602 ± 0.0187
0.7413 ± 0.0149
TimesNet
0.1341 ± 0.0018
0.1656 ± 0.0021
0.1593 ± 0.0057
0.7940 ± 0.0091
PatchTST
0.0890 ± 0.0002
0.0982 ± 0.0002
0.1097 ± 0.0062
0.6376 ± 0.0026
iTransformer
0.0931 ± 0.0010
0.1008 ± 0.0005
0.1277 ± 0.0029
0.6621 ± 0.0025
Appendix
Table 7: Performance of all methods on the four datasets over five independent runs. Results are reported as mean ± sample standard deviation. For each dataset–metric pair, the best and second-best mean results are shown in bold and underlined , respectively.
Dataset
Variant
A-PR
V-PR
R-F1
Aff-F1
SWaT
w/o Obs.
0.6552 ± 0.0652
0.3727 ± 0.0193
0.2446 ± 0.0110
0.7047 ± 0.0012
Concat fusion
0.7001 ± 0.1014
0.4170 ± 0.0508
0.2911 ± 0.0189
0.7270 ± 0.0090
LEARN-TS (Full)
0.7578 ± 0.0117
0.4859 ± 0.0346
0.3102 ± 0.0436
0.7277 ± 0.0069
SMD
w/o Obs.
0.4650 ± 0.0074
0.5266 ± 0.0047
0.3222 ± 0.0070
0.8137 ± 0.0102
Concat fusion
0.4829 ± 0.0122
0.5380 ± 0.0098
0.3262 ± 0.0057
0.8174 ± 0.0084
LEARN-TS (Full)
0.4930 ± 0.0056
0.5398 ± 0.0037
0.3203 ± 0.0033
0.8121 ± 0.0056
Appendix
Table 8: Dataset-wise cross-modal fusion ablation. Results are reported as mean ± sample standard deviation over five random seeds. The best result in each dataset–metric column is shown in bold .
Variant
Align.
Score
A-PR
V-PR
R-F1
Aff-F1
(a) SWaT
No normality
✗
✗
0.2088 ± 0.0076
0.1748 ± 0.0010
0.2971 ± 0.0611
0.7407 ± 0.0157
Scoring only
✗
✓
0.1460 ± 0.0531
0.1530 ± 0.0264
0.2711 ± 0.0635
0.7103 ± 0.0271
Alignment only
✓
✗
0.2004 ± 0.0094
0.1766 ± 0.0008
0.2646 ± 0.0211
0.7299 ± 0.0097
LEARN-TS (Full)
✓
✓
0.7578 ± 0.0117
0.4859 ± 0.0346
0.3102 ± 0.0436
0.7277 ± 0.0069
(b) SMD
Appendix
Table 9: Dataset-wise ablation of normality alignment and discrepancy-guided scoring. Results are reported as mean ± sample standard deviation over five random seeds. The best result in each dataset–metric column is shown in bold .
Normality reference
A-PR
V-PR
R-F1
Aff-F1
(a) SWaT
Random semantic-free
0.7431 ± 0.0426
0.4758 ± 0.0364
0.2933 ± 0.0182
0.7276 ± 0.0017
Dataset-agnostic semantic
0.7578 ± 0.0117
0.4859 ± 0.0346
0.3102 ± 0.0436
0.7277 ± 0.0069
(b) SMD
Random semantic-free
0.4919 ± 0.0052
0.5388 ± 0.0052
0.3153 ± 0.0058
0.8022 ± 0.0143
Dataset-agnostic semantic
0.4930 ± 0.0056
0.5398 ± 0.0037
0.3203 ± 0.0033
0.8121 ± 0.0056
Appendix
Table 10: Dataset-wise comparison between random semantic-free and dataset-agnostic semantic normality references. Results are reported as mean ± sample standard deviation over five random seeds. The best result in each dataset–metric column is shown in bold ; values tied at the displayed precision are both highlighted.
(a) Full-reconstruction weight λfull
Dataset
Value
A-PR
V-PR
R-F1
Aff-F1
SWaT
0.0
0.1766
0.1738
0.2920
0.7381
0.1
0.7654
0.5004
0.2793
0.7121
0.2
0.7662
0.5122
0.3293
0.7332
0.5
0.7684
0.5139
0.3784
0.7343
1.0
0.7665
0.5023
0.2655
0.7276
Appendix
Table 11: Dataset-wise sensitivity to (a) the full-reconstruction weight λfull and (b) the normality-alignment weight λnorm . Results are obtained using a fixed random seed. Bold rows denote the configurations used in the main experiments.
(a) Discrepancy-gating coefficient λgate
Dataset
Value
A-PR
V-PR
R-F1
Aff-F1
SWaT
0.00
0.2132
0.1763
0.2607
0.7387
0.01
0.7622
0.5137
0.2786
0.7330
0.05
0.7640
0.5138
0.2861
0.7338
0.10
0.7670
0.5153
0.3121
0.7337
0.50
0.7686
0.5138
0.3299
0.7394
Appendix
Table 12: Dataset-wise sensitivity to (a) the discrepancy-gating coefficient λgate and (b) the number of masked training patches. Results are obtained using a fixed random seed. Bold rows denote the configurations used in the main experiments.
Dataset
Backbone
Trainable
Peak Train GPU
Embed.
Scoring
Params ↓
(GiB) ↓
(prompts/s) ↑
(windows/s) ↑
SWaT
GPT-2
41.16M
0.795
80.90
1111.46
Qwen2.5-1.5B
41.75M
0.811
38.81
932.19
Llama-3.2-1B
42.14M
0.822
53.18
958.34
PSM
GPT-2
40.79M
0.792
80.15
1194.91
Qwen2.5-1.5B
41.38M
0.803
40.17
1152.70
Appendix
Table 13: Backbone efficiency and detection performance on four datasets. Detection results are mean ± sample standard deviation over five seeds. Embedding throughput measures prompt encoding; scoring throughput excludes text encoding and uses cached embeddings. Bold indicates the best result within each dataset and column.
Figure 6: Dataset-agnostic normality prompt used in LEARN-TS. The prompt is fixed before training and shared unchanged across all datasets and input windows.
Figure 7: Abridged SWaT observation prompt template. Blue placeholders denote window-dependent descriptors and patch-index lists; dataset-level descriptions and rules remain fixed. Exact measurements and explicit anomaly labels are excluded. Repeated entries are omitted for readability.