Organizations: School of Information Science and Technology, Beijing University of Technology, China · Beijing Key Laboratory of Computational Intelligence and Intelligent System, Beijing University of Technology, China
Reconstruction-based methods are a dominant paradigm in time series anomaly detection (TSAD), however, their near-universal reliance on Mean Squared Error (MSE) loss results in statistically flawed reconstruction residuals. This fundamental weakness leads to noisy, unstable anomaly scores, hindering reliable detection. To address this, we propose Constrained Gaussian-Noise Optimization and Smoothing (COGNOS), a universal, model-agnostic enhancement framework that tackles this issue at its source. COGNOS introduces a novel Gaussian-White Noise Regularization strategy during training, which directly constrains the model's output residuals to conform to a Gaussian white noise distribution. This engineered statistical property creates the ideal precondition for our second contribution: Adaptive Residual Kalman Smoother that operates as a statistically robust estimator to denoise the raw anomaly scores. Extensive experiments on multiple benchmarks demonstrate that COGNOS consistently enhances the performance of state-of-the-art backbones significantly, validating the efficacy of coupling statistical regularization with adaptive filtering.
Figures & tables
Figure 1 : We analyze the TimesNet model on the SWaT dataset: (a) The resulting anomaly scores are highly noisy, leading to a high noise floor during normal operation which makes actual anomalies difficult to detect. (b) The autocorrelation plot shows significant temporal correlations persist in the residuals, suggesting that the model failed to capture all predictable patterns. (c) The Q-Q plot reveals that reconstruction residuals are strongly non-Gaussian, indicating that the underlying noise was poorly modeled.
Figure 2 : COGNOS: Backbones are trained with GWNR Loss, and ARKS is used during inference to generate stable anomaly scores
Models
Datasets
MSL
PSM
SWAN
Metircs
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Autoformer
Std-F1
0.7724
0.9321
0.9098
0.9759
0.7336
0.7950
Aff-F1
0.4008
0.9388
0.5292
0.8743
0.1828
0.3604
R-A-R
0.6942
0.7084
0.6667
0.6954
0.9398
0.9433
R-A-P
0.2366
0.2446
0.4851
0.5235
0.9362
0.9382
V-R
0.6684
0.6884
0.6349
0.6650
0.9532
0.9542
Table 1: Multi-metric evaluation of anomaly detection performance between the Vanilla method and COGNOS (Ours) across three datasets. Best results are highlighted in bold .
Datasets
GECCO
MSL
PSM
SMAP
SWAN
SWaT
UCR
Models
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Autoformer
0.3929
0.9426
0.4008
0.9388
0.5292
0.8743
0.6266
0.8521
0.1828
0.3604
0.3967
0.9571
0.4702
0.7427
CrossAD
0.3115
0.9320
0.4135
0.9386
0.5904
0.8628
0.4871
0.7857
0.1829
0.5374
0.2973
0.9511
0.3383
0.7412
DLinear
0.4549
0.9384
0.5654
0.9432
0.7728
0.8615
0.6728
0.8437
0.2268
0.4402
0.4283
0.9542
0.4782
0.7325
KANAD
0.3785
0.9465
0.5360
0.9157
0.6223
0.8628
0.6666
0.8491
0.2066
0.3936
0.2905
0.9543
0.4409
0.7279
LSTMAE
0.3265
0.9463
0.5501
0.9460
0.6798
0.8629
0.6183
0.7770
0.2193
0.4329
0.4677
0.9521
0.4074
0.7328
Table 2: Comparison of anomaly detection performance between the Vanilla method and COGNOS (Ours) across seven datasets. The table reports the Affiliated-F1 metric. Best results are highlighted in bold .
Datasets
GECCO
MSL
PSM
GWNR
ARKS
MA
LP
Std-F1
Aff-F1
Std-F1
Aff-F1
Std-F1
Aff-F1
✗
0.4738
0.3785
0.8115
0.5360
0.8946
0.6223
✗
✓
0.4740
0.4845
0.8083
0.5285
0.9433
0.6689
✓
0.3440
0.4098
0.8704
0.5917
0.9160
0.4910
✓
✓
0.3531
0.4123
0.1626
0.2564
0.7979
0.1845
✓
✓
0.3539
0.4236
0.1662
0.2584
0.8006
0.2302
Table 3: Ablation Study of COGNOS Components and Filtering Methods. The best results are highlighted in bold , and the second-best are underlined . The full experiment results and more details can be found in the Appendix D.3 .
Figure 3 : Qualitative Analysis on Signal Purity and Statistical Physics. Top Row: Anomaly score dynamics on GECCO and PSM datasets. Bottom Row: Statistical diagnostics of reconstruction residuals.
Datasets
Models
Autoformer
KANAD
TimesNet
Vanilla
COGNOS
Vanilla
COGNOS
Vanilla
COGNOS
GECCO
Train
62.53
+41.648
39.94
+48.4
101.34
+171.49
Inference
0.1782
+0.1274
0.0847
+0.1467
0.1694
+0.1513
PSM
Train
53.30
+82.66
97.31
+206.68
103.89
+257.55
Inference
0.2150
+0.1031
0.2459
+0.1387
0.2054
+0.1327
SWAN
Train
87.64
+164.11
216.54
+355.01
153.57
+359.27
Table 4: Computational Efficiency Analysis. We report the average execution time (ms) per iteration (Training) and per sample point (Inference), (+) denotes the additional overhead introduced by our COGNOS framework. We set batch size =128 and sequence length =128
Figure 4 : The impact of the Confidence Level ( 1−α ) on KAN-AD backbone across the SMAP and SWAN datasets.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Category
Dimension
Training
Validation
Test
AR (%)
GECCO
Water treatment
9
55,408
13,852
69,261
1.25
MSL
Spacecraft
55
46,653
11,664
73,729
10.5
SMAP
Spacecraft
1
108,146
27,037
427,617
12.8
PSM
Server Machine
25
105,984
26,497
87,841
27.8
SWAN
Space Weather
38
48,000
12,000
60,000
23.8
SWaT
Water treatment
51
396,000
99,000
449,919
12.1
Appendix
Table 5 : Dataset Details
Miscellaneous Configurations
Datasets
Hyper-parameter
Traning Process
Window Length
dModel
dMLP
Layers
Num Heads
Confidence
Learning Rate 6 6 6 Initial Learning Rate
Batch Size
Epochs
GECCO
128
128
128
3
8
0.9
1e-4
128
10
UCR
3
SWaT
10
MSL
96
3
Appendix
Table 6 : Hyperparameters Configurations
Figure 5 : Residual analysis on the SWaT dataset using TimesNet. Left: vanilla method, Right: COGNOS.
Figure 6 : Residual analysis on the PSM dataset using TimeMixer++. Left: vanilla method, Right: COGNOS.
Figure 7 : Residual analysis on the SMAP dataset using CrossAD. Left: vanilla method, Right: COGNOS.
Figure 8 : Qualitative comparison of anomaly scores on PSM using TimeMixer++. Left: vanilla method, Right: COGNOS.
Figure 9 : Qualitative comparison of anomaly scores on GECCO using ModernTCN. Left: vanilla method, Right: COGNOS.
Figure 10 : Qualitative comparison of anomaly scores on PSM using KAN-AD. Left: vanilla method, Right: COGNOS.
Models
Datasets
GECCO
MSL
PSM
SMAP
SWAN
Metircs
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Autoformer
Std-F1
0.4826
0.7507
0.7724
0.9321
0.9098
0.9759
0.6847
0.7829
0.7336
0.7950
Aff-F1
0.3929
0.9426
0.4008
0.9388
0.5292
0.8743
0.6266
0.8521
0.1828
0.3604
R-A-R
0.9758
0.9825
0.6942
0.7084
0.6667
0.6954
0.5662
0.6109
0.9398
0.9433
R-A-P
0.5918
0.5471
0.2366
0.2446
0.4851
0.5235
0.1796
0.1652
0.9362
0.9382
V-R
0.9930
0.9920
0.6684
0.6884
0.6349
0.6650
0.5552
0.5882
0.9532
0.9542
Appendix
Table 7: Multi-metric evaluation of anomaly detection performance between the Vanilla method and COGNOS (Ours) across six datasets. Best results are highlighted in bold .
Datasets
Models
Autoformer
CrossAD
DLinear
KANAD
LSTMAE
Metircs
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
Vanilla
Ours
GECCO
Std-F1
0.4826
0.7507
0.4401
0.7228
0.4326
0.7400
0.4738
0.7466
0.4442
0.7619
Aff-F1
0.3929
0.9426
0.3115
0.9320
0.4549
0.9384
0.3785
0.9465
0.3265
0.9463
MSL
Std-F1
0.7724
0.9321
0.7320
0.9315
0.8196
0.9322
0.8115
0.9165
0.8108
0.9320
Aff-F1
0.4008
0.9388
0.4135
0.9386
0.5654
0.9432
0.5360
0.9157
0.5501
0.9460
PSM
Std-F1
0.9098
0.9759
0.9256
0.9740
0.9582
0.9747
0.8946
0.9744
0.9138
0.9759
Appendix
Table 8: Comparison of anomaly detection performance between the Vanilla method and Ours (Ours) across seven datasets. The table reports the Standard-F1 and Affiliated-F1 metric. Best results are highlighted in bold .
Datasets
GECCO
MSL
PSM
SMAP
SWAN
GWNR
ARKS
MA
LP
Std-F1
Aff-F1
Std-F1
Aff-F1
Std-F1
Aff-F1
Std-F1
Aff-F1
Std-F1
Aff-F1
✗
0.4738
0.3785
0.8115
0.5360
0.8946
0.6223
0.6237
0.6666
0.7384
0.2066
✗
✓
0.4740
0.4845
0.8083
0.5285
0.9433
0.6689
0.6035
0.6024
0.7403
0.2152
✗
✓
0.3886
0.4259
0.1601
0.2530
0.8546
0.3179
0.3529
0.3040
0.7321
0.1792
✗
✓
0.3985
0.4292
0.1951
0.2650
0.8566
0.4108
0.4372
0.4340
0.7321
0.1778
✓
0.3440
0.4098
0.8704
0.5917
0.9160
0.4910
0.6027
0.2980
0.7348
0.1872
Appendix
Table 9: Ablation Study of COGNOS Components and Filtering Methods. The best results are highlighted in bold , and the second-best are underlined .
Datasets
Models
Autoformer
KANAD
TimesNet
Vanilla
COGNOS
Vanilla
COGNOS
Vanilla
COGNOS
GECCO
Train
62.53
+41.648
39.94
+48.4
101.34
+171.49
Inference
0.1782
+0.1274
0.0847
+0.1467
0.1694
+0.1513
MSL
Train
65.32
+210.196
146.68
+354.92
97.57
+308.209
Inference
0.1829
+0.0831
0.2660
+0.144
0.1463
+0.1425
PSM
Train
53.30
+82.662
97.31
+206.687
103.89
+257.55
Appendix
Table 10: Computational Efficiency Analysis. We report the average execution time (ms) per iteration (Training) and per sample point (Inference). The value following “+” denotes the additional overhead introduced by our COGNOS framework. Note that the training overhead varies due to the data-dependent dynamic masking mechanism. We set batch size =128 and sequence length =128
Dataset
Precision
Recall
Std-F1
Aff-P
Aff-R
Aff-F1
PSM
0.9739
0.9550
0.9644
0.5429
0.7366
0.6251
MSL
0.9189
0.9515
0.9349
0.5187
0.9648
0.6746
SWaT
0.9024
0.9982
0.9479
0.5576
0.9760
0.7097
SMAP
0.9247
0.9745
0.9489
0.5125
0.9814
0.6733
GECCO
0.3737
0.4781
0.4195
0.5630
0.8612
0.6809
Appendix
Table 11: Anomaly Transformer performance under the unified quantile threshold protocol across five representative datasets.
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing 100876, China · China Telecom Research Institute Beijing, China · Department of Computer Science, Missouri University of Science and Technology, Rolla, MO 65409 USA