Time series anomaly detection (TSAD) is increasingly deployed in streaming settings, where data arrive sequentially and may exhibit non-stationarity. As a result, several works from the recent literature propose streaming anomaly detection methods that rely on incremental updates to adapt over time. However, most of these approaches originate from the streaming outlier detection literature and largely ignore core characteristics of time series anomalies. Moreover, their empirical evaluation is typically conducted on synthetic or small-scale benchmarks with limited diversity, making it unclear whether streaming methods are truly advantageous in realistic TSAD scenarios. In this work, we carry out the first large-scale experimental study comparing streaming and static TSAD methods under a unified streaming evaluation benchmark. We consider a realistic setting in which an initial batch of data is available for model training, followed by online evaluation of both detection accuracy and computational efficiency. In addition, we propose a distribution-drift dataset of real time series, called TSB- drift, to isolate scenarios where streaming updates are theoretically justified. Our results show that, contrary to common assumptions, static TSAD methods significantly outperform streaming approaches in most streaming settings. Such finding highlights a critical gap between the design of existing streaming methods and the requirements of modern TSAD, and calls for a rethinking of how streaming capabilities should be integrated into TSAD.
Figures & tables
Figure 1. Static , Online , and Streaming Time Series Anomaly Detection
Figure 2. Illustration of TSB- drift Construction
General
Drift
Dataset
Category (Field)
# TS
μ(D)
# TS
ratio
Type
Genesis
Sensor (Robotics)
1
18
0/1
0%
∅
MITDB
Medical
13
13
0/13
0%
∅
PSM
Facility
1
25
0/1
0%
∅
SVDB
Medical
31
2
0/31
0%
∅
MSL
Sensor (Aerospace)
16
16
0/16
0%
∅
Table 1. Datasets in StrAD . Drift : # of TS with Concept Drift, % of Dimensions with Drifts (ratio), and Type
Table 2. Characteristic Patterns Throughout TSB- drift (Training Batch B0 is Highlighted in Green)
Acronym
Method
Type
Complexity
Distance-based
LOF
LOF ( Breunig et al., 2000 )
Proximity
O(∣T∣×D)∗
KNN
k -NN ( Hawkins, 1980 )
Proximity
O(∣T∣×D)∗
KMAD
k -Means ( Hawkins, 1980 )
Clustering
O(D)
CBLOF
CBLOF ( He et al., 2003 )
Clustering
O(D)
Density-based
Table 3. Static/Online TSAD Methods in StrAD ( ∗ : For Online settings, ∣T∣ is the Size of the Initial Batch B0 )
Acronym
Method
UM (Sec 2.2.2 )
MM (Sec 2.2.3 )
Complexity
Numerical
LODA
LODA ( Pevný, 2016 )
Projections
Tumbling Window
O(D)
xS
xStream ( Manzoor et al., 2018 )
Projections
Tumbling Window
O(D)
RSH
RSHash ( Sathe and Aggarwal, 2018 )
Partitioning
Sliding Window (Point)
O(1)∗
HST
HSTree ( Tan et al., 2011 )
Partitioning
Tumbling Window
O(1)∗
SDOs
SDOstream ( Hartl et al., 2020 )
Partitioning
Soft Forgetting (Aging)
O(D)
Table 4. Streaming TSAD in StrAD ( ∗ : Independent to D but Dependent to model hyper-parameters)
Figure 3. Static vs. Online Accuracy Evaluation on TSB-AD-M: (1) Overall, (2) by Data Category.
Figure 4. (1) AUC-PR gain from Static to Online Settings; (2) Online vs Static on TSB-AD-M for (2a) LOF, (2b) AE, (2c) IF.
Figure 5. Online vs. Streaming Accuracy on TSB-AD-M: (1) Overall, (2) by Data Category, (3) Critical Difference Diagrams ( α=0.05 ) on (3a) TSB-AD-M and (3b) Point Anomalies Time Series.
Figure 6. (1) TSB- drift vs TSB- ¬ drift accuracy; Mean AUC-PR gain of TSB- drift on TSB- ¬ drift for (1a) Online models and (1b) Streaming models; (2) Mean performance on drift-types
Figure 7. Scalability Evaluation on TSB-AD-M: (1) Accuracy vs Throughput; (2) Throughput vs Number of Dimensions; (3) Inference time standard deviation on TSB- ¬ drift and TSB- drift for (3a) Online and (3b) Streaming models (Red line highlights median on TSB- ¬ drift and dotted line on TSB- drift )
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Hyperparameter
Search Range
RRCF
num_trees
{2, 4, 8, 16, 20, 25}
shingle_size
{2, 4, 8}
tree_size
{256, 512}
RSHash
sampling_points
{500, 1000}
decay
{0.01, 0.015, 0.02}
num_components
{50, 100, 200, 300}
Appendix
Table 5. Hyperparameter Search Grid for the Evaluated Streaming Models
EDF relies on continuous monitoring of its power plants to detect anomalies as soon as they occur. Given the absence of a universally optimal streaming method in unsupervised settings, we compare streaming methods with state-of-the-art TSAD models deployed online on a real nuclear power plant dataset. This work also evaluates Automated Anomaly Detection in a streaming context. Results show higher consistency for online TSAD and strong robustness from ensembling strategies.
Time series anomaly detection (TSAD) is essential for maintaining the reliability and security of IoT-enabled service systems. Existing methods require training one specific model for each dataset, which exhibits limited generalization capability across different target datasets, hindering anomaly detection performance in various scenarios with scarce training data. To address this limitation, foundation models have emerged as a promising direction. However, existing approaches either repurpose large language models (LLMs) or construct largescale time series datasets to develop general anomaly detection foundation models, and still face challenges caused by severe cross-modal gaps or in-domain heterogeneity. In this paper, we investigate the applicability of large-scale vision models to TSAD. Specifically, we adapt a visual Masked Autoencoder (MAE) pretrained on ImageNet to the TSAD task. However, directly transferring MAE to TSAD introduces two key challenges: overgeneralization and limited local perception. To address these challenges, we propose VAN-AD, a novel MAE-based framework for TSAD. To alleviate the over-generalization issue, we design an Adaptive Distribution Mapping Module (ADMM), which maps the reconstruction results before and after MAE into a unified statistical space to amplify discrepancies caused by abnormal patterns. To overcome the limitation of local perception, we further develop a Normalizing Flow Module (NFM), which combines MAE with normalizing flow to estimate the probability density of the current window under the global distribution. Extensive experiments on nine real-world datasets demonstrate that VAN-AD consistently outperforms existing state-of-the-art methods across multiple evaluation metrics.We make our code and datasets available at https://github.com/PenyChen/VAN-AD.
PengYu Chen, Shang Wan, Xiaohou Shi +3
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing 100876, China · China Telecom Research Institute Beijing, China · Department of Computer Science, Missouri University of Science and Technology, Rolla, MO 65409 USA
Although recent studies on time-series anomaly detection have increasingly adopted ever-larger neural network architectures such as transformers and foundation models, they incur high computational costs and memory usage, making them impractical for real-time and resource-constrained scenarios. Moreover, they often fail to demonstrate significant performance gains over simpler methods under rigorous evaluation protocols. In this study, we propose Patch-based representation learning for time-series Anomaly detection (PaAno), a lightweight yet effective method for fast and efficient time-series anomaly detection. PaAno extracts short temporal patches from time-series training data and uses a 1D convolutional neural network to embed each patch into a vector representation. The model is trained using a combination of triplet loss and pretext loss to ensure the embeddings capture informative temporal patterns from input patches. During inference, the anomaly score at each time step is computed by comparing the embeddings of its surrounding patches to those of normal patches extracted from the training time-series. Evaluated on the TSB-AD benchmark, PaAno achieved state-of-the-art performance, significantly outperforming existing methods, including those based on heavy architectures, on both univariate and multivariate time-series anomaly detection across various range-wise and point-wise performance measures.
Jinju Park, Seokho Kang
Department of Industrial Engineering, Sungkyunkwan University