Time series anomaly detection (TSAD) is increasingly deployed in streaming settings, where data arrive sequentially and may exhibit non-stationarity. As a result, several works from the recent literature propose streaming anomaly detection methods that rely on incremental updates to adapt over time. However, most of these approaches originate from the streaming outlier detection literature and largely ignore core characteristics of time series anomalies. Moreover, their empirical evaluation is typically conducted on synthetic or small-scale benchmarks with limited diversity, making it unclear whether streaming methods are truly advantageous in realistic TSAD scenarios. In this work, we carry out the first large-scale experimental study comparing streaming and static TSAD methods under a unified streaming evaluation benchmark. We consider a realistic setting in which an initial batch of data is available for model training, followed by online evaluation of both detection accuracy and computational efficiency. In addition, we propose a distribution-drift dataset of real time series, called TSB- drift, to isolate scenarios where streaming updates are theoretically justified. Our results show that, contrary to common assumptions, static TSAD methods significantly outperform streaming approaches in most streaming settings. Such finding highlights a critical gap between the design of existing streaming methods and the requirements of modern TSAD, and calls for a rethinking of how streaming capabilities should be integrated into TSAD.
Figures & tables
Figure 1. Static , Online , and Streaming Time Series Anomaly Detection
Figure 2. Illustration of TSB- drift Construction
General
Drift
Dataset
Category (Field)
# TS
μ(D)
# TS
ratio
Type
Genesis
Sensor (Robotics)
1
18
0/1
0%
∅
MITDB
Medical
13
13
0/13
0%
∅
PSM
Facility
1
25
0/1
0%
∅
SVDB
Medical
31
2
0/31
0%
∅
MSL
Sensor (Aerospace)
16
16
0/16
0%
∅
Table 1. Datasets in StrAD . Drift : # of TS with Concept Drift, % of Dimensions with Drifts (ratio), and Type
Table 2. Characteristic Patterns Throughout TSB- drift (Training Batch B0 is Highlighted in Green)
Acronym
Method
Type
Complexity
Distance-based
LOF
LOF ( Breunig et al., 2000 )
Proximity
O(∣T∣×D)∗
KNN
k -NN ( Hawkins, 1980 )
Proximity
O(∣T∣×D)∗
KMAD
k -Means ( Hawkins, 1980 )
Clustering
O(D)
CBLOF
CBLOF ( He et al., 2003 )
Clustering
O(D)
Density-based
Table 3. Static/Online TSAD Methods in StrAD ( ∗ : For Online settings, ∣T∣ is the Size of the Initial Batch B0 )
Acronym
Method
UM (Sec 2.2.2 )
MM (Sec 2.2.3 )
Complexity
Numerical
LODA
LODA ( Pevný, 2016 )
Projections
Tumbling Window
O(D)
xS
xStream ( Manzoor et al., 2018 )
Projections
Tumbling Window
O(D)
RSH
RSHash ( Sathe and Aggarwal, 2018 )
Partitioning
Sliding Window (Point)
O(1)∗
HST
HSTree ( Tan et al., 2011 )
Partitioning
Tumbling Window
O(1)∗
SDOs
SDOstream ( Hartl et al., 2020 )
Partitioning
Soft Forgetting (Aging)
O(D)
Table 4. Streaming TSAD in StrAD ( ∗ : Independent to D but Dependent to model hyper-parameters)
Figure 3. Static vs. Online Accuracy Evaluation on TSB-AD-M: (1) Overall, (2) by Data Category.
Figure 4. (1) AUC-PR gain from Static to Online Settings; (2) Online vs Static on TSB-AD-M for (2a) LOF, (2b) AE, (2c) IF.
Figure 5. Online vs. Streaming Accuracy on TSB-AD-M: (1) Overall, (2) by Data Category, (3) Critical Difference Diagrams ( α=0.05 ) on (3a) TSB-AD-M and (3b) Point Anomalies Time Series.
Figure 6. (1) TSB- drift vs TSB- ¬ drift accuracy; Mean AUC-PR gain of TSB- drift on TSB- ¬ drift for (1a) Online models and (1b) Streaming models; (2) Mean performance on drift-types
Figure 7. Scalability Evaluation on TSB-AD-M: (1) Accuracy vs Throughput; (2) Throughput vs Number of Dimensions; (3) Inference time standard deviation on TSB- ¬ drift and TSB- drift for (3a) Online and (3b) Streaming models (Red line highlights median on TSB- ¬ drift and dotted line on TSB- drift )
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Hyperparameter
Search Range
RRCF
num_trees
{2, 4, 8, 16, 20, 25}
shingle_size
{2, 4, 8}
tree_size
{256, 512}
RSHash
sampling_points
{500, 1000}
decay
{0.01, 0.015, 0.02}
num_components
{50, 100, 200, 300}
Appendix
Table 5. Hyperparameter Search Grid for the Evaluated Streaming Models
School of Computer Science (National Pilot Software Engineering School), Beijing University of Posts and Telecommunications, Beijing 100876, China · China Telecom Research Institute Beijing, China · Department of Computer Science, Missouri University of Science and Technology, Rolla, MO 65409 USA