Detect, Explain, Interpret: An End-to-End Benchmark for Time Series Anomaly Detection, Explainability and Interpretability
Organizations: Inria, ENS, CNRS, PSL Paris, France · Scality Paris, France · Inria, ENS, CNRS, PSL, EDF Paris, France
Abstract
Time Series Anomaly Detection has received increasing attention, driven by the growing availability of complex time series data. This surge has led to the development of numerous detection methods, as well as a variety of benchmarks aimed at thoroughly evaluating their performance. However, most existing detectors remain largely agnostic to domain context, overlooking explainability and interpretability. One of the main reasons for this gap is that current benchmarks primarily focus on detection accuracy, and only few of them evaluate spatial explainability. Moreover, no benchmark currently provides sufficiently rich semantic annotations to support the generation of human-understandable interpretations of anomalies. To address these limitations, we introduce SHAD (Scality High-dimensional Anomaly Detection benchmark), a fully annotated benchmark composed of 215 multivariate, high-dimensional time series collected from real-world distributed cloud storage systems operated by Scality. The proposed dataset includes rich contextual information, covering three families of anomalies with varying degrees of severity. As further contribution, we provide a foundation for future work by evaluating baseline methods for Detection, Explainability, and Interpretability, covering all stages of a TSAD pipeline. For Detection, we benchmark a wide range of existing anomaly detectors, testing their effectiveness on the proposed real-world dataset. Then, we consider explainability by evaluating whether measuring the contribution of each dimension in the generated anomaly score can provide accurate anomaly attributions. Finally, for interpretability, we investigate the effectiveness of frozen LLM baselines in localizing and interpreting anomalies.
Figures & tables
| Benchmark | Dataset | Evaluation | ||||
| Multi. | Ann. | Desc. | AD | Expl. | Interp. | |
| NAB [ 48 ] | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} |
| Exathlon [ 36 ] | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} |
| TODS [ 44 ] | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} |
| TimeEval [ 74 ] | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} |
| TSB-UAD [ 66 ] | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{0.0664,0.7266,0.0664}\checkmark} | {\color[rgb]{1,0,0}\times} | {\color[rgb]{1,0,0}\times} |
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
| Name | Category | Dimension | Avg Dim | # of time series | Avg Length | Domain |
| UCR [ 94 ] | P&S | I | 1 | 250 | 67818.7 | Multiple |
| Yahoo [ 47 ] | P&S | I | 1 | 367 | 1560.2 | Web Service |
| IOPS [ 35 ] | S | I | 1 | 58 | 72792.3 | IT/Web Services |
| MGAB [ 84 ] | S | I | 1 | 10 | 97777.8 | Synthetic |
| WSD [ 106 ] | S | I | 1 | 210 | 17444.5 | IT/Web Services |
| SED [ 14 ] | S | I | 1 | 6 | 23332.3 | Sensor |
| Metric | Unit | Description |
| cpu_usage_ratio | ratio [0,1] | Fraction of CPU time not idle, averaged across all servers. |
| cpu_load_ratio | ratio [ 0] | CPU load average normalized by CPU count across all servers. Can exceed 1.0 when the run queue exceeds CPU count. |
| memory_usage_ratio | ratio [0,1] | Fraction of total used memory, averaged across all servers. |
| network_in_bytes | bytes/s | Total incoming network throughput across all interfaces on all servers. |
| network_out_bytes | bytes/s | Total outgoing network throughput across all interfaces on all servers. |
| objects_count | count | Number of objects stored on the RING, including replicas. |
| Metric | Unit | Description |
| s3_errors_5xx | errors/s | Rate of HTTP 5xx server errors across all S3 services. |
| s3_request_get_operations | ops/s | GET request rate. |
| s3_request_post_operations | ops/s | POST request rate. Near-constant (no multipart uploads in the nominal workload). |
| s3_request_put_operations | ops/s | PUT request rate. |
| s3_request_delete_operations | ops/s | DELETE request rate. |
| s3_request_head_operations | ops/s | HEAD (metadata lookup) request rate. |
| Metric | Unit | Description |
| cpu_usage_ratio | ratio [0,1] | Fraction of CPU time not idle for this server. |
| cpu_load_ratio | ratio [ 0] | CPU load average normalized by CPU count. Can exceed 1.0. |
| memory_usage_ratio | ratio [0,1] | Fraction of total used memory. |
| network_in_bytes | bytes/s | All incoming network throughput (S3, RING inter-node, monitoring). |
| network_out_bytes | bytes/s | All outgoing network throughput. |
| operations_sent | ops/s | Rate of internal RING operations sent to other servers. |
| Metric | Unit | Description |
| disk_read_bytes | bytes/s | Disk read throughput. |
| disk_write_bytes | bytes/s | Disk write throughput. |
| disk_io_time_seconds | s/s [0,1] | Fraction of time spent on I/O. Approaching 1.0 implies saturation. |
| disk_usage_ratio | ratio [0,1] | Disk space utilization. |
| ID | Details | Notes |
|---|---|---|
| 1773251393 | 18h, 16 workers, 15k objects 100KiB, GET 95% / PUT 0.8% / STAT 4% / DEL 0.4% | |
| 1773251399 | — ibid — | |
| 1773343560 | — ibid — | |
| 1773343565 | — ibid — | |
| 1773343568 | — ibid — | |
| 1773521107 | — ibid — |
| ID | Details | Notes |
|---|---|---|
| Single Disk Failures | ||
| 1773343796 | s1/g2disk01, onset 3h30m, dur 1h30m | |
| 1773343802 | s3/g2disk02, onset 6h, dur 35m | |
| 1773343833 | s3/g1disk02, onset 7h15m, dur 2h30m | |
| 1773343804 | s1/g2disk01, onset 10h, dur 10m | |
| 1773343837 | s1/g2disk02, onset 8h45m, dur 3h15m | |
| ID | Details | Notes |
|---|---|---|
| Single Server Failures | ||
| 1773946658 | store 1, onset 6h, dur 9h45m | unplanned crash |
| 1774026192 | store 1, onset 6h, dur 20m | rerun of 1773946658 |
| 1774026217 | store 2, onset 10h, dur 15m | |
| 1774132105 | store 3, onset 3h, dur 5m | |
| 1774132113 | store 2, onset 14h, dur 1m | |
| ID | Details | Notes |
|---|---|---|
| Single-Server Software Node Failures | ||
| 1774656169 | 1 snode on s1, onset 6h, dur 1h | server metric loss |
| 1774656174 | 1 snode on s2, onset 9h, dur 1h30m | server metric loss |
| 1774656180 | 2 snodes on s1, onset 5h, dur 45m | server metric loss |
| 1774656183 | 2 snodes on s3, onset 8h, dur 1h | server metric loss |
| 1774764369 | 3 snodes on s2, onset 2h, dur 30m | server metric loss |
| Type ( Subtype ) | Onset | Duration | Parameters |
| Nominal | – | 18 h | normal behaviour |
| Single disk | 3.5–14.5 h | 10–240 min | 3 servers, 4 disk types |
| Simult. disk | 3–15.5 h | 30–215 min | 2-4 disks, 1-3 servers |
| Async. disk ( C/R/I ) | 1.5–15 h | 30–180 min | 2-4 events per exp. |
| Single server | 0.5–16 h | 1–25 min | 3 servers |
| Async. server | 1–16 h | 3–30 min | 2-3 events, same-server |
| Name | Legend | Supervision | Category | Description |
| AutoEncoder [ 73 ] | AE | S | Prediction-based | Projects data to the lower-dimensional latent space and then reconstruct it through the encoding-decoding phase, where anomalies are typically characterized by evident reconstruction deviations. |
| OmniAnomaly [ 79 ] | OA | S | Prediction-based | Stochastic recurrent neural network, which captures the normal patterns of time series by learning their robust representations with key techniques such as stochastic variable connection and planar normalizing flow, reconstructs input data by the representations, and use the reconstruction probabilities to determine anomalies. |
| USAD [ 5 ] | USAD | S | Prediction-based | based on adversely trained autoencoders, and the anomaly score is the combination of discriminator and reconstruction loss. |
| CNN [ 60 ] | CNN | S | Prediction-based | Employs Convolutional Neural Network (CNN) to predict the next time stamp on the defined horizon and then compare the difference with the original value. |
| LSTM [ 56 ] | LSTM | S | Prediction-based | Utilizes Long Short-Term Memory (LSTM) networks to model the relationship between current and preceding time series data, detecting anomalies through discrepancies between predicted and actual values. |
| OCSVM [ 75 ] | SVM | S | Density-based | Fits the dataset to find the normal data’s boundary by maximizing the margin between the origin and the normal samples. |
| Anomaly Type | Avg. Anomalous Dim. | CNN | KmeansAD | OmniAnomaly | Rand. |
| Single Disk | 4 | 0.00 | 0.00 | 0.00 | 0.02 |
| Single Server | 47 | 0.27 | 0.16 | 0.17 | 0.26 |
| Single Node | 47 | 0.24 | 0.18 | 0.29 | 0.26 |
| Simultaneous Disk | 12 | 0.05 | 0.11 | 0.00 | 0.08 |
| Simultaneous Node | 100.7 | 0.57 | 0.46 | 0.60 | 0.57 |
| Asynchronous Disk | 11.27 | 0.02 | 0.07 | 0.00 | 0.05 |
| Task | Model | Performance |
| Anomaly Criticality | Claude | 0.196400 |
| Anomaly Criticality | Gemini | 0.158063 |
| Anomaly Criticality | Llama | 0.149613 |
| Anomaly Criticality | Mistral | 0.147306 |
| Anomaly Criticality | Grok | 0.134204 |
| Dimension Localization | Grok | 0.047685 |
| Supervision | Task | Mean Gain |
| Level 1 | Anomaly Criticality | 0.0128 |
| Level 2 | Anomaly Criticality | -0.0078 |
| Level 1 | Dimension Localization | 0.0719 |
| Level 2 | Dimension Localization | N.A. (trivial) |
| Level 1 | Anomaly Presence | N.A. (trivial) |
| Level 2 | Anomaly Presence | N.A. (trivial) |