A Multi-Source Ultrasound Benchmark Revealing the Limits of Contemporary Self-Supervised Anomaly Detection Methods
Organizations: University of Augsburg
Abstract
Self-supervised anomaly detection is a promising paradigm for medical ultrasound, as normal images are often easier to obtain than exhaustive annotations of all possible pathologies. However, most existing evaluations are limited to a single anatomy or task, making it unclear whether models learn a robust notion of normal ultrasound appearance or only a source-specific representation. We introduce the SADUSI benchmark, a multi-source ultrasound dataset designed to train and evaluate anomaly detection methods across a broad range of anatomical regions, views, and acquisition protocols. The goal of SADUSI is to provide a diverse normal ultrasound distribution and a benchmark for visible structural anomalies that can be assessed from single images. We evaluate representative self-supervised anomaly detection methods and find that current approaches struggle in this setting. In particular, reconstruction-based diffusion methods such as AnoDDPM and DeCo-Diff achieve pixel-level AUROC values of 0.56-0.72 and maximum F1 scores of 0.10-0.26, indicating limited separation of pathology from normal image regions. Feature-based PatchCore variants perform better, reaching pixel-level AUROC values of 0.76-0.83, but remain limited with maximum F1 scores of 0.14-0.40. These findings suggest that broad multi-source ultrasound anomaly detection remains an open challenge and that SADUSI can serve as a resource for developing methods that generalize beyond anatomy-specific settings.
Figures & tables
| Source | Anatomy | Type | / / | Use |
|---|---|---|---|---|
| AbdomenUS [ 19 ] | Abdominal | Images | 616 / 616 / 0 | Train |
| AUL [ 20 ] | Liver | Images | 100 / 100 / 434 | Train / Test |
| Breast-Lesions-USG [ 21 , 22 ] | Breast | Images | 4 / 4 / 0 | Train |
| BUS-UCLM [ 23 , 10 ] | Breast | Images | 405 / 405 / 0 | Train |
| BUSI [ 9 , 24 ] | Breast | Images | 131 / 131 / 512 | Train / Test |
| COVID-19 [ 25 ] | Lung | Images | 62 / 62 / 0 | Train |
| Eval. set | Training | Method | AUROC | AP | AUPRO | F1 | Prec. | Rec. |
|---|---|---|---|---|---|---|---|---|
| TN3K | Filtered | AnoDDPM | 0.59 | 0.14 | 0.17 | 0.23 | 0.14 | 0.82 |
| TN3K | Filtered | DeCo-Diff | 0.66 | 0.16 | 0.20 | 0.26 | 0.16 | 0.64 |
| TN3K | Filtered | PatchCore-WRN | 0.76 | 0.24 | 0.40 | 0.33 | 0.22 | 0.62 |
| TN3K | Filtered | PatchCore-DINOv3 | 0.79 | 0.26 | 0.29 | 0.36 | 0.26 | 0.61 |
| TN3K | Filtered | PatchCore-DINOv3 (A) | 0.76 | 0.24 | 0.37 | 0.33 | 0.22 | 0.62 |
| TN3K | Filtered | EfficientAD | 0.50 | 0.11 | 0.07 | 0.20 | 0.11 | 0.94 |