Semi-supervised anomaly detection, which aims to improve the anomaly detection performance by using a small amount of labeled anomaly data in addition to unlabeled data, has attracted attention. Existing semi-supervised approaches assume that most unlabeled data are normal, and train anomaly detectors by minimizing the anomaly scores for the unlabeled data while maximizing those for the labeled anomaly data. However, in practice, the unlabeled data are often contaminated with anomalies. This weakens the effect of maximizing the anomaly scores for anomalies, and prevents us from improving the detection performance. To solve this, we propose the deep positive-unlabeled anomaly detection framework, which integrates positive-unlabeled learning with deep anomaly detection models such as autoencoders and deep support vector data descriptions. Our approach enables the approximation of anomaly scores for normal data using the unlabeled data and the labeled anomaly data. Therefore, without labeled normal data, our approach can train anomaly detectors by minimizing the anomaly scores for normal data while maximizing those for the labeled anomaly data. We also provide a theoretical analysis establishing a generalization error bound for the proposed objective, guaranteeing that the empirical minimizer converges asymptotically to the ideal minimizer. Our approach achieves better detection performance than existing approaches on various datasets.
Figures & tables
Figure 1: The comparison of the PU learning, the unsupervised anomaly detector (DAE), the semi-supervised anomaly detector (ABC), and the proposed method on the toy dataset. The unlabeled data in this dataset include both normal and anomaly data points. The yellow and blue in the contour maps represent abnormality and normality, respectively. The purple stars represent the examples of unseen anomalies, which are new types of anomalies unseen during training.
Term
Definition
Split
Unlabeled dataset ( U )
Contains normal data and seen anomalies without labels
Training
Anomaly dataset ( A )
Labeled seen anomaly samples
Training
Normal data
Non-anomalous samples
Training ( U ) + Test
Seen anomalies
Anomalies of the same types as those in A
Training ( A , U ) + Test
Unseen anomalies
Anomalies of types not represented in A
Test only
Table 1: Key terminology and dataset correspondence.
MNIST
FashionMNIST
SVHN
CIFAR10
IF
0.885 ± 0.062
0.916 ± 0.077
0.501 ± 0.014
0.610 ± 0.095
AE
0.912 ± 0.042
0.841 ± 0.102
0.562 ± 0.040
0.535 ± 0.120
DeepSVDD
0.937 ± 0.045
0.921 ± 0.088
0.582 ± 0.035
0.709 ± 0.054
LOE
0.945 ± 0.033
0.916 ± 0.094
0.624 ± 0.054
0.718 ± 0.062
ABC
0.916 ± 0.042
0.841 ± 0.104
0.562 ± 0.041
0.535 ± 0.119
DeepSAD
0.942 ± 0.041
0.928 ± 0.089
0.652 ± 0.034
0.726 ± 0.051
Table 2: Comparison of anomaly detection performance on MNIST, FashionMNIST, SVHN, and CIFAR10.
CIFAR100
Path
OCT
Tissue
IF
0.604 ± 0.004
0.809 ± 0.118
0.714 ± 0.004
0.472 ± 0.187
AE
0.589 ± 0.010
0.605 ± 0.240
0.860 ± 0.005
0.468 ± 0.178
DeepSVDD
0.587 ± 0.026
0.759 ± 0.148
0.726 ± 0.052
0.661 ± 0.055
LOE
0.576 ± 0.035
0.721 ± 0.160
0.783 ± 0.030
0.635 ± 0.086
ABC
0.590 ± 0.010
0.604 ± 0.241
0.857 ± 0.001
0.472 ± 0.177
DeepSAD
0.594 ± 0.012
0.763 ± 0.187
0.823 ± 0.038
0.683 ± 0.053
Table 3: Comparison of anomaly detection performance on CIFAR100, Path, OCT and Tissue.
Figure 2: Relationship between the anomaly detection performance and the hyperparameter α of the PUSVDD and the SOEL on each dataset. We used class 0 as unseen anomaly, class 1 as normal, and the remaining classes as seen anomalies. The semi-transparent area represents standard deviations.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: The example of the dataset in the case of MNIST.
Figure 4: Comparison of anomaly detection performance between the PUSVDD and the SOEL with various numbers of unlabeled anomalies on each dataset. We used class 0 as unseen anomaly, class 1 as normal, and the remaining classes as seen anomalies. The semi-transparent area represents standard deviations.
MNIST
FashionMNIST
SVHN
CIFAR10
IF
0.814 ± 0.096
0.915 ± 0.053
0.510 ± 0.016
0.585 ± 0.106
AE
0.843 ± 0.073
0.840 ± 0.092
0.563 ± 0.039
0.573 ± 0.131
DeepSVDD
0.925 ± 0.056
0.942 ± 0.040
0.600 ± 0.033
0.694 ± 0.069
LOE
0.942 ± 0.038
0.940 ± 0.042
0.645 ± 0.055
0.713 ± 0.077
ABC
0.852 ± 0.072
0.841 ± 0.092
0.564 ± 0.039
0.572 ± 0.131
DeepSAD
0.930 ± 0.052
0.956 ± 0.032
0.674 ± 0.031
0.716 ± 0.074
Appendix
Table 4: Seen anomaly detection performance on MNIST, FashionMNIST, SVHN, and CIFAR10.
MNIST
FashionMNIST
SVHN
CIFAR10
IF
0.955 ± 0.031
0.917 ± 0.111
0.492 ± 0.018
0.635 ± 0.095
AE
0.981 ± 0.013
0.842 ± 0.130
0.560 ± 0.045
0.497 ± 0.111
DeepSVDD
0.950 ± 0.044
0.900 ± 0.143
0.564 ± 0.043
0.724 ± 0.087
LOE
0.949 ± 0.034
0.892 ± 0.152
0.602 ± 0.058
0.724 ± 0.099
ABC
0.980 ± 0.014
0.840 ± 0.133
0.559 ± 0.046
0.497 ± 0.109
DeepSAD
0.954 ± 0.040
0.900 ± 0.153
0.631 ± 0.044
0.735 ± 0.078
Appendix
Table 5: Unseen anomaly detection performance on MNIST, FashionMNIST, SVHN, and CIFAR10.
CIFAR100
Path
OCT
Tissue
IF
0.577 ± 0.005
0.657 ± 0.211
0.679 ± 0.005
0.443 ± 0.189
AE
0.514 ± 0.013
0.562 ± 0.253
0.808 ± 0.005
0.443 ± 0.188
DeepSVDD
0.623 ± 0.036
0.769 ± 0.139
0.702 ± 0.043
0.692 ± 0.047
LOE
0.624 ± 0.035
0.746 ± 0.167
0.750 ± 0.031
0.657 ± 0.084
ABC
0.516 ± 0.014
0.564 ± 0.252
0.805 ± 0.002
0.448 ± 0.187
DeepSAD
0.628 ± 0.042
0.772 ± 0.165
0.798 ± 0.032
0.715 ± 0.040
Appendix
Table 6: Seen anomaly detection performance on CIFAR100, Path, OCT and Tissue.
CIFAR100
Path
OCT
Tissue
IF
0.632 ± 0.005
0.961 ± 0.055
0.785 ± 0.004
0.500 ± 0.187
AE
0.665 ± 0.009
0.647 ± 0.239
0.965 ± 0.005
0.493 ± 0.170
DeepSVDD
0.552 ± 0.033
0.749 ± 0.200
0.774 ± 0.072
0.629 ± 0.069
LOE
0.528 ± 0.047
0.696 ± 0.205
0.851 ± 0.030
0.613 ± 0.093
ABC
0.665 ± 0.009
0.644 ± 0.241
0.961 ± 0.002
0.496 ± 0.168
DeepSAD
0.561 ± 0.024
0.755 ± 0.242
0.874 ± 0.049
0.651 ± 0.075
Appendix
Table 7: Unseen anomaly detection performance on CIFAR100, Path, OCT and Tissue.
PUSVDD
SOEL
WSAD-DT
all
0.793 ± 0.016
0.669 ± 0.021
0.641 ± 0.029
seen
0.799 ± 0.013
0.688 ± 0.021
0.642 ± 0.030
unseen
0.787 ± 0.020
0.649 ± 0.027
0.640 ± 0.028
Appendix
Table 8: Comparison of anomaly detection performance on MVTec AD.
Figure 5: Relationship between the anomaly detection performance and the hyperparameter α of the PUSVDD and the SOEL on MVTec AD. The semi-transparent area represents standard deviations.
Dataset
α^
MNIST
0.048
FashionMNIST
0.056
SVHN
0.350
CIFAR-10
0.325
Appendix
Table 9: CPE estimates of α compared with the true value ( α=0.1 ).
Class
SOEL
PUSVDD
1
0.995 ± 0.001
0.999 ± 0.000
2
0.927 ± 0.047
0.988 ± 0.003
3
0.986 ± 0.009
0.996 ± 0.001
4
0.982 ± 0.013
0.997 ± 0.001
5
0.935 ± 0.024
0.989 ± 0.009
6
0.964 ± 0.005
0.967 ± 0.013
Appendix
Table 10: Per-class AUROC on MNIST. Class 0 (digit 0) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.983 ± 0.003
0.993 ± 0.002
2
0.930 ± 0.019
0.952 ± 0.010
3
0.911 ± 0.007
0.930 ± 0.015
4
0.920 ± 0.008
0.945 ± 0.014
5
0.985 ± 0.004
0.993 ± 0.002
6
0.750 ± 0.026
0.763 ± 0.019
Appendix
Table 11: Per-class AUROC on FashionMNIST. Class 0 (T-shirt/top) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.783 ± 0.040
0.850 ± 0.048
2
0.760 ± 0.014
0.801 ± 0.027
3
0.698 ± 0.032
0.682 ± 0.029
4
0.769 ± 0.033
0.788 ± 0.025
5
0.765 ± 0.012
0.799 ± 0.022
6
0.705 ± 0.013
0.690 ± 0.018
Appendix
Table 12: Per-class AUROC on SVHN. Class 0 (digit 0) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.793 ± 0.066
0.824 ± 0.030
2
0.712 ± 0.012
0.724 ± 0.010
3
0.724 ± 0.054
0.795 ± 0.026
4
0.731 ± 0.096
0.792 ± 0.015
5
0.790 ± 0.029
0.843 ± 0.025
6
0.856 ± 0.016
0.852 ± 0.023
Appendix
Table 13: Per-class AUROC on CIFAR10. Class 0 (airplane) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.965 ± 0.051
0.996 ± 0.005
2
0.526 ± 0.126
0.554 ± 0.123
3
0.950 ± 0.035
0.957 ± 0.007
4
0.737 ± 0.031
0.733 ± 0.068
5
0.741 ± 0.035
0.751 ± 0.096
6
0.812 ± 0.082
0.859 ± 0.086
Appendix
Table 14: Per-class AUROC on Path. Class 0 is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.623 ± 0.017
0.622 ± 0.010
2
0.788 ± 0.052
0.843 ± 0.029
3
0.739 ± 0.029
0.806 ± 0.030
4
0.693 ± 0.016
0.749 ± 0.018
5
0.686 ± 0.035
0.722 ± 0.040
6
0.739 ± 0.026
0.716 ± 0.011
Appendix
Table 15: Per-class AUROC on Tissue. Class 0 is the unseen anomaly.
Figure 6: The anomaly detection performance of the PUAE and the SOEL as the fraction β of unseen anomalies among the unlabeled anomalies increases, with the total contamination ratio fixed at α=0.1 . Both methods use the same DAE as the base anomaly detector, so that only the training objective differs. The three panels show the AUROC on the whole test set, on the seen anomalies, and on the unseen anomalies, respectively. The semi-transparent area represents standard deviations.
Visual anomaly detection requires adaptive representations and reliable decision boundaries, particularly when anomalous training samples are scarce and class distributions are highly imbalanced. Classical kernel-based methods yield principled geometric decision regions but typically operate on fixed features, while deep detectors learn task-specific representations but often fail to provide an explicit margin-aware kernel boundary. In this study, we propose DLM-SVDD, a deep large-margin novelty-detection framework that jointly learns convolutional features and an explicit kernel-based decision boundary. By drawing on the large-margin ℓp-Support Vector Data Description (ℓp-SVDD) approach, the proposed method performs explicit margin maximization and nonlinear slack penalization while adapting the representation to the target task. To train the proposed model, we present an optimization scheme that alternates between a Frank--Wolfe--based update of the convex dual boundary and a CNN update step operating on a smooth margin-violation loss induced by the recovered boundary. To improve scalability, we analyze the efficiency--accuracy trade-offs for different kernel approximation strategies, deriving practical propositions for large-scale anomaly detection. Extensive experiments on multiple standard benchmarks show consistent performance improvements over the baseline and strong overall performance compared with state-of-the-art methods while illustrating that the proposed joint representation--boundary learning scheme remains effective under severe imbalanced class distributions.
Anomaly detection is a critical and evolving field in Machine Learning, with applications targeting different domains such as cybersecurity, finance, healthcare, manufacturing and IoT (Internet of Things) systems. Traditionally, anomaly detection algorithms have been designed using both supervised and unsupervised learning paradigms. The fundamental challenge in real-world anomaly detection scenarios is related to the inherent class imbalance (anomalies are typically rare) and, for supervised methods, to the scarcity of labelled anomalous data. Indeed, labelling is both expensive and time-consuming. Conversely unsupervised methods do not require labelling, but may suffer from high false positive rates when deployed in safety-critical applications. In this work we introduce a novel unsupervised algorithm for anomaly detection in time series based on the Haar discrete wavelet and a suitably designed t-test. We establish the theoretical foundation of the proposed t-test and, through extensive experimentation across 343 datasets, demonstrate that our algorithm outperforms state-of-the-art unsupervised and self-supervised benchmarks.
Emanuele Mele, Massimo Cafaro, Angelo Coluccia +1
University of Salento, Via per Monteroni, Lecce, 73100, , Italy
This paper considers a practical few-shot anomaly detection (FSAD) setting, termed discriminative FSAD, where a limited number of both normal and anomalous examples are available as references during inference. Existing FSAD methods rely on normal-only references through normality matching, ignoring the discriminative clues in anomalous references, while directly fitting both references can overfit to the seen anomalies. We introduce IDEAL, an intrinsic deviation learning framework that leverages both reference types to learn intrinsic deviation patterns characterizing generalizable abnormality as deviations from normality. IDEAL decomposes the learning process into two novel components: 1) a Normal Variation Eraser to suppress nuisance normal variations that may lead to noisy deviations from normality, thereby highlighting anomaly-relevant deviation representations; 2) an Intrinsic Deviation Encoder to decompose these denoised deviation representations into intrinsic deviation vectors capturing the most discriminative orthogonal deviation directions. At inference, IDEAL scores query-to-normal deviations preserved after projection onto the learned intrinsic deviation vectors, enabling generalization for both seen and unseen anomalies. Extensive experiments on eight real-world datasets show that IDEAL generalizes effectively to unseen anomalies and consistently outperforms existing state-of-the-art FSAD methods. Code and data are available at https://github.com/mala-lab/IDEAL.
Huan Wang, Jun Shen, Jun Yan +1
Singapore Management University, Singapore · University of Wollongong, Australia