Semi-supervised anomaly detection, which aims to improve the anomaly detection performance by using a small amount of labeled anomaly data in addition to unlabeled data, has attracted attention. Existing semi-supervised approaches assume that most unlabeled data are normal, and train anomaly detectors by minimizing the anomaly scores for the unlabeled data while maximizing those for the labeled anomaly data. However, in practice, the unlabeled data are often contaminated with anomalies. This weakens the effect of maximizing the anomaly scores for anomalies, and prevents us from improving the detection performance. To solve this, we propose the deep positive-unlabeled anomaly detection framework, which integrates positive-unlabeled learning with deep anomaly detection models such as autoencoders and deep support vector data descriptions. Our approach enables the approximation of anomaly scores for normal data using the unlabeled data and the labeled anomaly data. Therefore, without labeled normal data, our approach can train anomaly detectors by minimizing the anomaly scores for normal data while maximizing those for the labeled anomaly data. We also provide a theoretical analysis establishing a generalization error bound for the proposed objective, guaranteeing that the empirical minimizer converges asymptotically to the ideal minimizer. Our approach achieves better detection performance than existing approaches on various datasets.
Figures & tables
Figure 1: The comparison of the PU learning, the unsupervised anomaly detector (DAE), the semi-supervised anomaly detector (ABC), and the proposed method on the toy dataset. The unlabeled data in this dataset include both normal and anomaly data points. The yellow and blue in the contour maps represent abnormality and normality, respectively. The purple stars represent the examples of unseen anomalies, which are new types of anomalies unseen during training.
Term
Definition
Split
Unlabeled dataset ( U )
Contains normal data and seen anomalies without labels
Training
Anomaly dataset ( A )
Labeled seen anomaly samples
Training
Normal data
Non-anomalous samples
Training ( U ) + Test
Seen anomalies
Anomalies of the same types as those in A
Training ( A , U ) + Test
Unseen anomalies
Anomalies of types not represented in A
Test only
Table 1: Key terminology and dataset correspondence.
MNIST
FashionMNIST
SVHN
CIFAR10
IF
0.885 ± 0.062
0.916 ± 0.077
0.501 ± 0.014
0.610 ± 0.095
AE
0.912 ± 0.042
0.841 ± 0.102
0.562 ± 0.040
0.535 ± 0.120
DeepSVDD
0.937 ± 0.045
0.921 ± 0.088
0.582 ± 0.035
0.709 ± 0.054
LOE
0.945 ± 0.033
0.916 ± 0.094
0.624 ± 0.054
0.718 ± 0.062
ABC
0.916 ± 0.042
0.841 ± 0.104
0.562 ± 0.041
0.535 ± 0.119
DeepSAD
0.942 ± 0.041
0.928 ± 0.089
0.652 ± 0.034
0.726 ± 0.051
Table 2: Comparison of anomaly detection performance on MNIST, FashionMNIST, SVHN, and CIFAR10.
CIFAR100
Path
OCT
Tissue
IF
0.604 ± 0.004
0.809 ± 0.118
0.714 ± 0.004
0.472 ± 0.187
AE
0.589 ± 0.010
0.605 ± 0.240
0.860 ± 0.005
0.468 ± 0.178
DeepSVDD
0.587 ± 0.026
0.759 ± 0.148
0.726 ± 0.052
0.661 ± 0.055
LOE
0.576 ± 0.035
0.721 ± 0.160
0.783 ± 0.030
0.635 ± 0.086
ABC
0.590 ± 0.010
0.604 ± 0.241
0.857 ± 0.001
0.472 ± 0.177
DeepSAD
0.594 ± 0.012
0.763 ± 0.187
0.823 ± 0.038
0.683 ± 0.053
Table 3: Comparison of anomaly detection performance on CIFAR100, Path, OCT and Tissue.
Figure 2: Relationship between the anomaly detection performance and the hyperparameter α of the PUSVDD and the SOEL on each dataset. We used class 0 as unseen anomaly, class 1 as normal, and the remaining classes as seen anomalies. The semi-transparent area represents standard deviations.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: The example of the dataset in the case of MNIST.
Figure 4: Comparison of anomaly detection performance between the PUSVDD and the SOEL with various numbers of unlabeled anomalies on each dataset. We used class 0 as unseen anomaly, class 1 as normal, and the remaining classes as seen anomalies. The semi-transparent area represents standard deviations.
MNIST
FashionMNIST
SVHN
CIFAR10
IF
0.814 ± 0.096
0.915 ± 0.053
0.510 ± 0.016
0.585 ± 0.106
AE
0.843 ± 0.073
0.840 ± 0.092
0.563 ± 0.039
0.573 ± 0.131
DeepSVDD
0.925 ± 0.056
0.942 ± 0.040
0.600 ± 0.033
0.694 ± 0.069
LOE
0.942 ± 0.038
0.940 ± 0.042
0.645 ± 0.055
0.713 ± 0.077
ABC
0.852 ± 0.072
0.841 ± 0.092
0.564 ± 0.039
0.572 ± 0.131
DeepSAD
0.930 ± 0.052
0.956 ± 0.032
0.674 ± 0.031
0.716 ± 0.074
Appendix
Table 4: Seen anomaly detection performance on MNIST, FashionMNIST, SVHN, and CIFAR10.
MNIST
FashionMNIST
SVHN
CIFAR10
IF
0.955 ± 0.031
0.917 ± 0.111
0.492 ± 0.018
0.635 ± 0.095
AE
0.981 ± 0.013
0.842 ± 0.130
0.560 ± 0.045
0.497 ± 0.111
DeepSVDD
0.950 ± 0.044
0.900 ± 0.143
0.564 ± 0.043
0.724 ± 0.087
LOE
0.949 ± 0.034
0.892 ± 0.152
0.602 ± 0.058
0.724 ± 0.099
ABC
0.980 ± 0.014
0.840 ± 0.133
0.559 ± 0.046
0.497 ± 0.109
DeepSAD
0.954 ± 0.040
0.900 ± 0.153
0.631 ± 0.044
0.735 ± 0.078
Appendix
Table 5: Unseen anomaly detection performance on MNIST, FashionMNIST, SVHN, and CIFAR10.
CIFAR100
Path
OCT
Tissue
IF
0.577 ± 0.005
0.657 ± 0.211
0.679 ± 0.005
0.443 ± 0.189
AE
0.514 ± 0.013
0.562 ± 0.253
0.808 ± 0.005
0.443 ± 0.188
DeepSVDD
0.623 ± 0.036
0.769 ± 0.139
0.702 ± 0.043
0.692 ± 0.047
LOE
0.624 ± 0.035
0.746 ± 0.167
0.750 ± 0.031
0.657 ± 0.084
ABC
0.516 ± 0.014
0.564 ± 0.252
0.805 ± 0.002
0.448 ± 0.187
DeepSAD
0.628 ± 0.042
0.772 ± 0.165
0.798 ± 0.032
0.715 ± 0.040
Appendix
Table 6: Seen anomaly detection performance on CIFAR100, Path, OCT and Tissue.
CIFAR100
Path
OCT
Tissue
IF
0.632 ± 0.005
0.961 ± 0.055
0.785 ± 0.004
0.500 ± 0.187
AE
0.665 ± 0.009
0.647 ± 0.239
0.965 ± 0.005
0.493 ± 0.170
DeepSVDD
0.552 ± 0.033
0.749 ± 0.200
0.774 ± 0.072
0.629 ± 0.069
LOE
0.528 ± 0.047
0.696 ± 0.205
0.851 ± 0.030
0.613 ± 0.093
ABC
0.665 ± 0.009
0.644 ± 0.241
0.961 ± 0.002
0.496 ± 0.168
DeepSAD
0.561 ± 0.024
0.755 ± 0.242
0.874 ± 0.049
0.651 ± 0.075
Appendix
Table 7: Unseen anomaly detection performance on CIFAR100, Path, OCT and Tissue.
PUSVDD
SOEL
WSAD-DT
all
0.793 ± 0.016
0.669 ± 0.021
0.641 ± 0.029
seen
0.799 ± 0.013
0.688 ± 0.021
0.642 ± 0.030
unseen
0.787 ± 0.020
0.649 ± 0.027
0.640 ± 0.028
Appendix
Table 8: Comparison of anomaly detection performance on MVTec AD.
Figure 5: Relationship between the anomaly detection performance and the hyperparameter α of the PUSVDD and the SOEL on MVTec AD. The semi-transparent area represents standard deviations.
Dataset
α^
MNIST
0.048
FashionMNIST
0.056
SVHN
0.350
CIFAR-10
0.325
Appendix
Table 9: CPE estimates of α compared with the true value ( α=0.1 ).
Class
SOEL
PUSVDD
1
0.995 ± 0.001
0.999 ± 0.000
2
0.927 ± 0.047
0.988 ± 0.003
3
0.986 ± 0.009
0.996 ± 0.001
4
0.982 ± 0.013
0.997 ± 0.001
5
0.935 ± 0.024
0.989 ± 0.009
6
0.964 ± 0.005
0.967 ± 0.013
Appendix
Table 10: Per-class AUROC on MNIST. Class 0 (digit 0) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.983 ± 0.003
0.993 ± 0.002
2
0.930 ± 0.019
0.952 ± 0.010
3
0.911 ± 0.007
0.930 ± 0.015
4
0.920 ± 0.008
0.945 ± 0.014
5
0.985 ± 0.004
0.993 ± 0.002
6
0.750 ± 0.026
0.763 ± 0.019
Appendix
Table 11: Per-class AUROC on FashionMNIST. Class 0 (T-shirt/top) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.783 ± 0.040
0.850 ± 0.048
2
0.760 ± 0.014
0.801 ± 0.027
3
0.698 ± 0.032
0.682 ± 0.029
4
0.769 ± 0.033
0.788 ± 0.025
5
0.765 ± 0.012
0.799 ± 0.022
6
0.705 ± 0.013
0.690 ± 0.018
Appendix
Table 12: Per-class AUROC on SVHN. Class 0 (digit 0) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.793 ± 0.066
0.824 ± 0.030
2
0.712 ± 0.012
0.724 ± 0.010
3
0.724 ± 0.054
0.795 ± 0.026
4
0.731 ± 0.096
0.792 ± 0.015
5
0.790 ± 0.029
0.843 ± 0.025
6
0.856 ± 0.016
0.852 ± 0.023
Appendix
Table 13: Per-class AUROC on CIFAR10. Class 0 (airplane) is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.965 ± 0.051
0.996 ± 0.005
2
0.526 ± 0.126
0.554 ± 0.123
3
0.950 ± 0.035
0.957 ± 0.007
4
0.737 ± 0.031
0.733 ± 0.068
5
0.741 ± 0.035
0.751 ± 0.096
6
0.812 ± 0.082
0.859 ± 0.086
Appendix
Table 14: Per-class AUROC on Path. Class 0 is the unseen anomaly.
Class
SOEL
PUSVDD
1
0.623 ± 0.017
0.622 ± 0.010
2
0.788 ± 0.052
0.843 ± 0.029
3
0.739 ± 0.029
0.806 ± 0.030
4
0.693 ± 0.016
0.749 ± 0.018
5
0.686 ± 0.035
0.722 ± 0.040
6
0.739 ± 0.026
0.716 ± 0.011
Appendix
Table 15: Per-class AUROC on Tissue. Class 0 is the unseen anomaly.
Figure 6: The anomaly detection performance of the PUAE and the SOEL as the fraction β of unseen anomalies among the unlabeled anomalies increases, with the total contamination ratio fixed at α=0.1 . Both methods use the same DAE as the base anomaly detector, so that only the training objective differs. The three panels show the AUROC on the whole test set, on the seen anomalies, and on the unseen anomalies, respectively. The semi-transparent area represents standard deviations.