RIFT: Relative Isolation From Trees For Anomaly Detection
Authors: Mark Daniel Szalai, Gabor Horvath
Organizations: Department of Networked Systems and Services, Faculty of Electrical Engineering and Informatics, Budapest University of Technology and Economics, M˝uegyetem rkp. 3, Budapest, H-1111, Hungary
Isolation Forest (IF) is a widely used baseline for unsupervised anomaly detection. Recent studies provide a closed-form expression for the infinite-forest limit for one-dimensional data. Inspired by the geometric interpretation of this formula, we introduce RIFT (Relative Isolation From Trees), a deterministic anomaly detection method that generates the minimum spanning tree and scores each point by the sum of the apparent sizes of tree edges as viewed from that point. For one-dimensional data, the RIFT score recovers the closed-form IF limit exactly. In higher dimensions, it provides a parameter-free generalization that is deterministic, robust to varying density and clustered anomalies and avoids the axis-parallel artifacts of IF. We further propose an ensemble variant for large datasets. Experiments on synthetic data and the ADBench benchmark demonstrate that the accuracy is comparable to IF, while the ensemble variant exhibits significantly lower variance across random seeds.
Figures & tables
Figure 1: Components of the explicit solution of the anomaly score for x4
Figure 2: Components of the proposed algorithm in 2D
Figure 3: Synthetic datasets used in the comparison
Method
Blob
Varying density
Clustered anomalies
Isolation Forest
0.995
0.961
0.956
LOF
0.999
1.000
0.365
KNN
0.995
0.846
0.002
HBOS
0.324
0.432
0.947
ECOD
0.526
0.656
0.983
RIFT (proposed)
0.999
0.982
0.993
Table 1: ROC AUC scores on the synthetic datasets
Figure 4: Anomaly scores on the “blobs” dataset
Figure 5: Anomaly scores on the “varying density” dataset
Figure 6: Anomaly scores on the “clustered anomalies” dataset
Figure 7: Performance of anomaly detection methods on two synthetic datasets
Dataset
Samples
Features
Category
annthyroid
7200
6
Healthcare
magic.gamma
19020
10
Physical
mammography
11183
6
Healthcare
mnist
7603
100
Image
satellite
6435
36
Astronautics
Table 2: Datasets used in the comparison
exact
Ensemble size =
Dataset
MST
4
8
16
32
64
128
256
Mean AUC scores across seeds
annthyroid
0.887
0.888
0.888
0.888
0.888
0.888
0.888
0.888
magic.gamma
0.693
0.695
0.694
0.694
0.695
0.695
0.695
0.695
mammography
0.836
0.858
0.859
0.859
0.859
0.859
0.859
0.859
mnist
0.845
0.840
0.840
0.840
0.840
0.840
0.840
0.840
Table 3: The effect of different ensemble sizes
exact
Number of samples per tree =
Dataset
MST
4
8
16
32
64
128
256
Mean AUC scores across seeds
annthyroid
0.887
0.865
0.875
0.881
0.885
0.887
0.888
0.888
magic.gamma
0.693
0.688
0.690
0.693
0.695
0.695
0.695
0.695
mammography
0.836
0.914
0.904
0.892
0.881
0.872
0.865
0.859
mnist
0.845
0.820
0.827
0.832
0.835
0.838
0.839
0.840
Table 4: The effect of the number of samples per tree
exact
Feature bagging ratio =
Dataset
MST
0.1
0.2
0.3
0.4
0.6
0.8
1.0
Mean AUC scores across seeds
annthyroid
0.887
0.840
0.840
0.840
0.858
0.877
0.887
0.888
magic.gamma
0.693
0.687
0.689
0.690
0.692
0.693
0.694
0.695
mammography
0.836
0.848
0.848
0.848
0.860
0.867
0.840
0.859
mnist
0.845
0.744
0.843
0.840
0.840
0.840
0.840
0.840
Table 5: The effect of feature bagging ratio
Mean AUC
Relative std of AUC
IF
0.7561
0.0220
RIFT exact
0.7526
0.0000
RIFT KNN
0.7526
0.0000
RIFT ensemble
0.7518
0.0010
KNN
0.7476
0.0000
CBLOF
0.7461
0.0384
Table 7: Results for datasets with at most 20,000 samples and 500 features
Mean AUC
Relative std of AUC
RIFT ensemble
0.7634
0.0014
IF
0.7632
0.0212
CBLOF
0.7507
0.0426
COPOD
0.7466
0.0000
ECOD
0.7418
0.0000
HBOS
0.7391
0.0000
Table 8: Results for all classic datasets of ADBench
Tsinghua Shenzhen International Graduate School, Tsinghua University, China · School of Computing and Information Technology, Great Bay University, China · School of Compuer Science, Nanjing University, China +3