Accurate and explainable out-of-distribution (OOD) detection is required to use machine learning systems safely. Previous work has shown that feature distance to decision boundaries can be used to identify OOD data effectively. In this paper, we build on this intuition and propose a post-hoc OOD detection method that, given an input, calculates the distance to decision boundaries by leveraging counterfactual explanations. Since computing explanations can be expensive for large architectures, we also propose strategies to improve scalability by computing counterfactuals directly in embedding space. Crucially, as the method employs counterfactual explanations, we can seamlessly use them to help interpret the results of our detector. We show that our method is in line with the state of the art on CIFAR-10, achieving 93.50% AUROC and 25.80% FPR95. Our method outperforms these methods on CIFAR-100 with 97.05% AUROC and 13.79% FPR95 and on ImageNet-200 with 92.55% AUROC and 33.55% FPR95 across four OOD datasets
Figures & tables
Figure 1: We present an instantiation of our framework with a toy example of a three-class classification problem. In-distribution points are represented by the coloured circles, triangles, and squares. We then consider two new inputs: an OOD input (black pentagon) and an ID input (white pentagon). Assume both inputs are initially classified as class 3. In this example, we compute two counterfactuals (the grey crosses) for each input, one for class 1 and one for class 2, and compute the distances to these counterfactuals. We observe that the OOD input is closer to the decision boundaries and thus results in a smaller average counterfactual distance. Instead, the ID input is farther from the decision boundaries and thus will have a greater average counterfactual distance than the anomaly point. In this paper, we leverage this intuition to build an OOD detection method that uses counterfactual distances as OOD scores to identify anomalies.
SVHN
iSUN
Textures
Places365
Average
Method
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
MSP
25.43
91.56
23.05
92.36
35.22
89.89
42.46
88.92
31.54
90.68
ODIN
68.54
84.70
27.20
94.63
67.89
86.94
70.40
85.08
58.51
87.84
Energy
34.20
91.98
27.69
93.24
52.01
89.46
54.69
89.25
42.15
90.98
ViM
19.01
94.58
21.12
94.20
21.14
95.16
41.44
89.50
25.68
93.36
MDS
25.88
91.17
32.09
90.11
28.03
92.69
47.69
84.91
33.42
89.72
Table 1: Evaluation on CIFAR-10. We use the OpenOOD implementation for all benchmarks and average the values over three training runs from their saved models. We give results on our method using two different counterfactual search methods, NICE and NNCE. The best results are in bold.
SVHN
iSUN
Textures
Places365
Average
Method
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
MSP
58.43
78.68
51.05
82.06
61.79
77.32
56.65
79.21
56.98
79.32
ODIN
67.18
74.72
35.50
90.65
62.41
79.34
59.73
79.45
56.21
81.04
Energy
53.19
82.28
45.53
85.82
62.39
78.35
57.69
79.50
54.70
81.49
ViM
46.37
82.87
46.64
82.13
47.22
85.75
61.57
75.86
50.45
81.65
MDS
67.47
70.32
77.48
63.77
70.21
76.42
79.19
63.41
73.59
68.48
Table 2: Evaluation on CIFAR-100. We use the OpenOOD implementation for all benchmarks and average the values over three training runs from their saved models. We give results on our method using two counterfactual search methods, NICE and NNCE. The best results are in bold.
SSB-hard
iNaturalist
Textures
OpenImage-O
Average
Method
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
FPR95 ↓
AUROC ↑
MSP
66.09
80.32
35.39
90.13
44.43
88.37
35.21
89.24
45.28
87.01
ODIN
73.52
77.17
22.49
94.34
42.96
90.67
37.29
90.10
44.07
88.07
Energy
69.97
79.72
26.41
92.52
41.29
90.80
36.80
89.20
43.62
88.06
ViM
71.59
74.00
26.81
91.38
20.19
94.77
33.96
88.21
38.14
87.09
MDS
83.60
58.52
58.13
75.34
58.13
79.25
67.87
70.15
66.93
70.81
Table 3: Evaluation on ImageNet-200. We use the OpenOOD implementation for all benchmarks and average the values over three training runs from their saved models. We give results on our method using the counterfactual search method NNCE. The best results are in bold.
Figure 2: Exemplar-based explanation for a novel class. The input (top) is an ‘8’ (OOD), misclassified as a ‘3’. The top row shows the Nearest Like Neighbours (closest 3’s). The bottom rows show Nearest Unlike Neighbours from other classes. Distances reveal the input is actually closer to class ‘2’ than the predicted class, highlighting the model’s uncertainty.
Appendix figures & tables2 assets
Supplementary material from the paper’s appendix.
Appendix
k
FPR95 ↓
AUROC ↑
Avg Time (ms)
25
14.93
96.85
2.24
50
14.06
97.01
4.53
75
14.10
96.95
6.65
100
13.79
97.05
8.74
Appendix
Table 4: Ablation study on CIFAR-100 OOD benchmarks using our method with NNCE averaged over OOD datasets SVHN, iSUN, Textures, and Places365 and averaged over three training runs, using the same ResNet-18 network as in Table 2 . Here, k is the number of classes we compute counterfactuals for and thus, how many decision boundaries we consider in our distance metric. Avg Time (ms) refers to the milliseconds it takes to compute our OOD score on a machine equipped with a 28-core CPU with 512GB RAM. We average the computation time over 100 inputs.
Figure 3: Explanation for a distribution shift. The input (top) is a ‘5’ misclassified as a ‘0’. Despite the prediction, the explanation indicates that the input is semantically closer to the ‘5’ class (Nearest Unlike Neighbours) than to the predicted ‘0’ class (Nearest Like Neighbours), thereby aiding error diagnosis.
Out-of-distribution (OOD) detection is essential for deploying machine learning models in open-world and safety-critical scenarios, where test inputs may deviate from the training distribution and overconfident predictions on unknown samples can lead to unreliable decisions. Outlier Exposure (OE) has emerged as a promising OOD detection paradigm by introducing auxiliary outliers during training to enlarge the margin between in-distribution (ID) and OOD samples. Existing OE-based methods typically enlarge this margin by employing uniform labels to maximize the entropy of OOD samples over ID categories. However, we theoretically show that uniform labels inevitably disregard the relations between OOD samples and ID categories, termed the over-softening effect, leading to a suboptimal margin bound. Our theoretical analysis further reveals that explicitly exploiting such relations can instead yield improved OOD detection performance. Motivated by this insight, we propose \underline{A}daptive Confidence \underline{OE} (AOE), a simple yet effective method that leverages temperature scaling to recalibrate outlier labels. Specifically, AOE generates adaptive soft targets from temperature-scaled model predictions for OOD samples, where the learnable temperature smooths the prediction distribution without fully erasing class-wise relational information. By supervising OOD samples with these adaptive soft targets, AOE preserves the semantic proximity between OOD samples and ID categories while encouraging the softened targets to approach a high-entropy distribution, thereby suppressing overconfident OOD predictions and enlarging the separation margin. Extensive experiments across diverse benchmarks demonstrate the effectiveness of AOE.
Fengqiang Wan, Qing-Yuan Jiang, Yang Yang
School of Computer Science and Engineering, Nanjing University of Science and Technology, Nanjing 210094, China.
Reliable detection of out-of-distribution (OOD) samples is crucial for the safe deployment of machine learning models. Neural networks often produce overconfident predictions for inputs that deviate from their training data, leading to significant degradation in performance. While many OOD detection methods focus on the final output layer, they neglect the rich hierarchical information present in intermediate network layers. This paper introduces a novel approach that leverages sparse autoencoders (SAEs) to learn interpretable features from these intermediate activations. We find that in-distribution (ID) and OOD data activate distinct sets of these sparse features. We propose a new OOD score derived from the cosine similarity between the sparse feature activations of a test sample and the mean activations of ID classes. Our post-hoc detection method not only achieves state-of-the-art performance on standard OOD detection benchmarks, but yields interpretable insights into how distribution shift affects learned representations.
Ayush Karmacharya, Luke Luschwitz, Lucia Romero +2
Reliable out-of-distribution (OOD) detection is a critical requirement for the safe deployment of machine learning systems. Despite recent progress, state-of-the-art OOD detectors are highly susceptible to adversarial attacks, which undermines their trustworthiness in automated systems. To address this vulnerability, we apply median smoothing to baseline OOD detection scores, balancing clean and adversarial accuracies. Our key insight is that the noisy samples generated for median smoothing can be repurposed to quantify the local instability of the base score. We observe that OOD samples exhibit higher instability under perturbation. Based on this, we propose ROSS, a novel and robust post-hoc OOD detector that leverages the instability of baseline scores to further distinguish between in-distribution (ID) and OOD samples. ROSS achieves symmetric robustness, performing strongly against both score-minimising and score-maximising attacks, unlike prior work. This symmetric defence leads to state-of-the-art robustness, outperforming prior methods by up to 40 AUROC points. We demonstrate ROSS's effectiveness on extensive experiments across CIFAR-10, CIFAR-100, and ImageNet. Code is available at: https://github.com/Abdu-Hekal/ROSS.