The adoption of machine learning for socially relevant tasks requires effective explainable artificial intelligence (XAI) methods to better understand the behavior of machine learning models. Attribution methods are a popular XAI approach in which input-output relationships are characterized by heat maps that reflect the relative importance of input features for a particular prediction. The quality of such maps is often assessed by measuring faithfulness based on the area under insertion and deletion curves, which measures changes in the model output as features are added and removed. In this study, we derive an objective function from this notion of faithfulness and a way to approximate its gradient. We establish the connection between insertion curves and top-k feature selection, which leads to a loss function measuring the quality of attributions. Randomization of the loss allows us to efficiently approximate its gradient. To show the effectiveness of the general approach, we combine the loss function with the neural explanation mask framework. The resulting method, termed Ra-NEM, can be used with any differentiable model without affecting the model's performance. Experiments demonstrate that Ra-NEM provides accurate attributions robustly and efficiently. Compared to other algorithms, the attributions have not only higher faithfulness but also perform well in terms of other XAI metrics. The high inference speed of Ra-NEM makes the method suitable for online applications. The code is available online: https://github.com/baerminator/Ra_Nem
Figures & tables
Fig. 1: The general Neural Explanation Method (NEM) framework’s operation during inference (black and red arrows) and training (black, red, and green arrows). During inference, the input is processed by the frozen network, which generates its output, and the masking network, which produces an attribution (explanation). Depending on the NEM architecture, representations from the frozen network may assist the masking network in generating the attribution. During training, the masking network generates a mask from the attribution. In [ 10 , 11 ] , this mapping is the identity, and in this work it is a differentiable top- k feature selector. The mask is applied to the input and passed through the frozen network to produce a masked output. Both masked and unmasked outputs are used in a loss function to optimize the masking network. Prior work has also incorporated the attribution into the loss.
Method
Faith. ↑
Robust. ↓
Rand. ↓0↑
Time ↓
RISE
0.422
0.362
-0.017
12.835
Grad-CAM++ ‡
0.390
0.599
-0.082
0.013
Integrated Gradients
0.374
1.298
- 0.016
0.092
Smooth Pixel Mask
0.386
0.730
- 0.138
4.672
Grad-SHAP
0.343
1.350
- 0.015
0.014
IBA †
0.455
0.189
-0.222
0.240
TABLE I: Aggregated results from running different XAI methods on four different models using 1000 samples of the validation split of the ImageNet dataset. The explained models are a ConvNeXt, a ResNet50, a VGG16, and a ViT. We compared Faithfulness of Ra-NEM with the other methods, and the differences in the table below are statistically highly significant (two-sided paired Wilcoxon rank-sum test, p<0.001 ). Results for each individual model are given in Appendix A in the supplementary material. † IBA was only evaluated for the three CNN architectures. ‡ Grad-CAM++ was only evaluated for Robustness on the three CNNs and NEMt was only evaluated on ViT and ResNet50. Please see Appendix A for further information.
Resnet50
ConvNeXt
RISE
0.242
0.438
Integrated Gradients
0.313
0.532
Ra-NEM
0.168
0.212
TABLE II: Average ratio of input that has to be perturbed before the prediction of the base model differs from the the original. Lower is better. Perturbation method is zero imputation.
Method
Faithfulness ↑
Grad-SHAP
0.343
Grad-SHAP (FORgrad)
0.354
Integrated Gradients
0.374
Integrated Gradients (FORgrad)
0.407
Saliency
0.294
Saliency (FORgrad)
0.308
TABLE III: Comparing faithfulness on ResNet50 for Ra-NEM and gradient-based methods with and without FORgrad. We can see that FORgrad improves faithfullness at the cost of complexity. Ra-NEM is still outperforming the gradient-based methods even after they have been “repaired” by FORgrad.
Fig. 2: Model explanations generated by different XAI methodologies (see appendix for more examples). The explained model is a ResNet50 architecture trained on the training split of the ImageNet dataset. The example image is taken from the ImageNet validation dataset. Compared to other methods, such as RISE and Grad-CAM++, Ra-NEM generates a ranking more closely aligned with the central object (dog), while the others exhibit a circular bias, likely due to smoothing. This may explain the higher faithfulness of Ra-NEM (see section V ).
Method
Noise
Zero
Blur
RISE
0.266
0.267
0.198
Integrated Gradients
0.138
0.191
0.160
Ra-NEM
0.429
0.466
0.261
TABLE IV: Effect of the choice of perturbation when calculating the deletion/insertion AUC, the shown values are AUCins−AUCdel (ConvNeXt).
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Faith. ↑
Robust. ↓
Rand. ↓0↑
Time ↓
RISE
0.568
0.263
-0.034
9.235
Grad-CAM++
0.484
0.284
0.240
0.009
Integrated Gradients
0.397
1.360
0.017
0.046
Smooth Pixel Mask
0.436
0.729
0.032
3.349
Grad-SHAP
0.389
1.088
0.017
0.008
IBA
0.531
0.184
-0.068
0.117
Appendix
TABLE V: Metrics from running nine different XAI methods on a ResNet50 architecture using 1000 samples of the validation split of the ImageNet dataset. We compared Faithfulness, Complexity and Sparsity results of Ra-NEM with the other methods, and the differences in the table below, excluding RISE when measuring faithfulness, are statistically highly significant (two-sided paired Wilcoxon rank-sum test, p<0.001 ).
Method
Faith. ↑
Robust. ↓
Rand. ↓0↑
Time ↓
RISE
0.253
0.395
-0.015
14.203
Grad-CAM++
0.321
0.107
-0.002
0.015
Integrated Gradients
0.332
1.756
0.012
0.081
Smooth Pixel Mask
0.251
0.744
0.275
5.026
Grad-SHAP
0.311
2.073
0.010
0.012
IBA
0.311
0.079
-0.296
0.486
Appendix
TABLE VI: Metrics from running nine different XAI methods on a ConvNeXt architecture using 1000 samples of the validation split of the ImageNet dataset. We compared Faithfulness, Complexity and Sparsity results of Ra-NEM with the other methods, and the differences in the table below are statistically highly significant (two-sided paired Wilcoxon rank-sum test, p<0.001 ).
Method
Faith. ↑
Robust. ↓
Rand. ↓0↑
Time ↓
RISE
0.532
0.354
-0.017
16.599
Grad-CAM++
0.502
0.378
-0.485
0.013
Integrated Gradients
0.313
0.975
0.017
0.115
Smooth Pixel Mask
0.474
0.670
0.033
3.873
Grad-SHAP
0.31
1.019
0.017
0.018
IBA
0.522
0.305
-0.301
0.358
Appendix
TABLE VII: Metrics from running nine different XAI methods on a VGG16 architecture using 1000 samples of the validation split of the ImageNet dataset. We compared Faithfulness, Complexity and Sparsity results of Ra-NEM with the other methods, and the differences in the table below, excluding NEMt when measuring faithfulness, are statistically highly significant (two-sided paired Wilcoxon rank-sum test, p<0.001 ).
Method
Faith. ↑
Robust. ↓
Rand. ↓0↑
Time ↓
RISE
0.336
0.438
-0.003
11.303
Grad-CAM++
0.254
1.626
n/a
0.012
Integrated Gradients
0.456
1.104
0.017
0.124
Smooth Pixel Mask
0.384
0.777
0.213
6.441
Grad-SHAP
0.362
1.221
0.016
0.017
IBA
n/a
n/a
n/a
n/a
Appendix
TABLE VIII: Metrics from running nine different XAI methods on a ViT architecture using 1000 samples of the validation split of the ImageNet dataset. We compared Faithfulness, Complexity and Sparsity results of Ra-NEM with the other methods, and the differences in the table below, excluding Integrated Gradients when measuring faithfulness, are statistically highly significant (two-sided paired Wilcoxon rank-sum test, p<0.001 ).
Method
NEMt (sec.) ↓
Ra-NEM (sec.) ↓
ResNet50
335
700
ConvNeXt
360
869
VGG16
437
980
ViT
403
913
Appendix
TABLE IX: Time spent training NEM models for different explained models. The models are trained for 10 epochs on 10000 images extracted from the validation split of the ImageNet Model.
Round
Faithfulness ↑
1
0.602
2
0.581
3
0.595
4
0.597
5
0.582
Appendix
TABLE X: Variability in faithfulness for a Ra-NEM explaining a ResNet50 across five runs with the same hyperparameters. The results are very stable.
k
Training time ↓
Faithfullness ↑
1
252
0.520
2
310
0.560
4
463
0.596
6
700
0.607
8
967
0.617
Appendix
TABLE XI: The effect of the number of samples on faithfulness and training time for the Ra-NEM method when explaining a ResNet50 model. All hyperparameters except for number of samples are identical across runs. It can be seen that increasing k can improve faithfulness but also increase training time.
blur
noise
zero imputation
all
Faithfulness
0.450
0.465
0.464
0.468
Appendix
TABLE XII: Impact of Ra-NEM training perturbation choices on validation performance measured by faithfulness (ConvNeXt).
Fig. 3: Model explanations generated by various different XAI methodologies. The explained model is a ResNet50 architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. Comparing Ra-NEM with other occlusion-based methods, we observe that its finer granularity more precisely reveals specific image features. While all methods highlight the importance of the eye, Ra-NEM identifies only a small region of the eye and its outline as necessary.
Fig. 4: Model explanations generated by various different XAI methodologies. The explained model is a ResNet50 architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset.
Fig. 5: Model explanations generated by different XAI methodologies. The explained model is a ConvNeXt architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. Occlusion-based methods generally agree on the area of interest, but only Ra-NEM with its high granularity can outline specific features. For example, it highlights both the shaft of the hammer and the outline of the hammerhead, which indicates that most of the hammerhead is not needed for classification.
Fig. 6: Model explanations generated by different XAI methodologies. The explained model is a ConvNeXt architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. It can be seen that all occlusion-based methods agree that the face of the dog is important, but Ra-NEM also emphasizes the ears.
Fig. 7: Model explanations generated by different XAI methodologies. The explained model is a ConvNeXt architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. This example illustrates Ra-NEM’s ability to discover fine details and as such indicate that the outline of the dog is central to classification, whereas other methods biased toward smooth explanations need to indicate the entire animal to be important.
Fig. 8: Model explanations generated by different XAI methodologies. The explained model is a VGG16 architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. Notably, Ra-NEM excludes the neck brace, whereas the smoothing of the other occlusion based methods somewhat include it as an important feature. Furthermore, it can be seen that the gradient based methods all generally target the neck brace.
Fig. 9: Model explanations generated by different XAI methodologies. The explained model is a VGG16 architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset.
Fig. 10: Model explanations generated by different XAI methodologies. The explained model is a VGG16 architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset.
Fig. 11: Model explanations generated by different XAI methodologies. The explained model is a ViT architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. There is no adaptation of IBA for transformer-based image models, so IBA results have not been generated. This exemplary image illustrates a number of important points. We can see that the methods which leverage smoothing (Grad-CAM++ and RISE) generate attributions that are very scattered over the input, which might be due to an unfortunate combination of the smoothing and how the ViT processes an image. Given the ViT relies on attention mechanisms instead of convolutions, it could be that the architecture’s bias toward spatial cohesion is much lower and therefore enforcing a spatial bias in the attributions might have an adverse effect. Additionally, while Integrated Gradients may achieve high faithfulness, its attributions offer limited interpretability for end users. In contrast, Ra-NEM provides a balanced trade-off between interpretability and faithfulness.
Fig. 12: Model explanations generated by different XAI methodologies. The explained model is a ViT architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. There is no adaptation of IBA for transformer-based image models, so IBA results have not been generated.
Fig. 13: Model explanations generated by different XAI methodologies. The explained model is a ViT architecture trained on the training split of the ImageNet dataset. The example image is taken from the validation split of the ImageNet dataset. As there is no adaptation of IBA for transformer-based image models, IBA results have not been generated. This example underpins many of the points already discussed in Figure 11 , a.i. the confusion of methods leveraging smoothing and the limited interpretability of Integrated Gradients, despite its higher faithfulness score. Additionally, NEMt occasionally generates very large attributions for the ViT.
Fig. 14: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a ConVNeXt trained on the training split of the ImageNet dataset.
Fig. 15: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a ConVNeXt trained on the trainings split of the ImageNet dataset.
Fig. 16: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a ViT trained on the trainings split of the ImageNet dataset. Here Ra-NEM appears to not properly distinguish between different elements in the images.
Fig. 17: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We includeIntegrated Gradients and RISE for comparison. The explained model is a ViT trained on the trainings split of the ImageNet dataset. Here Ra-NEM again does not appear to properly distinguish between objects in the image.
Fig. 18: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a VGG16 trained on the trainings split of the ImageNet dataset. Here it can be seen, that Ra-NEM has switched foreground for background.
Fig. 19: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a VGG16 trained on the trainings split of the ImageNet dataset. Here it can be seen, that Ra-NEM has again switched foreground for background.
Fig. 20: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a ResNet50 trained on the trainings split of the ImageNet dataset. Here Ra-NEM predicts that the wrong object is the most important.
Fig. 21: Example of image from ImageNet test set, where Ra-NEM performs relatively poor compared to other methods. Score is Faithfulness. We include Integrated Gradients and RISE for comparison. The explained model is a ResNet50 trained on the trainings split of the ImageNet dataset. Here Ra-NEM again focuses on the wrong object in the image.