Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks
Authors: Yannick Lunk, Atell Yehor Krasnopolsky, Damien Garreau, Leon Bungert
Organizations: Institute of Mathematics University of Würzburg, Germany · Technical University of Munich, Germany · Institute of Computer Science, CAIDAS University of Würzburg, Germany · Institute of Mathematics, CAIDAS University of Würzburg, Germany
Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the regularization parameters are sent to zero at appropriate rates, we prove that minimizers converge to a Bayes classifier with minimal gradient and Barron norm. Sparsity appears since for adversarial training with ℓ∞-attacks the gradient norm is anisotropic and favors axis-aligned / sparse gradients. We illustrate our theoretical findings experimentally by evaluating the gradient ℓ1-norm and thresholded sparsity of naturally versus adversarially trained models.
Figures & tables
Figure 1: Comparison of saliency maps for a naturally trained ResNet50 on ImageNet versus adversarially trained models across varying ℓ∞ perturbation budgets ε . As predicted by Theorem 1 , with decreasing adversarial budget ε the gradient saliency maps converge to a sparse one whereas natural training (corresponding to ε=0 ) produces a noisy saliency map.
Figure 2: Decision boundaries and training data for a data distribution with non-uniquely determined Bayes classifier in the strip [−0.5,0.5]2 : Natural training (left) has a highly oscillatory decision boundary arising from fitting the training data. Nonlocal TV regularization (center, analyzed here) and adversarial training (right) have a decision boundary with small anisotropic perimeter and exhibit approximate gradient sparsity, as indicated by the axis-parallel parts of the decision boundary. See Appendix F for more details.
Figure 3: Accuracy and gradient metrics for ResNet18 on CIFAR-10 (top) and ResNet50 on ImageNet (bottom). From left to right, we display clean accuracy, robust accuracy, the average ℓ1 -norm, and the average sparsity (thresholded at different values) of the input gradients.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Clean and adversarial accuracies of the models from Figure 2 , trained on the toy dataset using the architecture ( 1 ). We used n=1,000 data points, mn=2,000 neurons, εn=0.24 , and λn=0.01⋅εn . The Bayes (i.e., best achievable) accuracy for this dataset is 80% .
Figure 5: Decision boundaries, using ℓ2 -attacks instead of ℓ∞ -attacks: As expected, no gradient sparsity in the undetermined region [−0.5,0.5]2 is observed, but the decision boundaries of TV regularization and adversarial training still have much lower complexity and are shorter than the one of natural training.
Figure 6: Clean and adversarial accuracies of the two-layer neural network models trained on two classes of CIFAR-10. We used n=10,000 , m=4,096 neurons, εn=0.0587 , and λn=0.0055⋅εn .
Figure 7: Sparsity of the saliency maps of the two-layer models. Larger value of the bar indicates more values below the threshold τ .
Figure 8: Distribution of input gradient magnitudes for naturally and adversarially trained ResNet50 models on the ImageNet dataset. Notably, the adversarially trained model exhibits a higher frequency of near-zero values.
Figure 9: Off-target saliency maps across varying adversarial budgets. Left: The original image (ground truth: hummingbird ). Right: Input gradients from a naturally trained ResNet50 and adversarially trained models with increasing budgets. The gradients are computed with respect to incorrect target classes rather than the ground truth: goldfish (top row) and tennis ball (bottom row).
College of Information and Communication Technology, Can Tho University, Vietnam · College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam · Center of Digital Transformation and Communication, Can Tho University, Vietnam