Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks
Authors: Yannick Lunk, Atell Yehor Krasnopolsky, Damien Garreau, Leon Bungert
Organizations: Institute of Mathematics University of Würzburg, Germany · Technical University of Munich, Germany · Institute of Computer Science, CAIDAS University of Würzburg, Germany · Institute of Mathematics, CAIDAS University of Würzburg, Germany
Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the regularization parameters are sent to zero at appropriate rates, we prove that minimizers converge to a Bayes classifier with minimal gradient and Barron norm. Sparsity appears since for adversarial training with ℓ∞-attacks the gradient norm is anisotropic and favors axis-aligned / sparse gradients. We illustrate our theoretical findings experimentally by evaluating the gradient ℓ1-norm and thresholded sparsity of naturally versus adversarially trained models.
Figures & tables
Figure 1: Comparison of saliency maps for a naturally trained ResNet50 on ImageNet versus adversarially trained models across varying ℓ∞ perturbation budgets ε . As predicted by Theorem 1 , with decreasing adversarial budget ε the gradient saliency maps converge to a sparse one whereas natural training (corresponding to ε=0 ) produces a noisy saliency map.
Figure 2: Decision boundaries and training data for a data distribution with non-uniquely determined Bayes classifier in the strip [−0.5,0.5]2 : Natural training (left) has a highly oscillatory decision boundary arising from fitting the training data. Nonlocal TV regularization (center, analyzed here) and adversarial training (right) have a decision boundary with small anisotropic perimeter and exhibit approximate gradient sparsity, as indicated by the axis-parallel parts of the decision boundary. See Appendix F for more details.
Figure 3: Accuracy and gradient metrics for ResNet18 on CIFAR-10 (top) and ResNet50 on ImageNet (bottom). From left to right, we display clean accuracy, robust accuracy, the average ℓ1 -norm, and the average sparsity (thresholded at different values) of the input gradients.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: Clean and adversarial accuracies of the models from Figure 2 , trained on the toy dataset using the architecture ( 1 ). We used n=1,000 data points, mn=2,000 neurons, εn=0.24 , and λn=0.01⋅εn . The Bayes (i.e., best achievable) accuracy for this dataset is 80% .
Figure 5: Decision boundaries, using ℓ2 -attacks instead of ℓ∞ -attacks: As expected, no gradient sparsity in the undetermined region [−0.5,0.5]2 is observed, but the decision boundaries of TV regularization and adversarial training still have much lower complexity and are shorter than the one of natural training.
Figure 6: Clean and adversarial accuracies of the two-layer neural network models trained on two classes of CIFAR-10. We used n=10,000 , m=4,096 neurons, εn=0.0587 , and λn=0.0055⋅εn .
Figure 7: Sparsity of the saliency maps of the two-layer models. Larger value of the bar indicates more values below the threshold τ .
Figure 8: Distribution of input gradient magnitudes for naturally and adversarially trained ResNet50 models on the ImageNet dataset. Notably, the adversarially trained model exhibits a higher frequency of near-zero values.
Figure 9: Off-target saliency maps across varying adversarial budgets. Left: The original image (ground truth: hummingbird ). Right: Input gradients from a naturally trained ResNet50 and adversarially trained models with increasing budgets. The gradients are computed with respect to incorrect target classes rather than the ground truth: goldfish (top row) and tennis ball (bottom row).
The vulnerability of ML models to adversarial examples has recently emerged as a major concern. While adversarial training is one of the most effective countermeasures to this issue, its high computational cost remains an obstacle to practical deployment. Recent progress in reducing this cost has relied, in the case of linear models, on a formal equivalence between the adversarial risk and a simpler form of regularized risk. This enabled significantly more efficient training procedures, which naturally raises the question of whether such an equivalence can be extended beyond linear models. In this work, we formally show that no such equivalence is possible for two-layer networks. Our proofs proceed via a reduction to key properties that fundamentally separate the adversarial risk from any simple regularized risk which would only exhibit a weak form of data dependence. Beyond this setting, we provide empirical evidence on Wide-ResNets indicating that the same type of impossibility persists in deeper and more expressive architectures.
David A. R. Robin, Rafael Pinot, Yann Chevaleyre
LAMSADE, Dauphine · Universit´e Paris Dauphine PSL Research University · LPSM, Jussieu +1
Spiking Neural Networks (SNNs) have recently received increasing attention in both computational neuroscience and artificial intelligence owing to their potential for energy-efficient computation and reduced memory requirements. Despite these advantages, improving adversarial robustness in SNNs (particularly for vision-based applications) remains an emerging and relatively underexplored research problem. Recent work has suggested that encouraging sparse gradients can act as a regularization mechanism to improve resistance against adversarial perturbations. In this study, we report an unexpected observation: under certain architectural configurations, SNNs inherently exhibit sparse gradients and can attain state-of-the-art adversarial defense performance without requiring any explicit regularization strategy. Further investigation reveals an inherent trade-off between robustness and generalization. Specifically, increased gradient sparsity enhances resistance to adversarial attacks but may reduce the model's generalization capability, whereas denser gradients tend to improve generalization while simultaneously increasing susceptibility to adversarial perturbations. These findings provide new perspectives on the role of gradient sparsity in the training dynamics of SNNs.
Nhan Trong Luu, Duong Trung Luu
College of Information and Communication Technology, Can Tho University, Vietnam · College of Engineering and Computer Science, VinUniversity, Hanoi, Vietnam · Center of Digital Transformation and Communication, Can Tho University, Vietnam
Generating adversarial examples at scale is a core primitive for robustness evaluation, adversarial training, and red-teaming, yet even "fast" attacks such as FGSM remain throughput-limited by the cost of a backward pass. We introduce a family of attacks that eliminates the backward pass by predicting the input gradient from forward-pass hidden states via a lightweight linear regression. Theoretically, we derive exact affine conditional gradient means, showing optimality in the (idealized) NTK regime. Empirically, our methods work when applied to practical finite-width models; we recover much of FGSM's attack performance while using only a small fraction of the time, corresponding to a 532% increase in throughput. These results suggest gradient prediction as a simple and general route to significantly faster adversarial generation under realistic wall-clock constraints.
Kamil Ciosek, Aleksandr V. Petrov, Nicolò Felicioni +1