cs.LGOct 7, 2026

Explaining the Saliency Map Sparsity of Adversarially-Trained Neural Networks

Authors: Yannick Lunk, Atell Yehor Krasnopolsky, Damien Garreau, Leon Bungert

Organizations: Institute of Mathematics University of Würzburg, Germany · Technical University of Munich, Germany · Institute of Computer Science, CAIDAS University of Würzburg, Germany · Institute of Mathematics, CAIDAS University of Würzburg, Germany

Abstract

Understanding why deep neural networks make a given prediction is of great importance for their safe deployment. In computer vision, saliency maps, which highlight the image region most influential for a prediction, remain a widely-used form of explanation. An empirical observation is the apparent sparsity of gradient saliency maps of adversarially-trained neural networks. In this paper, we propose a theoretical explanation of this phenomenon for two-layer ReLU networks. We build on the established equivalence of adversarial training to the minimization of the empirical risk with weight-decay penalization and an added adversarial total variation term -- valid for certain loss functions. As the number of data points and neurons grows and the regularization parameters are sent to zero at appropriate rates, we prove that minimizers converge to a Bayes classifier with minimal gradient and Barron norm. Sparsity appears since for adversarial training with ℓ∞\ell_\infty-attacks the gradient norm is anisotropic and favors axis-aligned / sparse gradients. We illustrate our theoretical findings experimentally by evaluating the gradient ℓ1\ell_1-norm and thresholded sparsity of naturally versus adversarially trained models.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Robustness Cannot be Reduced to Regularization: Studying Adversarial Training Beyond the Linear Case

    Jun 19, 2026David A. R. Robin, Rafael Pinot, Yann ChevaleyreAdversarial TrainingNeural Network Robustness

  2. Quantifying How Training Gradient Sparsity Affect Spiking Neural Network Accuracy And Robustness

    Sep 28, 2025Nhan Trong Luu, Duong Trung LuuNeural Network GeneralizationNeural Network Robustness

  3. Fast Adversarial Attacks with Gradient Prediction

    May 14, 2026Kamil Ciosek, Aleksandr V. Petrov, Nicolò Felicioni +1Adversarial AttacksNeural Network Robustness