On the Interaction of Compressibility and Adversarial Robustness
Authors: Melih Barsbey, Antônio H. Ribeiro, Umut Şimşekli, Tolga Birdal
Organizations: Department of Computing, Imperial College London, UK · Department of Information Technology, Uppsala University, Sweden · INRIA, CNRS, Département d’Informatique de l’Ecole Normale Supérieure / PSL, France
As demands for resource efficiency and safety in modern neural networks intensify, substantial research effort has gone into model compression and adversarial robustness. Yet despite progress on each in isolation, a systematic understanding of how compressibility shapes robustness remains elusive. In this paper, we develop a principled framework to analyze how different forms of structured compressibility - such as neuron-level and spectral compressibility - affect adversarial robustness. We show that structured compressibility can induce a small number of highly sensitive directions in the representation space, which adversaries can exploit to construct effective perturbations. Our analysis yields a robustness bound that reveals how neuron and spectral compressibility impact ℓ∞ and ℓ2 robustness via their effects on the learned representations. Crucially, the vulnerabilities we identify arise irrespective of how compressibility is achieved - whether via regularization, architectural bias, or learning dynamics. Through empirical evaluations across synthetic and realistic tasks, we confirm our theoretical predictions, and further demonstrate that these vulnerabilities persist under adversarial training and transfer learning, and contribute to the emergence of universal adversarial examples. Our findings show a fundamental tension between structured compressibility and robustness and highlight new pathways for designing models that are efficient and safe.
Figures & tables
Figure 1: A visual preview of our findings. (Left) Sparsification expedites compression but creates sensitive latent directions. (Center) Adversaries exploit these sensitive directions to increase their potency. (Right) This leads to decreased adversarial robustness.
Figure 2: Decision boundaries under compressibility.
Figure 3: Corollary 3.3 vs. empirical robustness gap.
Figure 4: Model statistics under increasing strength of nuclear norm regularization ( α ).
Figure 5: Results with FCN (top) and ResNet18 (bottom) trained on CIFAR-10 dataset.
Figure 6: Results with ViT (left) and CLIP (right).
Figure 7: (Left) Effects of compressibility under adversarial training. UAEs under increasing (center left) compressibility vs. (center right) parameter scale. (Right) Robustness under transfer learning.
Figure 8: Robustness under compression. SA/RA: Standard/Robust Acc. LW/Glob.: Layerwise vs. global pruning.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Optimizing for ℓ∞ (top) and ℓ2 (bottom) operator norms.
Figure 10: Empirically investigating the implications of Theorem 3.2 .
Figure 11: Adversarial fine-tuning (left) and training (center). Robust accuracy under increasing learning rate (right).
Figure 12: Effects of standard group lasso on compressibility and adversarial robustness.
Figure 14: Unstructured alongside structured compressibility, for row/neuron (top) and spectral compressibility (bottom).
Figure 15: Results with CIFAR-10, FCN (top) and ResNet18 (bottom), with alternative attack norms.
Figure 16: Results on CIFAR-10, ResNet18 and attacks with FGSM (top left), AutoCG (top right), Square Attack (bottom left), AutoAttack (bottom right).
Figure 17: Post-pruning fine-tuning and robustness.
ρ=0.0
ρ=0.05
ρ=0.1
ρ=0.25
ρ=0.5
δ=2/255
0.111
0.333
0.479
0.519
0.510
δ=4/255
0.002
0.061
0.263
0.371
0.390
δ=8/255
0.000
0.002
0.032
0.113
0.179
δ=16/255
0.000
0.000
0.000
0.005
0.019
Appendix
Table 1: Robust accuracy of a ViT model trained on CIFAR-10, under increasing adversarial sample ratio in training ( ρ ) vs. increasing ℓ∞ attack budgets ( δ ).
α=0.0
α=0.001
α=0.005
α=0.01
α=0.05
α=0.1
Rob. Acc.
0.383
0.362
0.369
0.219
0.123
0.111
Std. Acc.
0.920
0.926
0.921
0.893
0.873
0.829
RA/SA
0.416
0.401
0.391
0.245
0.141
0.134
Appendix
Table 2: Robust and standard accuracies of pretrained ViT models fine-tuned on CIFAR-10 dataset under varying neuron sparsification regularization strength ( α ), i.e. group lasso.
α=0.0
α=0.001
α=0.005
α=0.01
α=0.05
α=0.1
Rob. Acc.
0.384
0.360
0.357
0.326
0.155
0.083
Std. Acc.
0.889
0.877
0.887
0.880
0.881
0.875
RA/SA
0.432
0.410
0.402
0.370
0.176
0.095
Appendix
Table 3: Robust and standard accuracies of pretrained Swin Transformer models fine-tuned on SVHN dataset under varying neuron sparsification regularization strength ( α ).
α=0.0
α=0.001
α=0.005
α=0.01
α=0.05
α=0.1
Rob. Acc.
0.384
0.360
0.357
0.326
0.155
0.083
Std. Acc.
0.889
0.877
0.887
0.880
0.881
0.875
Adv. Loss - Test Loss
0.505
0.517
0.530
0.554
0.726
0.792
Appendix
Table 4: Robust and standard accuracies and loss differences for pretrained Swin Transformer models fine-tuned on SVHN dataset under varying neuron sparsification regularization strength ( α ).
Figure 18: Adversarial post-pruning fine-tuning and robustness.
Figure 19: Comparing singular values of a baseline (top) vs. compressible (bottom) model.
Figure 20: Examining alignment of a single adversarial perturbation with first 20 singular directions.
Figure 21: Comparing singular directions exploited by adversaries in baseline (left) vs. compressible (right) model.
Figure 22: Comparing pre-activation ( ∥Wa∥2/∥Wx∥2 ) and post-activation ( ∥zadv−z∥2/∥z∥2 ) representations of baseline (left) vs. compressible (right) models.
Figure 23: Utilization of singular directions by a white box PGD (top) vs. black box NES (bottom) attack under a compressible model.
Figure 24: Effects of regularizing interlayer alignment (ILA).
Figure 25: Operator norms of models under adversarial pruning vs. baselines.
Figure 26: Robustness against standard vs. universal adversarial attacks under changing Frobenius norm coefficient (left) vs. group lasso (right).
Graz University of Technology Graz, Austria · Samsung AI Center-Cambridge University of Cambridge Cambridge, United Kingdom · University of Cambridge Cambridge, United Kingdom +1