On the Interaction of Compressibility and Adversarial Robustness
Authors: Melih Barsbey, Antônio H. Ribeiro, Umut Şimşekli, Tolga Birdal
Organizations: Department of Computing, Imperial College London, UK · Department of Information Technology, Uppsala University, Sweden · INRIA, CNRS, Département d’Informatique de l’Ecole Normale Supérieure / PSL, France
As demands for resource efficiency and safety in modern neural networks intensify, substantial research effort has gone into model compression and adversarial robustness. Yet despite progress on each in isolation, a systematic understanding of how compressibility shapes robustness remains elusive. In this paper, we develop a principled framework to analyze how different forms of structured compressibility - such as neuron-level and spectral compressibility - affect adversarial robustness. We show that structured compressibility can induce a small number of highly sensitive directions in the representation space, which adversaries can exploit to construct effective perturbations. Our analysis yields a robustness bound that reveals how neuron and spectral compressibility impact ℓ∞ and ℓ2 robustness via their effects on the learned representations. Crucially, the vulnerabilities we identify arise irrespective of how compressibility is achieved - whether via regularization, architectural bias, or learning dynamics. Through empirical evaluations across synthetic and realistic tasks, we confirm our theoretical predictions, and further demonstrate that these vulnerabilities persist under adversarial training and transfer learning, and contribute to the emergence of universal adversarial examples. Our findings show a fundamental tension between structured compressibility and robustness and highlight new pathways for designing models that are efficient and safe.
Figures & tables
Figure 1: A visual preview of our findings. (Left) Sparsification expedites compression but creates sensitive latent directions. (Center) Adversaries exploit these sensitive directions to increase their potency. (Right) This leads to decreased adversarial robustness.
Figure 2: Decision boundaries under compressibility.
Figure 3: Corollary 3.3 vs. empirical robustness gap.
Figure 4: Model statistics under increasing strength of nuclear norm regularization ( α ).
Figure 5: Results with FCN (top) and ResNet18 (bottom) trained on CIFAR-10 dataset.
Figure 6: Results with ViT (left) and CLIP (right).
Figure 7: (Left) Effects of compressibility under adversarial training. UAEs under increasing (center left) compressibility vs. (center right) parameter scale. (Right) Robustness under transfer learning.
Figure 8: Robustness under compression. SA/RA: Standard/Robust Acc. LW/Glob.: Layerwise vs. global pruning.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9: Optimizing for ℓ∞ (top) and ℓ2 (bottom) operator norms.
Figure 10: Empirically investigating the implications of Theorem 3.2 .
Figure 11: Adversarial fine-tuning (left) and training (center). Robust accuracy under increasing learning rate (right).
Figure 12: Effects of standard group lasso on compressibility and adversarial robustness.
Figure 14: Unstructured alongside structured compressibility, for row/neuron (top) and spectral compressibility (bottom).
Figure 15: Results with CIFAR-10, FCN (top) and ResNet18 (bottom), with alternative attack norms.
Figure 16: Results on CIFAR-10, ResNet18 and attacks with FGSM (top left), AutoCG (top right), Square Attack (bottom left), AutoAttack (bottom right).
Figure 17: Post-pruning fine-tuning and robustness.
ρ=0.0
ρ=0.05
ρ=0.1
ρ=0.25
ρ=0.5
δ=2/255
0.111
0.333
0.479
0.519
0.510
δ=4/255
0.002
0.061
0.263
0.371
0.390
δ=8/255
0.000
0.002
0.032
0.113
0.179
δ=16/255
0.000
0.000
0.000
0.005
0.019
Appendix
Table 1: Robust accuracy of a ViT model trained on CIFAR-10, under increasing adversarial sample ratio in training ( ρ ) vs. increasing ℓ∞ attack budgets ( δ ).
α=0.0
α=0.001
α=0.005
α=0.01
α=0.05
α=0.1
Rob. Acc.
0.383
0.362
0.369
0.219
0.123
0.111
Std. Acc.
0.920
0.926
0.921
0.893
0.873
0.829
RA/SA
0.416
0.401
0.391
0.245
0.141
0.134
Appendix
Table 2: Robust and standard accuracies of pretrained ViT models fine-tuned on CIFAR-10 dataset under varying neuron sparsification regularization strength ( α ), i.e. group lasso.
α=0.0
α=0.001
α=0.005
α=0.01
α=0.05
α=0.1
Rob. Acc.
0.384
0.360
0.357
0.326
0.155
0.083
Std. Acc.
0.889
0.877
0.887
0.880
0.881
0.875
RA/SA
0.432
0.410
0.402
0.370
0.176
0.095
Appendix
Table 3: Robust and standard accuracies of pretrained Swin Transformer models fine-tuned on SVHN dataset under varying neuron sparsification regularization strength ( α ).
α=0.0
α=0.001
α=0.005
α=0.01
α=0.05
α=0.1
Rob. Acc.
0.384
0.360
0.357
0.326
0.155
0.083
Std. Acc.
0.889
0.877
0.887
0.880
0.881
0.875
Adv. Loss - Test Loss
0.505
0.517
0.530
0.554
0.726
0.792
Appendix
Table 4: Robust and standard accuracies and loss differences for pretrained Swin Transformer models fine-tuned on SVHN dataset under varying neuron sparsification regularization strength ( α ).
Figure 18: Adversarial post-pruning fine-tuning and robustness.
Figure 19: Comparing singular values of a baseline (top) vs. compressible (bottom) model.
Figure 20: Examining alignment of a single adversarial perturbation with first 20 singular directions.
Figure 21: Comparing singular directions exploited by adversaries in baseline (left) vs. compressible (right) model.
Figure 22: Comparing pre-activation ( ∥Wa∥2/∥Wx∥2 ) and post-activation ( ∥zadv−z∥2/∥z∥2 ) representations of baseline (left) vs. compressible (right) models.
Figure 23: Utilization of singular directions by a white box PGD (top) vs. black box NES (bottom) attack under a compressible model.
Figure 24: Effects of regularizing interlayer alignment (ILA).
Figure 25: Operator norms of models under adversarial pruning vs. baselines.
Figure 26: Robustness against standard vs. universal adversarial attacks under changing Frobenius norm coefficient (left) vs. group lasso (right).
Deep neural networks (DNNs) deployed on resource-constrained neuromorphic hardware face three concurrent challenges: the need for model compression through pruning, vulnerability to adversarial input perturbations, and susceptibility to hardware-induced weight faults such as stuck-at-zero errors. While each of these factors has been studied in isolation, their combined effects on model reliability have received little attention. This paper presents an empirical investigation of how pruning, adversarial training, and hardware fault injection interact to affect the robustness of convolutional neural networks. Using a compact three-layer CNN trained on MNIST, we conduct three experiments: (1) comparing the fault tolerance of naturally and adversarially trained models under simultaneous hardware faults and adversarial attacks, (2) evaluating how pruning affects adversarial robustness, and (3) characterizing the joint accuracy surface across fault rates, adversarial perturbation magnitudes, and pruning levels. Our results show that adversarial training improves robustness against input perturbations but increases sensitivity to stuck-at-zero weight faults. Contrary to intuition, pruning did not significantly increase fault sensitivity, and varying the pruning level had little effect across fault rates and attack strengths. These results highlight the need to jointly consider adversarial robustness and hardware reliability.
Manali Dangarikar, Cory Merkel
Brain Lab Rochester Institute of Technology Rochester, NY, USA
Deep neural networks deployed in the wild must be both efficient and adaptable, requiring model compression and test-time adaptation (TTA). While both are well studied in isolation, their interaction remains poorly understood. We systematically analyze how structured compression affects a model's ability to adapt under distribution shift. Using ResNet-18 and ViT-Base on CIFAR-10-C and ImageNet-C, we evaluate multiple compression methods combined with standard TTA techniques. We introduce a diagnostic framework that examines representational expressivity and adaptation subspace compatibility. Our results reveal a consistent gap: although compressed models retain high accuracy under supervised adaptation, their TTA performance degrades significantly with increasing compression. We show that this stems from reduced representational diversity and structural constraints that limit recoverability. These effects strongly depend on the compression method, highlighting the need to design compression strategies that preserve adaptability.
Francesco Corti, Dong Wang, Young D. Kwon +2
Graz University of Technology Graz, Austria · Samsung AI Center-Cambridge University of Cambridge Cambridge, United Kingdom · University of Cambridge Cambridge, United Kingdom +1
The vulnerability of ML models to adversarial examples has recently emerged as a major concern. While adversarial training is one of the most effective countermeasures to this issue, its high computational cost remains an obstacle to practical deployment. Recent progress in reducing this cost has relied, in the case of linear models, on a formal equivalence between the adversarial risk and a simpler form of regularized risk. This enabled significantly more efficient training procedures, which naturally raises the question of whether such an equivalence can be extended beyond linear models. In this work, we formally show that no such equivalence is possible for two-layer networks. Our proofs proceed via a reduction to key properties that fundamentally separate the adversarial risk from any simple regularized risk which would only exhibit a weak form of data dependence. Beyond this setting, we provide empirical evidence on Wide-ResNets indicating that the same type of impossibility persists in deeper and more expressive architectures.
David A. R. Robin, Rafael Pinot, Yann Chevaleyre
LAMSADE, Dauphine · Universit´e Paris Dauphine PSL Research University · LPSM, Jussieu +1