This work investigates a novel approach to boost adversarial robustness and generalization by incorporating structural prior into the design of deep learning models. Specifically, our study surprisingly reveals that existing dictionary learning-inspired convolutional neural networks (CNNs) are robust against random noise but remain highly vulnerable to adversarial attacks. To address this, we propose Elastic Dictionary Learning Networks (EDLNets), a novel ResNet architecture that significantly enhances adversarial robustness and generalization. Extensive and reliable experiments demonstrate consistent improvements in adversarial robustness across multiple datasets, backbone architectures, and threat models. To the best of our knowledge, this is the first work to discover and validate that dictionary structure can reliably enhance deep learning robustness under strong adaptive attacks, unveiling a promising direction for future research.
Figures & tables
Figure 1: Overview of Elastic Dictionary Learning Networks (EDLNets). EDLNets are constructed by replacing the convolutional layers in conventional backbones (e.g., ResNets) with EDL layers obtained by unrolling the proposed efficient RISTA algorithm. Each EDL layer incorporates a dictionary-structure prior, assuming that the input signal x(l) can be represented by a sparse code x(l+1) using a small number of atoms from the dictionary A(l) . During backward propagation, the task loss is propagated through all unrolled RISTA iterations to jointly optimize the layer-wise dictionaries {A(l)}l=0L−1 and balancing parameters {β(l)}l=0L−1 , as well as the classifier parameters θ .
Model \ Noise Level
L-1
L-2
L-3
L-4
L-5
PGD
Classic ConvNet (ResNet18)
81.44
57.23
48.32
32.49
16.98
0.00
Vanilla DL (Fixed λ=0.1 )
82.39
68.90
59.28
40.8
23.83
0.01
Vanilla DL (Tuned λ )
82.39
68.90
59.28
43.71
33.43
0.13
Table 1: Preliminary study on CIFAR10 comparing standard ResNet18 with a Vanilla DL variant under varying levels of random noise and adaptive PGD attack ( ϵ=2558 ). Vanilla DL exhibits strong robustness against weak random noise but is easily broken by the strong PGD attack.
Natural Acc.
Robust Acc.
Method
Final
Best
Diff
Final
Best
Diff
Vanilla
78.98
79.90
0.92
44.90
48.01
3.11
ℓ1 Reg.
64.84
65.71
0.87
40.94
41.97
1.03
ℓ2 Reg.
78.88
79.39
0.51
42.73
48.26
5.53
ℓ2 + ℓ1 Reg.
66.86
67.62
0.76
42.53
43.33
0.80
Cutout
75.11
75.58
0.47
47.12
48.23
1.11
Table 2: Natural and robust performance of PGD-based adversarial training with different methods to mitigate the overfitting on CIFAR10. BEST represents the highest test accuracy achieved during training, while FINAL is the average accuracy over the last five epochs. DIFF, the difference between BEST and FINAL, measures the ability to mitigate overfitting. The best performance is highlighted in bold , while the second-best is underlined .
Figure 4
Figure 3: Adversarial robustness under various settings. Our Elastic DL outperforms Vanilla DL across various datasets (CIFAR10 / CIFAR100 / Tiny-ImageNet), backbones (ResNet10 / ResNet18 / ResNet34 / ResNet50) and attacks (PGD / FGSM / CW / AA).
Figure 6
Figure 6: Hidden embedding visualization under clean and attacked scenarios. The difference between clean and attacked embeddings in Elastic DL is smaller compared to Vanilla DL, with this effect becoming more significant in deeper layers. Consequently, while an adversarial attack alters the Vanilla DL output from ” SHIP ” to ” FROG ”, Elastic DL successfully preserves the correct prediction.
Figure 8
Method
Train (h)
Params (M)
FLOPs (G)
Size (MB)
Vanilla DL
13.8
11.174
4.776
42.71
Elastic DL
15.2
11.174
4.776
42.71
Table 6: Computational cost comparison.
Appendix figures & tables20 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10: Comparison of training-from-scratch and pretraining–finetuning strategies for Elastic DL. Pretraining–finetuning achieves more stable optimization and better final performance.
Figure 11: Performance of the Vanilla DL variant of ResNet18 under random impulse noise at different levels.
Figure 12: Adversarial training curve of our Elastic DL. During epochs 100–150, the model experiences a catastrophic robust overfitting problem. Introducing the Elastic DL structural prior at epoch 150 and fine-tuning effectively mitigates overfitting and yields significantly improved robustness and generalization.
Figure 13: Training curves of baselines.
Figure 14: Comparison of training curves across all methods.
Method
Natural
PGD
FGSM
C&W
AA
Vanilla DL + ResNet10
81.55
45.48
52.53
45.85
41.60
Elastic DL + ResNet10
82.69
49.54
64.52
57.37
46.30
Vanilla DL + ResNet18
83.28
45.64
53.88
41.22
43.70
Elastic DL + ResNet18
83.57
53.22
69.35
60.80
52.90
Vanilla DL + ResNet34
82.45
45.37
54.32
42.12
44.40
Elastic DL + ResNet34
82.95
55.88
70.19
61.74
53.80
Appendix
Table 7: Adversarial robustness on CIFAR10 with different backbones.
Method
Natural
PGD
FGSM
C&W
AA
Vanilla DL + ResNet10
55.94
22.45
26.57
18.90
21.00
Elastic DL + ResNet10
55.20
26.30
35.34
26.45
22.60
Vanilla DL + ResNet18
57.24
22.17
26.81
17.43
21.60
Elastic DL + ResNet18
57.70
27.27
37.62
28.87
26.30
Vanilla DL + ResNet34
56.18
21.77
26.14
16.38
20.80
Elastic DL + ResNet34
56.38
32.67
43.39
39.34
29.20
Appendix
Table 8: Adversarial robustness on CIFAR100 with different backbones.
Method
Natural
PGD
FGSM
C&W
AA
Vanilla DL + ResNet10
49.60
27.17
32.46
37.91
20.20
Elastic DL + ResNet10
50.12
32.93
39.64
40.10
24.90
Vanilla DL + ResNet18
50.22
31.45
36.46
39.02
30.90
Elastic DL + ResNet18
50.52
37.60
43.10
46.64
36.30
Vanilla DL + ResNet34
50.03
33.54
37.24
37.19
29.30
Elastic DL + ResNet34
50.40
34.80
41.75
44.72
34.60
Appendix
Table 9: Adversarial robustness on Tiny-ImageNet with different backbones.
Figure 15: Different adversarial training methods. Our Elastic DL is orthogonal to existing adversarial training methods and can be combined with them to further improve performance.
Figure 16: Different attack measurements. Our Elastic DL consistently outperforms Vanilla DL across attacks (PGD- ℓ∞ , PGD- ℓ2 , SparseFool) evaluated under various metrics ( ℓ∞,ℓ2,ℓ0 norms).
Num. of Layers
3
5
8
10
15
20
25
30
40
50
Relative Error
0.00066
0.00095
0.00152
0.00092
0.00118
0.00098
0.00112
0.00082
0.00060
0.00093
Appendix
Table 10: Zero-order gradient analysis: relative difference between autograd and zero-order gradients across network depths.
Budget
0
8/255
16/255
32/255
64/255
Elastic DL
83.57
53.29
34.13
23.84
12.45
Elastic DL + BPDA
83.57
53.20
34.08
23.51
12.44
Appendix
Table 11: Adversarial robustness under standard adaptive attack versus BPDA attack, showing comparable performance and ruling out gradient obfuscation.
Table 12: Reconstruction error between the recovered noise ϵ^ and various input noises, including random noise ( ϵrandom ), transfer noise from ResNet ( ϵresnet ) and Vanilla DL ( ϵvanilla ), and adaptive noise from our Elastic DL ( ϵelastic ). Our Elastic DL achieves the smallest reconstruction error, indicating that our approach can adaptively recover and neutralize the input perturbation, thereby mitigating its impact.
Figure 20: Reconstruction process.
Figure 21: Reconstruction process (ImageNet).
Figure 22: Reconstruction process (CIFAR10, Part 1).
Figure 23: Reconstruction process (CIFAR10, Part 2).