Adversarial training has been widely studied in recent years due to its role in improving model robustness against adversarial attacks. This paper focuses on comparing different distributed adversarial training algorithms--including centralized and decentralized strategies--within multi-agent learning environments. Previous studies have highlighted the importance of model flatness in determining robustness. To this end, we develop a general theoretical framework to study the escaping efficiency of these algorithms from local minima, which is closely related to the flatness of the resulting models. We show that when the perturbation bound is sufficiently small (i.e., when the attack strength is relatively mild) and a large batch size is used, decentralized adversarial training algorithms--including consensus and diffusion--are guaranteed to escape faster from local minima than the centralized strategy, thereby favoring flatter minima. However, as the perturbation bound increases, this trend may no longer hold. In the simulation results, we illustrate our theoretical findings and systematically compare the performance of models obtained through decentralized and centralized adversarial training algorithms. The results highlight the potential of decentralized strategies to enhance the robustness of models in distributed settings.
Figures & tables
Fig. 1: Flatness visualization of the final models trained on the CIFAR-10 dataset. Wider valleys correspond to flatter minima. The terminology random in the title of each plots mean the randomly generated graph. From these figures, we observe that when adversarial perturbations are bounded by 128/255 under the ℓ2 norm or by 3/255 under ℓ∞ norm–i.e., when the attack strength is relatively mild–decentralized adversarial training methods consistently yield flatter solutions compared to the centralized counterpart. However, when the attack strength increases to ϵ=8/255 under the ℓ∞ norm, this principle could be broken. For example, in panels (c) and (f), the models obtained via diffusion are noticeably sharper than those obtained via centralized training.
Fig. 2: Flatness visualization of the final models trained on the CIFAR-100 dataset. The observed phenomenon is consistent with Fig. 1 .
Fig. 3: Flatness visualization of the models trained on the Tiny ImageNet-200 dataset. The observed phenomenon is consistent with Fig. 1 .
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-10
random
16
ℓ∞
8/255
128
centralized
71.11
35.89
71.44
35.86
consensus
80.95
45.16
82.74
43.05
diffusion
84.02
44.46
84.38
38.71
256
centralized
65.95
33.93
66.35
34.06
consensus
81.18
43.15
82.05
42.75
TABLE I: Clean and robust accuracy of the obtained models evaluated using AutoAttack on CIFAR-10
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-100
random
16
ℓ∞
8/255
128
centralized
46.95
16.28
47.09
15.47
consensus
58.24
23.70
58.85
21.52
diffusion
58.92
20.92
57.67
18.86
256
centralized
39.96
15.32
40.06
15.40
consensus
55.33
22.49
57.09
21.07
TABLE II: Clean and robust accuracy of the obtained models evaluated using AutoAttack on CIFAR-100
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
Tiny ImageNet-200
random
16
ℓ∞
3/255
128
centralized
47.11
22.22
46.11
19.67
consensus
57.34
32.54
56.19
27.75
diffusion
54.92
27.11
54.03
25.50
ℓ2
128/255
128
centralized
47.13
24.41
46.32
22.44
consensus
60.01
37.60
58.27
31.67
TABLE III: Clean and robust accuracy of the obtained models evaluated using AutoAttack on Tiny ImageNet-200
Fig. 4: The evolution of the training error on CIFAR-10 using a random graph topology with K=16 nodes. We observe that under strong adversarial attacks, specifically when the perturbation bound is ϵ=8/255 , the optimization performance of centralized models in the large-batch setting can be bad.
Fig. 5: The evolution of the training error on CIFAR-100 using the random graph structure with K=16 nodes. Similar to Figure 4 , when ϵ=8/255 , the optimization performance of centralized models are bad.
Fig. 6: The evolution of the training error on Tiny ImageNet-200 using the random graph structure with K=16 nodes.
Figure 10Figure 11Figure 12
Appendix figures & tables24 assets
Supplementary material from the paper’s appendix.
Appendix
Data
Neural Network
Initial Learning Rate(LR)
Weight Decay
Optimizer
K
B
Epoch
LR Milestones
CIFAR-10
WideResNet-28-10
0.1
5e-4
SGD
8
128
350
[140, 210, 280]
16
128
350
[140, 210, 280]
256
400
[200, 280, 360]
32
128
400
[200, 280, 360]
SGD Momentum
16
128
350
[140, 210, 280]
256
400
[200, 280, 360]
Appendix
TABLE IV: Training Hyperparameter Settings
Data
K
B
Method
Hardware
Epoch Time
CIFAR-10
8
128
Centralized
1*A100
220s
Consensus
1*A100
270s
Diffusion
1*H100
139s
16
128
Centralized
1*H100
106s
Consensus
1*H100
164s
Diffusion
1*H100
162s
Appendix
TABLE V: The training time of each epoch
Fig. 7: Flatness visualization of the final models trained on the CIFAR-10 dataset using the ring topology with K={8,16,32} . Under relatively mild attacks, i.e., ℓ∞ perturbations bounded by ϵ=3/255 and ℓ2 perturbations bounded by 128/255 , decentralized adversarial training methods, including consensus and diffusion, converge to flatter models than the centralized strategy. However, when the attack strength increases to ϵ=8/255 with the ℓ∞ norm, diffusion can yield sharper models than centralized training. These observations remain consistent across different network sizes.
Fig. 8: Flatness visualization of the final models trained on the CIFAR-10 dataset using the grid topology with K=16 . In (c), diffusion yields a sharper model than centralized. Combining these results with those in Figures 1 and 7 , we conclude that a similar phenomenon holds across different graph topologies.
Fig. 9: Flatness visualization of the final models trained on the CIFAR-100 dataset using the ring topology with K={8,16,32} . The results shows a similar phenomenon to that observed in Figure 7 .
Fig. 10: Flatness visualization of the final models trained on the CIFAR-100 dataset using the grid topology with K=16 . The results are similar to Figure 8 .
Fig. 11: Flatness visualization of the best models trained on the CIFAR-10 dataset using the random topology with K=16 . The results are similar to Figure 1 .
Fig. 12: Flatness visualization of the best models trained on the CIFAR-100 dataset using the random topology with K=16 nodes. The results are similar to Figure 2 .
Fig. 13: Flatness visualization of the best models trained on the Tiny ImageNet-200 dataset using the random topology with K=16 nodes. The results are similar to Figure 3 .
Fig. 14: Flatness visualization of the best models trained on the CIFAR-10 dataset using the ring topology with K={8,16,32} . The results are similar to Figure 7 .
Fig. 15: Flatness visualization of the best models trained on the CIFAR-10 dataset using the grid topology with K=16 . The results are similar to Figure 8 .
Fig. 16: Flatness visualization of the best models trained on the CIFAR-100 dataset using the ring topology with K={8,16,32} . The results are similar to Figure 9 .
Fig. 17: Flatness visualization of the best models trained on the CIFAR-100 dataset using the grid topology with K=16 . The results are similar to Figure 10 .
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-10
ring
8
ℓ∞
8/255
128
centralized
78.36
37.32
78.68
33.84
consensus
84.63
44.99
84.97
38.64
diffusion
85.44
39.96
84.85
37.38
3/255
128
centralized
86.49
62.04
86.45
61.97
consensus
91.53
71.40
91.17
69.50
Appendix
TABLE VI: Clean and robust accuracy on CIFAR-10 for models trained using a ring topology
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-10
grid
16
ℓ∞
8/255
128
centralized
71.11
35.89
71.44
35.86
consensus
80.82
44.85
82.99
42.96
diffusion
84.55
44.67
84.85
39.65
3/255
128
centralized
84.91
60.97
84.50
59.55
consensus
90.18
71.77
91.00
69.22
Appendix
TABLE VII: Clean and robust accuracy on CIFAR-10 for models trained using a grid topology
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-100
ring
8
ℓ∞
8/255
128
centralized
40.51
15.23
48.97
14.30
consensus
59.08
20.55
57.68
19.19
diffusion
57.02
19.29
56.92
18.91
3/255
128
centralized
56.73
31.10
57.28
31.28
consensus
67.60
40.17
67.38
39.56
Appendix
TABLE VIII: Clean and robust accuracy on CIFAR-100 for models trained using a ring topology
dataset
graph
K
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-100
grid
16
ℓ∞
8/255
128
centralized
46.95
16.28
47.09
15.47
consensus
58.63
23.04
59.40
21.50
diffusion
60.25
21.63
58.03
19.35
3/255
128
centralized
55.13
29.07
54.57
28.27
consensus
68.55
42.99
68.02
39.18
Appendix
TABLE IX: Clean and robust accuracy on CIFAR-100 for models trained using a grid topology
Fig. 18: The evolution of the training error on CIFAR-10 using the ring graph with K={8,16,32} nodes. The centralized method exhibits particularly high training error when ϵ=8/255 .
Fig. 19: The evolution of the training error on CIFAR-10 using a grid graph with K=16 nodes. The observed phenomenon is similar to Figure 18 .
Fig. 20: The evolution of the training error on CIFAR-100 using the ring graph with K={8,16,32} nodes. The observed phenomenon is similar to Figure 18 .
Fig. 21: The evolution of the training error on CIFAR-100 using a grid graph with K=16 nodes. The observed phenomenon is similar to Figure 18 .
Fig. 22: Flatness visualization of models in the pretrained initialization scenario. Basically, when initialized from models pretrained by the centralized approach, decentralized converges to flatter regions than centralized ones.
dataset
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-10
ℓ∞
8/255
128
centralized
75.67
39.34
83.88
38.86
consensus
85.23
42.87
84.89
40.18
diffusion
76.53
41.34
84.75
40.36
256
centralized
76.87
38.70
83.27
37.25
consensus
83.60
39.98
83.35
39.04
Appendix
TABLE X: Clean and robust accuracy of models trained using PGD and the SGD momentum optimizer.
dataset
norm
ϵ
local batch
method
best
final
clean( % )
AA( % )
clean( % )
AA( % )
CIFAR-10
ℓ∞
8/255
128
centralized
81.59
43.47
81.68
40.10
consensus
82.46
45.54
83.05
41.20
diffusion
82.54
45.39
83.07
41.46
256
centralized
80.07
42.20
80.16
37.51
consensus
81.21
43.83
82.03
40.52
Appendix
TABLE XI: Clean and robust accuracy of the models trained with Trades.