Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework
Authors: Yaohua Liu, Yifan Guo, Jiaxin Gao
Organizations: The University of Hong Kong Hong Kong, China · International School of Information Science & Engineering Dalian University of Technology, Dalian, China · The Hong Kong Polytechnic University Hong Kong, China
A key bottleneck in adversarial transfer is a trajectory-level geometric disconnect: ambient gradients often drift away from the intrinsic data manifold, causing surrogate-specific overfitting. To rectify this, we propose Manifold Anchored Bilevel Transfer (MABT), a unified framework that anchors adversarial trajectories to the shared semantic subspace. MABT introduces a relaxed manifold-anchoring operator as a semantic rectifier to suppress off-manifold noise. With this constraint, we cast transfer attack generation as a distributional bilevel optimization problem that learns a geometry-aligned initialization by minimizing expected transfer risk under a surrogate uncertainty distribution. We further develop a Hessian-free solver with linear-time complexity to handle the resulting hierarchy. Experiments demonstrate improved transferability for 10 baseline attackers across 28 attack configurations, diverse victim architectures, and defense mechanisms.
Figures & tables
Figure 1: Motivation. (a) Manifold hypothesis : Natural images concentrate near a low-dimensional manifold M capturing intrinsic semantics, where decision boundaries of diverse models tend to partially align. (b) Off-manifold divergence : Standard iterative attacks inherently drift into high-dimensional noise away from M ; in this region, boundaries are less consistent, so the trajectory may cross the surrogate boundary but miss the victim boundary . (c) Manifold anchoring : MABT uses a manifold-anchoring projector with a geometry-aligned initialization to suppress off-manifold deviations while preserving task-relevant semantics to cross shared boundaries .
Figure 2: Feature space divergence. We quantify victim-side feature shift during the attack process as 1−cos(fclean,fadv) , where x1→x10 denotes the adversarial trajectory over 10 attack iterations. Compared with GAA [ 16 ] +AWT [ 6 ] , MABT induces substantially larger semantic shifts on victim models, with relative gains of 99.8% on Transformers and 32.9% on CNNs.
Untargeted Attack Scenario, ResNet-50 backbone, Average Success Rate (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
16.38
13.04
37.52
36.50
12.80
17.10
10.18
9.40
5.86
10.98
7.50
5.56
15.24
Ours
34.96
24.58
70.00
69.22
27.84
33.94
19.72
16.38
10.28
17.02
11.54
7.62
28.59
GHOST
18.34
14.44
41.56
41.72
13.60
19.78
10.98
10.00
6.02
12.10
7.14
5.90
16.80
Ours
37.38
27.52
74.66
72.58
30.84
36.72
21.18
17.66
10.50
18.24
11.92
7.72
30.58
Table 1: ASR (%) of MABT across 16 attack configurations on the ResNet-50 surrogate. Results cover 4 base attackers combined with ensemble/model-related strategies. Best results are in bold .
Method
Average ASR (%) ↑ , ResNet-50
Avg.
CNN
CNN Ensemble
Transformer
PGD
N/A
22.22
8.48
8.01
15.24
BETAK
30.87
11.75
10.47
20.99
Ours
43.42
15.46
12.06
28.59
MI
N/A
33.74
13.15
12.99
23.41
BETAK
43.86
19.74
15.97
30.86
Table 2: Comparison with initialization-based BETAK which uses auxiliary pseudo-victims (Inc-v3).
Defense Mechanism, ResNet-50 backbone, Average Success Rate (%) ↑
Method
HGD
JPEG
RS
R&P
NIPS-r3
Avg.
Method
HGD
JPEG
RS
R&P
NIPS-r3
Avg.
N/A
16.94
11.60
7.96
12.12
14.28
12.58
N/A
27.12
17.54
9.44
13.40
23.10
18.12
Ours
36.16
21.38
10.42
13.90
29.32
22.24
Ours
59.82
39.68
14.80
18.28
53.68
37.25
GHOST
19.34
11.84
8.24
12.70
15.76
13.58
GHOST
30.94
19.02
9.40
13.92
24.80
19.62
Ours
38.86
21.94
10.40
14.26
31.44
23.38
Ours
63.02
41.56
15.08
19.40
56.08
39.03
MBA
22.22
13.22
9.22
12.84
19.80
15.46
MBA
40.10
25.82
11.94
15.52
35.80
25.84
Table 3: ASR (%) results against five representative defenses using the ResNet-50 surrogate.
Figure 3: Adversarial loss landscape . Surfaces are plotted along the gradient direction and a random orthogonal direction . Unlike baselines that suffer from landscape collapse on unknown victims, MABT preserves higher victim-side loss basins with relatively flatter surrounding curvature across architectures, which indicates that the learned trajectory better retains transfer-relevant directions.
Figure 4: Convergence and efficiency analysis. Left & Mid: Increasing T improves geometry-aligned initialization to mitigate surrogate overfitting, thereby ensuring sustained victim loss growth during the subsequent attack. Right: Resize provides a favorable low-cost MAP realization, while generative priors achieve stronger anchoring at higher computational cost.
Figure 5: Feature shift consistency. We plot surrogate-side (ResNet-50) feature shift against the average victim-side shift. Shaded regions denote the standard deviation over four heterogeneous victims (Inc-v3, DenseNet, DeiT-s and ConViT-b).
Method
Average ASR (%) ↑ , ResNet-50
Avg.
CNN
CNN Ensemble
Transformer
N/A
22.22
8.48
8.01
15.24
w/ DBO,w/o MAP
30.82
10.11
8.40
20.04
w/o DBO,w/ MAP
29.24
12.67
9.67
20.21
w/ DBO,w/ MAP
43.42
15.46
12.06
28.59
AWT
25.99
11.57
9.05
18.15
Table 4: Ablation results by implementing DBO and MAP based on PGD and PGD+AWT.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
16.38
13.04
37.52
36.50
12.80
17.10
10.18
9.40
5.86
10.98
7.50
5.56
15.24
FFT
32.42
23.00
64.06
64.22
24.34
28.66
17.96
15.18
8.98
14.44
10.90
7.12
25.94
Bit-Depth
34.68
25.50
69.78
70.78
25.38
33.76
19.08
15.44
9.50
20.28
11.08
7.50
28.56
Resize
34.96
24.58
70.00
69.22
27.84
33.94
19.72
16.38
10.28
17.02
11.54
7.62
28.59
DnCNN
24.62
16.92
56.50
59.88
16.12
23.96
12.94
10.54
6.10
13.56
8.36
5.86
21.28
Appendix
Table 5: Ablation study on the effectiveness of different MAP components based on PGD attacker.
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
16.38
13.04
37.52
36.50
12.80
17.10
10.18
9.40
5.86
10.98
7.50
5.56
15.24
w/ DBO,w/o MAP
22.64
15.76
52.58
57.24
14.52
22.18
13.02
10.80
6.50
12.26
7.92
5.02
20.04
w/o DBO,w/ MAP
22.78
18.28
47.94
43.84
20.70
21.92
14.80
14.30
8.90
11.98
10.02
7.02
20.21
w/ DBO,w/ MAP
34.96
24.58
70.00
69.22
27.84
33.94
19.72
16.38
10.28
17.02
11.54
7.62
28.59
GHOST
18.34
14.44
41.56
41.72
13.60
19.78
10.98
10.00
6.02
12.10
7.14
5.90
16.80
Appendix
Table 6: Ablation study on the effectiveness of individual DBO and MAP components.
Figure 6: Hyperparameter sensitivity on ResNet-50 and VGG-19 surrogates. We evaluate the effects of the warm-up iteration T , lower-level step size β , and gap coefficient λ .
Untargeted Attack Scenario, ResNet-50 backbone, ASR (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
36.98
31.44
62.64
59.84
33.42
39.52
23.08
20.34
14.10
25.96
16.78
13.18
31.44
Ours
63.64
51.00
90.96
89.84
53.10
61.86
40.50
34.10
22.98
37.46
23.44
16.88
48.81
GHOST
39.58
32.42
67.82
64.70
34.50
42.10
24.38
21.74
14.34
26.64
17.38
12.84
33.20
Ours
66.22
53.82
92.64
91.62
56.46
65.62
41.76
35.72
24.20
41.26
26.08
18.48
51.16
Appendix
Table 7: Comparison of untargeted ASR (%) by implementing 3 base attackers under MABT, including VMI, DI, and SI. We integrate and compare MABT with different method combinations. The best results are highlighted in bold .
Untargeted Attack on ImageNet, VGGNet-19 backbone, ASR (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
12.84
9.54
25.50
32.44
13.06
13.04
7.96
7.46
4.18
9.52
5.10
3.08
11.98
Ours
29.74
23.50
52.56
58.34
37.58
32.70
15.38
12.92
7.94
19.84
10.28
6.28
25.59
GHOST
12.60
9.64
25.18
32.52
13.46
13.78
7.96
7.42
4.10
9.54
5.30
3.02
12.04
Ours
29.94
23.22
52.50
58.46
37.38
32.02
15.26
12.94
8.16
19.88
10.14
6.42
25.53
Appendix
Table 8: Comparison of untargeted ASR (%) using a VGG-19 surrogate. We implement 4 base attackers along with 3 ensembles attacks. The best results in each category are highlighted in bold .
Targeted Attack on ImageNet, ResNet-50 backbone, ASR (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
0.02
0.00
0.04
0.08
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.01
Ours
0.08
0.10
1.32
0.74
0.22
0.30
0.06
0.02
0.00
0.04
0.02
0.00
0.24
GHOST
0.02
0.00
0.14
0.08
0.00
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.02
Ours
0.32
0.22
2.94
1.30
0.50
0.50
0.06
0.00
0.00
0.14
0.02
0.00
0.50
Appendix
Table 9: Comparative results of targeted ASR against 12 victim architectures.
The University of Hong Kong, Hong Kong SAR, China · International School of Information Science & Engineering, Dalian University of Technology, Dalian, China · The Hong Kong Polytechnic University, Hong Kong SAR, China
College of Computer Science, Sichuan University, China · Department of Computer and Information Sciences, Northumbria University, UK · School of Artificial Intelligence, Sichuan University, China