Anchoring Adversarial Trajectories to Data Manifolds: A Bilevel Transfer Optimization Framework
Authors: Yaohua Liu, Yifan Guo, Jiaxin Gao
Organizations: The University of Hong Kong Hong Kong, China · International School of Information Science & Engineering Dalian University of Technology, Dalian, China · The Hong Kong Polytechnic University Hong Kong, China
A key bottleneck in adversarial transfer is a trajectory-level geometric disconnect: ambient gradients often drift away from the intrinsic data manifold, causing surrogate-specific overfitting. To rectify this, we propose Manifold Anchored Bilevel Transfer (MABT), a unified framework that anchors adversarial trajectories to the shared semantic subspace. MABT introduces a relaxed manifold-anchoring operator as a semantic rectifier to suppress off-manifold noise. With this constraint, we cast transfer attack generation as a distributional bilevel optimization problem that learns a geometry-aligned initialization by minimizing expected transfer risk under a surrogate uncertainty distribution. We further develop a Hessian-free solver with linear-time complexity to handle the resulting hierarchy. Experiments demonstrate improved transferability for 10 baseline attackers across 28 attack configurations, diverse victim architectures, and defense mechanisms.
Figures & tables
Figure 1: Motivation. (a) Manifold hypothesis : Natural images concentrate near a low-dimensional manifold M capturing intrinsic semantics, where decision boundaries of diverse models tend to partially align. (b) Off-manifold divergence : Standard iterative attacks inherently drift into high-dimensional noise away from M ; in this region, boundaries are less consistent, so the trajectory may cross the surrogate boundary but miss the victim boundary . (c) Manifold anchoring : MABT uses a manifold-anchoring projector with a geometry-aligned initialization to suppress off-manifold deviations while preserving task-relevant semantics to cross shared boundaries .
Figure 2: Feature space divergence. We quantify victim-side feature shift during the attack process as 1−cos(fclean,fadv) , where x1→x10 denotes the adversarial trajectory over 10 attack iterations. Compared with GAA [ 16 ] +AWT [ 6 ] , MABT induces substantially larger semantic shifts on victim models, with relative gains of 99.8% on Transformers and 32.9% on CNNs.
Untargeted Attack Scenario, ResNet-50 backbone, Average Success Rate (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
16.38
13.04
37.52
36.50
12.80
17.10
10.18
9.40
5.86
10.98
7.50
5.56
15.24
Ours
34.96
24.58
70.00
69.22
27.84
33.94
19.72
16.38
10.28
17.02
11.54
7.62
28.59
GHOST
18.34
14.44
41.56
41.72
13.60
19.78
10.98
10.00
6.02
12.10
7.14
5.90
16.80
Ours
37.38
27.52
74.66
72.58
30.84
36.72
21.18
17.66
10.50
18.24
11.92
7.72
30.58
Table 1: ASR (%) of MABT across 16 attack configurations on the ResNet-50 surrogate. Results cover 4 base attackers combined with ensemble/model-related strategies. Best results are in bold .
Method
Average ASR (%) ↑ , ResNet-50
Avg.
CNN
CNN Ensemble
Transformer
PGD
N/A
22.22
8.48
8.01
15.24
BETAK
30.87
11.75
10.47
20.99
Ours
43.42
15.46
12.06
28.59
MI
N/A
33.74
13.15
12.99
23.41
BETAK
43.86
19.74
15.97
30.86
Table 2: Comparison with initialization-based BETAK which uses auxiliary pseudo-victims (Inc-v3).
Defense Mechanism, ResNet-50 backbone, Average Success Rate (%) ↑
Method
HGD
JPEG
RS
R&P
NIPS-r3
Avg.
Method
HGD
JPEG
RS
R&P
NIPS-r3
Avg.
N/A
16.94
11.60
7.96
12.12
14.28
12.58
N/A
27.12
17.54
9.44
13.40
23.10
18.12
Ours
36.16
21.38
10.42
13.90
29.32
22.24
Ours
59.82
39.68
14.80
18.28
53.68
37.25
GHOST
19.34
11.84
8.24
12.70
15.76
13.58
GHOST
30.94
19.02
9.40
13.92
24.80
19.62
Ours
38.86
21.94
10.40
14.26
31.44
23.38
Ours
63.02
41.56
15.08
19.40
56.08
39.03
MBA
22.22
13.22
9.22
12.84
19.80
15.46
MBA
40.10
25.82
11.94
15.52
35.80
25.84
Table 3: ASR (%) results against five representative defenses using the ResNet-50 surrogate.
Figure 3: Adversarial loss landscape . Surfaces are plotted along the gradient direction and a random orthogonal direction . Unlike baselines that suffer from landscape collapse on unknown victims, MABT preserves higher victim-side loss basins with relatively flatter surrounding curvature across architectures, which indicates that the learned trajectory better retains transfer-relevant directions.
Figure 4: Convergence and efficiency analysis. Left & Mid: Increasing T improves geometry-aligned initialization to mitigate surrogate overfitting, thereby ensuring sustained victim loss growth during the subsequent attack. Right: Resize provides a favorable low-cost MAP realization, while generative priors achieve stronger anchoring at higher computational cost.
Figure 5: Feature shift consistency. We plot surrogate-side (ResNet-50) feature shift against the average victim-side shift. Shaded regions denote the standard deviation over four heterogeneous victims (Inc-v3, DenseNet, DeiT-s and ConViT-b).
Method
Average ASR (%) ↑ , ResNet-50
Avg.
CNN
CNN Ensemble
Transformer
N/A
22.22
8.48
8.01
15.24
w/ DBO,w/o MAP
30.82
10.11
8.40
20.04
w/o DBO,w/ MAP
29.24
12.67
9.67
20.21
w/ DBO,w/ MAP
43.42
15.46
12.06
28.59
AWT
25.99
11.57
9.05
18.15
Table 4: Ablation results by implementing DBO and MAP based on PGD and PGD+AWT.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
16.38
13.04
37.52
36.50
12.80
17.10
10.18
9.40
5.86
10.98
7.50
5.56
15.24
FFT
32.42
23.00
64.06
64.22
24.34
28.66
17.96
15.18
8.98
14.44
10.90
7.12
25.94
Bit-Depth
34.68
25.50
69.78
70.78
25.38
33.76
19.08
15.44
9.50
20.28
11.08
7.50
28.56
Resize
34.96
24.58
70.00
69.22
27.84
33.94
19.72
16.38
10.28
17.02
11.54
7.62
28.59
DnCNN
24.62
16.92
56.50
59.88
16.12
23.96
12.94
10.54
6.10
13.56
8.36
5.86
21.28
Appendix
Table 5: Ablation study on the effectiveness of different MAP components based on PGD attacker.
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
16.38
13.04
37.52
36.50
12.80
17.10
10.18
9.40
5.86
10.98
7.50
5.56
15.24
w/ DBO,w/o MAP
22.64
15.76
52.58
57.24
14.52
22.18
13.02
10.80
6.50
12.26
7.92
5.02
20.04
w/o DBO,w/ MAP
22.78
18.28
47.94
43.84
20.70
21.92
14.80
14.30
8.90
11.98
10.02
7.02
20.21
w/ DBO,w/ MAP
34.96
24.58
70.00
69.22
27.84
33.94
19.72
16.38
10.28
17.02
11.54
7.62
28.59
GHOST
18.34
14.44
41.56
41.72
13.60
19.78
10.98
10.00
6.02
12.10
7.14
5.90
16.80
Appendix
Table 6: Ablation study on the effectiveness of individual DBO and MAP components.
Figure 6: Hyperparameter sensitivity on ResNet-50 and VGG-19 surrogates. We evaluate the effects of the warm-up iteration T , lower-level step size β , and gap coefficient λ .
Untargeted Attack Scenario, ResNet-50 backbone, ASR (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
36.98
31.44
62.64
59.84
33.42
39.52
23.08
20.34
14.10
25.96
16.78
13.18
31.44
Ours
63.64
51.00
90.96
89.84
53.10
61.86
40.50
34.10
22.98
37.46
23.44
16.88
48.81
GHOST
39.58
32.42
67.82
64.70
34.50
42.10
24.38
21.74
14.34
26.64
17.38
12.84
33.20
Ours
66.22
53.82
92.64
91.62
56.46
65.62
41.76
35.72
24.20
41.26
26.08
18.48
51.16
Appendix
Table 7: Comparison of untargeted ASR (%) by implementing 3 base attackers under MABT, including VMI, DI, and SI. We integrate and compare MABT with different method combinations. The best results are highlighted in bold .
Untargeted Attack on ImageNet, VGGNet-19 backbone, ASR (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
12.84
9.54
25.50
32.44
13.06
13.04
7.96
7.46
4.18
9.52
5.10
3.08
11.98
Ours
29.74
23.50
52.56
58.34
37.58
32.70
15.38
12.92
7.94
19.84
10.28
6.28
25.59
GHOST
12.60
9.64
25.18
32.52
13.46
13.78
7.96
7.42
4.10
9.54
5.30
3.02
12.04
Ours
29.94
23.22
52.50
58.46
37.38
32.02
15.26
12.94
8.16
19.88
10.14
6.42
25.53
Appendix
Table 8: Comparison of untargeted ASR (%) using a VGG-19 surrogate. We implement 4 base attackers along with 3 ensembles attacks. The best results in each category are highlighted in bold .
Targeted Attack on ImageNet, ResNet-50 backbone, ASR (%) ↑
Method
CNN
CNN Ensemble
Transformer
Avg.
Inc-v3
Inc-Res-v2
DenseNet
MobileNet
PNASNet
SENet
Inc-v3ens3
Inc-v3ens4
IncRes-v2ens
Visformer-s
DeiT-s
ConViT-b
N/A
0.02
0.00
0.04
0.08
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.01
Ours
0.08
0.10
1.32
0.74
0.22
0.30
0.06
0.02
0.00
0.04
0.02
0.00
0.24
GHOST
0.02
0.00
0.14
0.08
0.00
0.04
0.00
0.00
0.00
0.00
0.00
0.00
0.02
Ours
0.32
0.22
2.94
1.30
0.50
0.50
0.06
0.00
0.00
0.14
0.02
0.00
0.50
Appendix
Table 9: Comparative results of targeted ASR against 12 victim architectures.
Transfer-based adversarial attacks craft adversarial examples using surrogate models to mislead black-box victim models. Beyond perturbation generation, transferability is fundamentally governed by the coupling of initialization, surrogate adaptation, and gradient dynamics. We revisit this challenge from a bilevel-minimax perspective and propose BMAT (Bilevel-Minimax Adversarial Transfer). The bilevel formulation captures the dependency between initialization and perturbation, while the inner minimax problem promotes surrogate robustness for cross-architecture generalization. Algorithmically, we develop an integrated bottom-up solver that combines a Soft Weight Modulator and an Implicit Gradient Approximator to enable ternary coupling among initialization, surrogate adaptation, and perturbation optimization. We further provide theoretical insights into the optimization dynamics of the proposed bilevel-minimax framework. Extensive experiments on classification and segmentation benchmarks show that BMAT outperforms more than 10 strong baselines across more than 30 victim models, improving both intra- and cross-architecture transfer and yielding up to a 2x reduction in mIoU. Code is available at https://github.com/callous-youth/BMAT.
Yaohua Liu, Yifan Guo, Jiaxin Gao
The University of Hong Kong, Hong Kong SAR, China · International School of Information Science & Engineering, Dalian University of Technology, Dalian, China · The Hong Kong Polytechnic University, Hong Kong SAR, China
Transfer-based adversarial attacks often transfer poorly across heterogeneous architectures because CNNs favor local textures while Vision Transformers (ViTs) rely on global shapes. We propose Season, a spectrum-aware orthogonal gradient refinement framework for L-infinity transfer attacks against black-box target models on ImageNet, using a white-box surrogate. Season decomposes each update into a low-frequency branch capturing structural cues and a high-frequency branch capturing textures. A low-saliency guidance scheme reallocates high-frequency energy to background regions, preserving foreground structures that ViTs depend on. An orthogonal projection then forces the textural update to lie in the orthogonal complement of the structural direction, mitigating feature interference. As a training-free plug-and-play wrapper, Season enhances eight gradient-stabilization and input-enhancement attacks without modifying their cores. Across eight CNN, ViT, and MLP targets, Season improves transfer success rate by 6.6 percentage points on average and up to 16.0 points over strong baselines under a unified protocol.
Tianyi Wang, Zhenghao Gao, Shengjie Xu
Tongji University Shanghai, China · Huazhong University of Science and Technology · Wuhan LightRead Intelligent Technology Co., Ltd. Wuhan, China
Transfer-based adversarial attacks rely on surrogate models to craft perturbations, yet often overfit the surrogate's decision boundary. To address this problem, we propose Inverse Knowledge Distillation (IKD), a simple and attack-agnostic mechanism that maximizes the prediction-distribution discrepancy between benign and adversarial samples on the surrogate model. IKD uses a CE/KL-equivalent soft-label objective to push adversarial predictions away from a fixed benign prediction anchor and enrich the attack with Fisher-sensitive surrogate directions. We prove that, under a matched fixed-anchor implementation, soft-label cross-entropy and KL divergence differ only by a constant entropy term and therefore induce identical gradients, Hessians, and adversarial optimization trajectories. Our information-geometric analysis further derives a quantitative lower bound on dominant Fisher-subspace overlap between surrogate and target models from local same-task stability and a Fisher eigengap, and establishes a sufficient target-margin crossing condition under oriented gradient coherence and target smoothness. This analysis connects IKD's surrogate Fisher sensitivity to cross-model transfer. In contrast, mean squared error uses a different Euclidean pullback in output probability space. IKD integrates seamlessly with standard gradient-based attacks without modifying their optimization pipelines. Extensive ImageNet experiments demonstrate consistent black-box gains across CNN, ViT, and defended models, while ablations confirm CE and KL equivalence and the pronounced disadvantage of MSE. These results establish IKD as an effective and lightweight component for improving adversarial transferability. Code is available at https://github.com/ImmortalTing/IKD.
Wenyuan Wu, Yuan Sun, Yingke Chen +4
College of Computer Science, Sichuan University, China · Department of Computer and Information Sciences, Northumbria University, UK · School of Artificial Intelligence, Sichuan University, China