One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs
Authors: Sen Nie, Jie Zhang, Zhongqi Wang, Shiguang Shan, Xilin Chen
Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, and the University of Chinese Academy of Sciences.
Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.
Figures & tables
Fig. 1: Black-box attack success rates (ASR) on 10 frontier MLLMs. All methods use same surrogate models (CLIP-B/16, CLIP-B/32, and CLIP-G/14). O-Attack (ours) achieves an average ASR of 77.1% , compared with 36.2% for M-Attack [ 1 ] and 44.0% for FOA-Attack [ 2 ] .
Fig. 2: An example of O-Attack transferring across models and prompts. A single adversarial image elicits responses aligned with the target market scene from six frontier commercial MLLMs under both image-description and question-answering prompts.
Fig. 3: Late-layer CKA analysis on surrogate CLIP models and victim MLLMs. Late-layer representations show strong intra-modal similarity, while cross-modal alignment and surrogate–victim similarity extend beyond the final layer, revealing a transferable high-level semantic space.
Fig. 4: Overview of O-Attack . (a) CKA trajectories guide the selection of late visual and textual anchor layers. (b) Progressive dropout sampling explores diverse semantic conditions within the anchored space. (c) Semantic consensus optimization promotes high mean target alignment and low variance across surrogates and augmented views. Here, fθ1 and gϕ1 denote a surrogate’s frozen visual and text encoders.
Model
AnyAttack
COA
M-Attack
FOA-Attack
M-Attack-V2
MPCAttack
O-Attack (ours)
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
10 frontier commercial MLLMs
GPT 5.4
2.3
0.071
1.8
0.026
29.1
0.325
35.0
0.381
46.5
0.485
43.5
0.439
77.2
0.650
Claude 4.6
1.4
0.044
0.9
0.014
42.8
0.452
50.2
0.495
76.4
0.643
76.7
0.627
81.6
0.693
Gemini 3.1
1.7
0.055
0.0
0.028
38.2
0.419
48.3
0.482
66.7
0.581
77.6
0.649
80.9
0.678
Grok 4.3
1.5
0.054
0.9
0.015
37.6
0.469
50.1
0.512
79.6
0.637
66.3
0.620
85.7
0.700
TABLE I: Black-box adversarial attack performance on 10 frontier commercial MLLMs and 14 widely used MLLMs.
ϵ
Method
GPT 5.4
Claude 4.6
Gemini 3.1
Grok 4.3
Qwen 3.5
Kimi K2.5
Average
Imperceptibility
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ℓ1norm↓
ℓ2norm↓
8/255
M-Attack
16.0
0.212
4.0
0.143
21.0
0.249
19.0
0.278
11.0
0.214
2.0
0.079
12.2
0.196
0.0207
0.0233
FOA-Attack
14.0
0.227
6.0
0.131
21.0
0.272
26.0
0.325
16.0
0.231
5.0
0.096
14.7
0.214
0.0208
0.0234
M-Attack-V2
32.0
0.337
23.0
0.320
27.0
0.371
48.0
0.506
49.0
0.503
20.0
0.264
33.2
0.383
0.0233
0.0254
MPCAttack
25.0
0.309
10.0
0.212
48.0
0.469
54.0
0.509
46.0
0.492
23.0
0.266
34.3
0.376
0.0235
0.0257
O-Attack (ours)
52.0
0.494
40.0
0.436
53.0
0.509
57.0
0.557
59.0
0.569
25.0
0.342
47.7
0.484
0.0224
0.0247
TABLE II: Performance under different perturbation budgets ϵ , with normalized ℓ1 and ℓ2 perturbation magnitudes.
Fig. 5: Attack success rate (ASR) and average similarity (AvgSim) under different levels of reasoning effort. The results compare O-Attack and MPCAttack on GPT-5.5 and Claude Opus 4.8, with reasoning effort ranging from no reasoning to extra-high reasoning.
Fig. 7: Qualitative safety-moderation failures on UnsafeBench. Adversarial images generated by O-Attack cause GPT-5.5 and Claude Opus 4.8 to classify unsafe content from four representative categories—illegal activity, shocking, sexual, and violence—as safe.
Method
GPT 5.4
Claude 4.6
Gemini 3.1
Grok 4.1
Qwen 3.5
Kimi K2.5
Avg. (6)
Acc* ( ↓ )
EM ( ↓ )
Acc* ( ↓ )
EM ( ↓ )
Acc* ( ↓ )
EM ( ↓ )
Acc* ( ↓ )
EM ( ↓ )
Acc* ( ↓ )
EM ( ↓ )
Acc* ( ↓ )
EM ( ↓ )
Acc* ( ↓ )
EM ( ↓ )
Clean
69.6
55.1
45.2
38.8
67.6
55.1
45.2
35.7
78.8
68.4
57.5
49.0
60.7
50.3
M-Attack
35.9
25.5
16.1
14.3
42.3
34.4
22.4
17.3
36.7
29.6
22.2
17.7
29.3
23.1
FOA-Attack
36.8
25.9
14.2
11.9
42.4
34.4
24.1
18.0
35.9
27.6
23.6
18.4
29.5
22.7
M-Attack-V2
34.2
25.9
10.8
8.5
34.4
27.6
21.4
15.3
31.0
22.8
13.1
9.6
24.1
18.3
MPCAttack
31.7
23.8
10.9
8.2
28.7
22.8
19.2
13.6
26.4
21.1
15.0
12.2
22.0
17.0
TABLE IV: VQA v2 evaluation under different attacks on six frontier commercial MLLMs. Lower Acc* and EM indicate stronger attacks.
Fig. 8: Qualitative comparison of adversarial examples and their perturbation textures. The left panel shows source images and the adversarial examples generated by different attacks, while the right panel shows target images and the corresponding amplified perturbations. Compared with the baselines, O-Attack produces cleaner and more coherent target-related semantic structures, including under the smaller ϵ=12/255 budget.
Fig. 9: Qualitative example illustrating the effect of reasoning effort shown in Fig. 5 . Given the same adversarial image generated by O-Attack , the attack on GPT-5.5 fails under direct answering without reasoning but succeeds when high-effort reasoning is enabled, with the model producing a response semantically aligned with the target image.
Method
CSA
PSS
SCO
GPT 5.4
Claude 4.6
Gemini 3.1
Grok 4.3
Qwen 3.5
Kimi K2.5
Avg.(6)
Space
Text
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
M-Attack
×
×
×
×
29.1
0.325
42.8
0.452
38.2
0.419
37.6
0.469
40.9
0.439
49.5
0.504
39.7
0.435
O-Attack
×
×
×
✓
39.0
0.442
67.0
0.584
53.0
0.534
65.0
0.574
62.0
0.583
73.0
0.625
59.8
0.557
O-Attack
×
✓
×
✓
41.0
0.451
68.0
0.585
52.0
0.532
65.0
0.589
76.0
0.624
62.0
0.553
60.7
0.556
O-Attack
✓
×
✓
✓
52.0
0.520
70.0
0.625
68.0
0.593
59.0
0.579
74.0
0.617
62.0
0.529
64.2
0.577
O-Attack
✓
✓
×
✓
59.0
0.541
77.0
0.675
66.0
0.622
72.0
0.632
76.0
0.661
77.0
0.650
71.2
0.630
TABLE V: Component ablation of O-Attack on six frontier commercial MLLMs. CSA: cross-modal semantic space anchoring (Space: visual encoders; Text: textual encoders). PSS: progressive semantic space sampling; SCO: semantic consensus optimization.
Method
Surrogate CLIP Models
GPT 5.4
Claude 4.6
Gemini 3.1
Grok 4.3
Qwen 3.5
Kimi K2.5
Avg.(6)
G/14
B/32
B/16
B/32(l)
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
M-Attack-V2
✓
✓
✓
✓
46.5
0.485
76.4
0.643
66.7
0.581
79.6
0.637
77.9
0.664
67.8
0.628
69.2
0.606
FOA-Attack
✓
✓
✓
×
35.0
0.381
50.2
0.495
48.3
0.482
50.1
0.512
41.4
0.466
52.2
0.548
46.2
0.481
O-Attack
✓
×
×
×
48.6
0.434
56.5
0.500
64.0
0.546
71.3
0.580
61.0
0.510
58.5
0.512
60.0
0.514
O-Attack
✓
×
✓
×
67.6
0.571
66.9
0.584
74.8
0.633
76.6
0.618
67.4
0.549
68.5
0.560
70.3
0.586
O-Attack
✓
✓
×
×
67.6
0.587
82.7
0.676
68.5
0.589
78.4
0.640
74.6
0.629
77.4
0.634
74.9
0.626
TABLE VI: Ablation study of surrogate CLIP backbones on six frontier commercial MLLMs. B/32(l) denotes the LAION version of CLIP-B/32.
Method
Settings
#Num.
Time
GPT 5.4
Claude 4.6
Gemini 3.1
Grok 4.3
Qwen 3.5
Kimi K2.5
Avg. (6)
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
MPCAttack
–
6
∼ 79s
43.5
0.439
76.7
0.627
77.6
0.649
66.3
0.620
77.7
0.663
76.2
0.632
69.7
0.605
M-Attack-V2
K=10
4
∼ 283s
46.5
0.485
76.4
0.643
66.7
0.581
79.6
0.637
77.9
0.664
67.8
0.628
69.2
0.606
FOA-Attack
C=[3,5,8,10]
3
∼ 284s
35.0
0.381
50.2
0.495
48.3
0.482
50.1
0.512
41.4
0.466
52.2
0.548
46.2
0.481
O-Attack
K=3
3
∼ 86s
62.0
0.574
74.0
0.640
74.0
0.643
79.0
0.666
74.0
0.648
67.0
0.611
71.7
0.630
O-Attack
K=5
3
∼ 145s
64.0
0.594
71.0
0.630
79.0
0.661
82.0
0.692
80.0
0.672
75.0
0.639
75.2
0.648
TABLE VII: Ablation study of source augmentation crop number K on six frontier commercial MLLMs. #Num. denotes the number of surrogate models used by each method, and Time reports the end-to-end testing time per image on a single RTX 4090 GPU.
Dropout cap
GPT 5.4
Claude 4.6
Gemini 3.1
Grok 4.3
Qwen 3.5
Kimi K2.5
Avg. (6)
G/14
B/16, B/32
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
ASR
AvgSim
0.05→0.20
0.0→0.15
66.0
0.613
76.0
0.660
78.0
0.671
76.0
0.672
74.0
0.654
66.0
0.612
72.7
0.647
0.05→0.10
0.0→0.05
68.0
0.584
77.0
0.681
75.0
0.667
79.0
0.679
75.0
0.653
70.0
0.625
74.0
0.648
0.15→0.10
0.10→0.05
73.0
0.619
76.0
0.648
74.0
0.640
77.0
0.660
76.0
0.663
79.0
0.657
75.8
0.648
0.10→0.15
0.05→0.10
77.2
0.650
81.6
0.693
80.9
0.678
85.7
0.700
83.2
0.668
79.7
0.652
81.4
0.673
TABLE VIII: Ablation study of progressive dropout sampling in O-Attack on six frontier commercial MLLMs. Arrows indicate the initial and final dropout probability caps for each surrogate group.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 10: Targeted transfer from a butterfly image to a scene of a person resting in bed with cats. Across the displayed models, responses to the adversarial image emphasize the person, bedding, and cats in the target scene. The source, adversarial, and target images appear above the response columns.
Fig. 11: Targeted transfer from a camera image to an urban street scene. Responses emphasize buses, pedestrians, storefronts, and signs associated with the target. Some descriptions also refer to composite or distorted visual content, revealing differences in how models interpret the same adversarial input.
Fig. 12: Targeted transfer toward a marina scene. Most displayed responses describe boats, docks, and water, with variation in colors and spatial details. The Gemma-3 response instead describes a street scene, illustrating that target alignment is not uniform across all model responses.
Fig. 13: Targeted transfer from an ostrich image to a scene of people riding an elephant. The displayed models consistently mention an elephant and its riders, while their descriptions of water, vegetation, and nearby structures vary.
Fig. 14: Targeted transfer from a toy fire-engine image to geese swimming on water. Responses identify ducks or geese but often combine them with colorful docks or boat-like structures. These mixed descriptions retain target-related birds alongside details associated with the source image.
Fig. 15: Targeted transfer from a motorboat image to sheep in a grassy field. The responses broadly agree on sheep but often place them in a vehicle or behind an opening. References to a Yamaha vehicle or boat illustrate the coexistence of target-related animal content and source-related details.
Fig. 16: Attack success rate (%) as the semantic similarity threshold increases from 0.1 to 0.9 . Each model panel compares O-Attack (red) with FOA-Attack (blue); the final panel reports the mean over the displayed models. Higher thresholds require closer agreement with the target semantics.
Model
Access
Model identifier
LLaVA-v1.6
HF
llava-hf/llava-v1.6-mistral-7b-hf
Qwen3-VL
HF
Qwen/Qwen3-VL-8B-Instruct
Gemma-3
HF
google/gemma-3-12b-it
InternVL3.5
HF
OpenGVLab/InternVL3_5-8B-Instruct
MiniCPM-V 4.5
HF
openbmb/MiniCPM-V-4_5
DeepSeek-VL2
HF
deepseek-ai/deepseek-vl2-tiny
Appendix
TABLE IX: Victim model identifiers and access modes. HF denotes locally evaluated public weights; API denotes access through OpenRouter. GPT-5.5 and Claude Opus 4.8 are used in the reasoning-effort study.
Surrogate
Model identifier
Used by
CLIP-B/32
openai/clip-vit-base-patch32
All seven attacks
CLIP-B/16
openai/clip-vit-base-patch16
Five attacks (see caption)
CLIP-G/14
laion/CLIP-ViT-g-14-laion2B-s12B-b42K
Five attacks (see caption)
CLIP-B/32 (LAION)
laion/CLIP-ViT-B-32-laion2B-s34B-b79K
M-Attack-V2
EVA-02-L/14
timm/eva02_large_patch14_448.mim_m38m_ft_in1k
AnyAttack
ViT-B/16
torchvision/vit_b_16-imagenet1k
AnyAttack
Appendix
TABLE X: Surrogate checkpoints and the attacks that use them. The shared CLIP-B/16 and CLIP-G/14 models are used by M-Attack, FOA-Attack, M-Attack-V2, MPCAttack, and O-Attack.
Adversarial perturbations can mislead Multimodal Large Language Models (MLLMs) recognize a benign image as a specific target object, posing serious risks in safety-critical scenarios such as autonomous driving and medical diagnosis. This makes transfer-based targeted attacks crucial for understanding and improving black-box MLLM robustness. Existing transfer-based targeted attack methods typically rely on the final global features of the surrogate encoder and anchor optimization to original-resolution target crops, leading to their limited transferability and robustness. To address these challenges, we propose Progressive Resolution Processing and Adaptive Feature Alignment (PRAF-Attack), a targeted transfer-based attack framework that integrates multi-scale global semantic guidance with robust intermediate-layer local alignment. Unlike prior methods that align only the surrogate encoder's final layer, we design an adaptive feature alignment strategy that leverages intermediate representations to enhance transferability. Specifically, we introduce an adaptive intermediate layer selection mechanism to identify transferable hierarchical features across surrogate ensembles via gradient consistency, along with an adaptive patch-level optimization strategy that preserves highly correlated local regions through efficient patch filtering. To overcome the reliance on fixed original-resolution target crops, we propose a progressive resolution processing strategy that gradually refines optimization from coarse to fine, enabling the attack to better exploit target information at multiple scales and achieve stronger transferability. We evaluate PRAF-Attack on a diverse suite of black-box MLLMs, including six open-source models and six closed-source commercial APIs. Compared with seven state-of-the-art targeted attack baselines, the proposed PRAF-Attack consistently achieves superior transferability.
Haobo Wang, Xiaorong Ma, Weiqi Luo +2
Sun Yat-sen University, China · Nanyang Technological University, Singapore · Shenzhen MSU-BIT University, China
Multimodal large language models (MLLMs) remain vulnerable to transfer-based targeted attacks, where perturbations optimized on open-source surrogate encoders can generalize to closed-source MLLMs. A key challenge for improving adversarial transferability is to effectively capture the intrinsic visual focus shared across different models, such that perturbations align with transferable semantic cues rather than surrogate-specific behaviors. However, existing methods suffer from spatial-domain feature redundancy and surrogate-specific gradient signals, thereby hindering cross-model transferability. In this paper, we propose FRA-Attack, which addresses both challenges from a unified frequency-domain regularization perspective. For feature alignment, a high-pass DCT objective on patch features suppresses redundant global structures and concentrates the loss on the high-frequency band that carries the MLLMs' intrinsic visual focus. For gradient optimization, we introduce Frequency-domain Gradient Regularization (FGR), a \textit{model-agnostic} low-pass regularizer that modulates the surrogate gradient using only the geometric frequency coordinate, \textit{i.e.}, no surrogate-derived statistic is involved, so that FGR is model-agnostic by construction, removing surrogate-specific high-frequency artifacts while preserving transferable low-frequency directions. Together, the two components form a unified frequency-domain treatment of transferability. Extensive experiments on 15 flagship MLLMs across 7 vendors show that FRA-Attack achieves superior cross-model transferability, particularly with state-of-the-art performance on GPT-5.4, Claude-Opus-4.6 and Gemini-3-flash.
Leitao Yuan, Qinghua Mao, Daizong Liu +5
Shanghai Artificial Intelligence Laboratory · Zhejiang University · Shanghai Jiao Tong University +3
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, traditional adversarial attacks are typically limited to single, predefined objectives, tightly coupling each attack to a specific model or task, which restricts their scalability and flexibility in real-world scenarios. In this work, we present DarkLLM, a novel attack framework that trains an LLM to translate natural-language attack instructions into latent attack vectors, which are then decoded into visual adversarial perturbations. By leveraging natural-language instruction tuning, DarkLLM not only unifies targeted, untargeted, segmentation, and multi-model attacks within a single framework, but also achieves flexible and controllable adversarial generation, enabling each instruction to produce a perturbation that induces desired behaviors across heterogeneous models. Through extensive experiments across 4 tasks, 13 datasets, and 15 models, we demonstrate that DarkLLM with only 1B parameters can follow attacker instructions and generate highly effective attacks against CLIP, SAM, and frontier LLMs, revealing a systemic vulnerability in modern foundation models.
Ye Sun, Xin Wang, Jiaming Zhang +7
Fudan University · Nanyang Technological University · Tongji University