Large vision-language models (LVLMs) have demonstrated their incredible capability in visual question answering. However, this rich visual interaction also makes LVLMs vulnerable to adversarial examples. In this paper, we formulate a novel and practical targeted attack scenario that the adversary knows only the vision encoder of the victim LVLM, without the knowledge of its prompts and its underlying large language model. This practical setting poses challenges to the cross-prompt and cross-model transferability of targeted adversarial attack, which aims to confuse the LVLM to output a response that is semantically similar to the attacker's chosen target text. To this end, we propose an instruction-tuned targeted attack (dubbed InstructTA) to deliver the targeted adversarial attack on LVLMs with high transferability. Initially, we utilize a public text-to-image generative model to reverse the target response into a target image, and employ GPT-4 to infer a reasonable instruction p′ from the target response. We then form a local surrogate model (sharing the same vision encoder with the victim LVLM) to extract instruction-aware features of an adversarial image example and the target image, and minimize the distance between these two features to optimize the adversarial example. To further improve the transferability with instruction tuning, we augment the instruction p′ with instructions paraphrased from GPT-4. Extensive experiments on 6 victim LVLMs demonstrate the superiority of our proposed method in targeted attack performance and transferability. In particular, InstructTA achieves an attack success rate of 51.9% on BLIP-2, outperforming the strongest baseline by 10.5%, and consistently yields the highest attack success rates across all evaluated models. The code is available at https://github.com/xunguangwang/InstructTA.
Figures & tables
Figure 1 : The framework of our instruction-tuned targeted attack ( InstructTA ). The framework comprises three key modules: (1) a text-to-image model hξ , which generates a target image xt from the target text yt ; (2) GPT-4, which serves dual roles of inferring a plausible instruction p′ for yt and augmenting p′ into a set of semantically consistent instruction variants; and (3) a surrogate model M (sharing the same vision encoder as the victim LVLM), which acts as the substitute model for crafting adversarial examples. Given a target text yt , we first transform it into the target image xt with hξ . Simultaneously, GPT-4 infers a reasonable instruction p′ . Upon providing the augmented instructions pi′ and pj′ rephrased from p′ using GPT-4, the surrogate model M extracts instruction-aware features of xt and the AE x′ , respectively. Finally, we minimize the L2 distance between these two features to optimize x′ .
Figure 2 : The architecture of LVLM.
Figure 3 : An example of rephrasing an instruction. Given the target response yt , GPT-4 infers an instruction, i.e. , “ Can you describe what’s in the picture you’re looking at related to food? ”. The real instruction assigned to this target text is “What are the essential components depicted in this image?”
LVLM
Attack method
Text encoder (pretrained) for evaluation
NoS ( ↑ )
RN50
RN101
ViT-B/16
ViT-B/32
ViT-L/14
Ensemble
BLIP-2
Clean image
0.281
0.395
0.323
0.333
0.184
0.303
1
MF-tt Zhao et al. (2023)
0.284
0.400
0.326
0.337
0.185
0.306
0
MF-it Zhao et al. (2023)
0.533
0.634
0.575
0.581
0.469
0.559
401
MF-ii Zhao et al. (2023)
0.536
0.634
0.575
0.583
0.471
0.560
414
M-Attack Li et al. (2026)
0.324
0.438
0.366
0.375
0.228
0.346
17
Table 1 : Targeted attacks against victim LVLMs. Bold entries denote the best performance among all compared methods.
Figure 4 : Visualization examples of various targeted attack methods on InstructBLIP. “What are the essential components depicted in this image?” is a real instruction.
Figure 5 : Failure cases of InstructTA on LLaVA-1.5. The first column is the target text, and the second column is the target image generated by Stable Diffusion. The other columns are the generated responses of AEs constructed by MF-tt, MF-it, MF-ii, and InstructTA , respectively. The instruction is “What are the essential components depicted in this image?”.
Attack method
CLIP-score ( ↑ )
ASR ( ↑ )
InstructTA
0.588
0.519
InstructTA -woMF
0.574
0.472
InstructTA -woMFG
0.521
0.324
Table 2 : Ablation study of the proposed method on BLIP-2. CLIP-score Zhao et al. (2023) is computed by ensemble CLIP text encoders. Attack success rate (ASR) is the ratio between the No. of attack successes and the total No. of samples.
Figure 6 : To explore the impact of varying ϵ values within the InstructTA , we conducted experiments aiming to achieve different levels of perturbed images on BLIP-2, i.e. , referred to as the AE x′ . Our findings indicate a degradation in the visual quality of x′ , as quantified by the LPIPS Zhang et al. (2018) distance between the original image x and the adversarial image x′ . Simultaneously, the effectiveness of targeted response generation reaches a saturation point. Consequently, it is crucial to establish an appropriate perturbation budget, such as ϵ=8 , to effectively balance the image quality and the targeted attack performance.
Parameter n
Text encoder (pretrained) for evaluation
ASR ( ↑ )
RN50
RN101
ViT-B/16
ViT-B/32
ViT-L/14
Ensemble
1
0.560
0.658
0.600
0.606
0.498
0.584
0.504
5
0.560
0.657
0.600
0.606
0.498
0.584
0.504
10
0.562
0.660
0.604
0.611
0.502
0.588
0.519
50
0.562
0.659
0.601
0.607
0.499
0.586
0.529
Table 4 : Targeted attack performance for different n .
Instruction
Rephrased instruction pi/j′
Real instruction p
Inferred instruction p′
0.943
0.829
Real instruction p
0.828
1.000
Table 5 : Average CLIP-score ( ↑ ) between different versions of instructions. Each type of instruction has 1,000 samples.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
LVLM
Attack method
Text encoder (pretrained) for evaluation
NoS ( ↑ )
RN50
RN101
ViT-B/16
ViT-B/32
ViT-L/14
Ensemble
BLIP-2
Clean image
0.275
0.397
0.319
0.328
0.177
0.299
2
MF-tt
0.277
0.400
0.320
0.331
0.177
0.301
3
MF-it
0.529
0.634
0.574
0.578
0.465
0.556
419
MF-ii
0.529
0.634
0.570
0.578
0.463
0.555
419
InstructTA
0.556
0.658
0.597
0.604
0.493
0.582
561
Appendix
Table A6 : Targeted attack performance with shuffled instructions.
Targeted adversarial attacks on Large Vision-Language Models (LVLMs) test whether small image perturbations can steer model responses toward attacker-specified content. Under the standard L-infinity constraint, targeted attacks become a regional perturbation budget allocation problem: attack success depends not only on the perturbation objective, but also on which regions receive updates and in what order. Existing localized attacks improve over global perturbations but rely on stochastic spatial sampling, often updating weakly influential regions. We address this limitation through an attention-based analysis showing that cross-modal attention identifies adversarially sensitive regions and that perturbing high-attention hotspots induces predictable redistribution toward subsequent salient regions. These findings motivate attention-guided region sequencing, which begins from dominant hotspots and progressively moves the update support toward next-salient regions. Based on these principles, we propose Stage-wise Attention-Guided Attack (SAGA), a black-box region-sequencing framework that uses a fixed attention map from an open-source LVLM to guide perturbation updates without accessing target-model parameters, gradients, or attention maps. Across ten closed-source and open-source LVLMs, SAGA achieves state-of-the-art attack success rates and the best overall imperceptibility. The source code is available at https://github.com/jaehyun-kwak/SAGA.
While vision and multimodal foundation models underpin critical tasks from perception to complex reasoning, they remain highly vulnerable to adversarial attacks. However, traditional adversarial attacks are typically limited to single, predefined objectives, tightly coupling each attack to a specific model or task, which restricts their scalability and flexibility in real-world scenarios. In this work, we present DarkLLM, a novel attack framework that trains an LLM to translate natural-language attack instructions into latent attack vectors, which are then decoded into visual adversarial perturbations. By leveraging natural-language instruction tuning, DarkLLM not only unifies targeted, untargeted, segmentation, and multi-model attacks within a single framework, but also achieves flexible and controllable adversarial generation, enabling each instruction to produce a perturbation that induces desired behaviors across heterogeneous models. Through extensive experiments across 4 tasks, 13 datasets, and 15 models, we demonstrate that DarkLLM with only 1B parameters can follow attacker instructions and generate highly effective attacks against CLIP, SAM, and frontier LLMs, revealing a systemic vulnerability in modern foundation models.
Ye Sun, Xin Wang, Jiaming Zhang +7
Fudan University · Nanyang Technological University · Tongji University
Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the same perturbed input is paired with different textual queries. This paper studies cross-query response manipulation, where a single adversarial example is expected to remain effective across diverse user queries. We first analyze the limitations of existing attacks and find that successful transfer is closely associated with preserving an image-dominant attention pattern during response generation. Motivated by the observation, we propose \textbf{Attention Hijacking}, a novel adversarial attack that explicitly steers internal attention distributions toward a persistent image-dominant pattern. By amplifying the influence of visual tokens on target response tokens while suppressing the competing influence of textual tokens, our method reduces the dependence of the manipulated output on the specific wording of the query. Extensive experiments on widely used VLMs show that Attention Hijacking substantially improves cross-query transferability across diverse target responses and unseen queries. The method also extends effectively to multiple attack scenarios, offering new insights into the role of attention stability in transferable response manipulation for VLMs.
Zhiqiang Wang, Dongrui Liu, Yan Li +4
Hong Kong University of Science and Technology · Shanghai Jiao Tong University · Beihang University