Large vision-language models (LVLMs) have demonstrated their incredible capability in visual question answering. However, this rich visual interaction also makes LVLMs vulnerable to adversarial examples. In this paper, we formulate a novel and practical targeted attack scenario that the adversary knows only the vision encoder of the victim LVLM, without the knowledge of its prompts and its underlying large language model. This practical setting poses challenges to the cross-prompt and cross-model transferability of targeted adversarial attack, which aims to confuse the LVLM to output a response that is semantically similar to the attacker's chosen target text. To this end, we propose an instruction-tuned targeted attack (dubbed InstructTA) to deliver the targeted adversarial attack on LVLMs with high transferability. Initially, we utilize a public text-to-image generative model to reverse the target response into a target image, and employ GPT-4 to infer a reasonable instruction p′ from the target response. We then form a local surrogate model (sharing the same vision encoder with the victim LVLM) to extract instruction-aware features of an adversarial image example and the target image, and minimize the distance between these two features to optimize the adversarial example. To further improve the transferability with instruction tuning, we augment the instruction p′ with instructions paraphrased from GPT-4. Extensive experiments on 6 victim LVLMs demonstrate the superiority of our proposed method in targeted attack performance and transferability. In particular, InstructTA achieves an attack success rate of 51.9% on BLIP-2, outperforming the strongest baseline by 10.5%, and consistently yields the highest attack success rates across all evaluated models. The code is available at https://github.com/xunguangwang/InstructTA.
Figures & tables
Figure 1 : The framework of our instruction-tuned targeted attack ( InstructTA ). The framework comprises three key modules: (1) a text-to-image model hξ , which generates a target image xt from the target text yt ; (2) GPT-4, which serves dual roles of inferring a plausible instruction p′ for yt and augmenting p′ into a set of semantically consistent instruction variants; and (3) a surrogate model M (sharing the same vision encoder as the victim LVLM), which acts as the substitute model for crafting adversarial examples. Given a target text yt , we first transform it into the target image xt with hξ . Simultaneously, GPT-4 infers a reasonable instruction p′ . Upon providing the augmented instructions pi′ and pj′ rephrased from p′ using GPT-4, the surrogate model M extracts instruction-aware features of xt and the AE x′ , respectively. Finally, we minimize the L2 distance between these two features to optimize x′ .
Figure 2 : The architecture of LVLM.
Figure 3 : An example of rephrasing an instruction. Given the target response yt , GPT-4 infers an instruction, i.e. , “ Can you describe what’s in the picture you’re looking at related to food? ”. The real instruction assigned to this target text is “What are the essential components depicted in this image?”
LVLM
Attack method
Text encoder (pretrained) for evaluation
NoS ( ↑ )
RN50
RN101
ViT-B/16
ViT-B/32
ViT-L/14
Ensemble
BLIP-2
Clean image
0.281
0.395
0.323
0.333
0.184
0.303
1
MF-tt Zhao et al. (2023)
0.284
0.400
0.326
0.337
0.185
0.306
0
MF-it Zhao et al. (2023)
0.533
0.634
0.575
0.581
0.469
0.559
401
MF-ii Zhao et al. (2023)
0.536
0.634
0.575
0.583
0.471
0.560
414
M-Attack Li et al. (2026)
0.324
0.438
0.366
0.375
0.228
0.346
17
Table 1 : Targeted attacks against victim LVLMs. Bold entries denote the best performance among all compared methods.
Figure 4 : Visualization examples of various targeted attack methods on InstructBLIP. “What are the essential components depicted in this image?” is a real instruction.
Figure 5 : Failure cases of InstructTA on LLaVA-1.5. The first column is the target text, and the second column is the target image generated by Stable Diffusion. The other columns are the generated responses of AEs constructed by MF-tt, MF-it, MF-ii, and InstructTA , respectively. The instruction is “What are the essential components depicted in this image?”.
Attack method
CLIP-score ( ↑ )
ASR ( ↑ )
InstructTA
0.588
0.519
InstructTA -woMF
0.574
0.472
InstructTA -woMFG
0.521
0.324
Table 2 : Ablation study of the proposed method on BLIP-2. CLIP-score Zhao et al. (2023) is computed by ensemble CLIP text encoders. Attack success rate (ASR) is the ratio between the No. of attack successes and the total No. of samples.
Figure 6 : To explore the impact of varying ϵ values within the InstructTA , we conducted experiments aiming to achieve different levels of perturbed images on BLIP-2, i.e. , referred to as the AE x′ . Our findings indicate a degradation in the visual quality of x′ , as quantified by the LPIPS Zhang et al. (2018) distance between the original image x and the adversarial image x′ . Simultaneously, the effectiveness of targeted response generation reaches a saturation point. Consequently, it is crucial to establish an appropriate perturbation budget, such as ϵ=8 , to effectively balance the image quality and the targeted attack performance.
Parameter n
Text encoder (pretrained) for evaluation
ASR ( ↑ )
RN50
RN101
ViT-B/16
ViT-B/32
ViT-L/14
Ensemble
1
0.560
0.658
0.600
0.606
0.498
0.584
0.504
5
0.560
0.657
0.600
0.606
0.498
0.584
0.504
10
0.562
0.660
0.604
0.611
0.502
0.588
0.519
50
0.562
0.659
0.601
0.607
0.499
0.586
0.529
Table 4 : Targeted attack performance for different n .
Instruction
Rephrased instruction pi/j′
Real instruction p
Inferred instruction p′
0.943
0.829
Real instruction p
0.828
1.000
Table 5 : Average CLIP-score ( ↑ ) between different versions of instructions. Each type of instruction has 1,000 samples.
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
LVLM
Attack method
Text encoder (pretrained) for evaluation
NoS ( ↑ )
RN50
RN101
ViT-B/16
ViT-B/32
ViT-L/14
Ensemble
BLIP-2
Clean image
0.275
0.397
0.319
0.328
0.177
0.299
2
MF-tt
0.277
0.400
0.320
0.331
0.177
0.301
3
MF-it
0.529
0.634
0.574
0.578
0.465
0.556
419
MF-ii
0.529
0.634
0.570
0.578
0.463
0.555
419
InstructTA
0.556
0.658
0.597
0.604
0.493
0.582
561
Appendix
Table A6 : Targeted attack performance with shuffled instructions.