Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing representation-alignment attacks, which make a VLM perceive a target image, achieve limited success at ε≤4/255. Therefore, VLMs seems robust to perturbations in this range. We show that this robustness does not hold, as targeted semantic substitution succeeds within the same range. Specifically, we align each stream of the source image with its counterpart in the target image in the victim VLM's post-merger token space, operating under a white-box threat model. We evaluate under a strict success criterion, requiring the model to simultaneously name the target, confirm its presence, and deny the source. In images, target semantics appear at ε=2/255 and complete replacement reaches 38% at ε=4/255. On video, complete replacement reaches 35.9% at ε=1/255. We also observe a phenomenon of \textit{semantic fusion}, where Large Language Model (LLM) rationalizes contradictory visual signals into a coherent narrative.
Figures & tables
Figure 1: Qualitative image example on Qwen3-VL (dog → cat, ϵ = 4/255). Adversarial images are omitted: at ϵ = 4/255, the maximum per-pixel change is 1.6% of the full intensity range. All questions are posed independently to the VLM, each on the same adversarial image.
Figure 2: VLM pipeline and two attack surfaces. Both methods only perturb the input pixels. Token-level Representation Alignment (TokRA) (B1) rewrites what LLM sees by aligning every source and target stream at each spatial position. Output-space PGD (B2) forces what LLM says by shifting its output distribution towards target token.
Goal
Format
Success criteria
In strict ASR
Q1
Main object identification
Open-ended
Names the target
Yes
Q2
Free-form description
One sentence
Describes the target as the main object
No
Q3
Target confirmation
Yes/no
Answers Yes
Yes
Q4
Source denial
Yes/no
Answers No
Yes
Table 1: Evaluation questions. Each questions is posed independently on the same adversarial input.
State
Condition
Degeneration
Any response fails linguistic validity
Source retention
All responses valid; Q3 fails
Partial replacement
All responses valid; Q3 passes; Q1 or Q4 fails
Complete replacement
All responses valid; Q1, Q3, and Q4 pass
Table 2: Outcome states. Every sample falls into exactly one state.
Figure 5
Qwen2.5-VL
Qwen3-VL
ϵ
Strict ASR
Q3
Strict ASR
Q3
2/255
6.8
38.8
8.7
35.0
4/255
33.0
79.5
38.2
80.6
8/255
66.1
96.4
64.4
98.3
Table 3: Strict ASR and Q3 pass rate (%) of TokRA on Qwen2.5-VL and Qwen3-VL (image, 469 pairs).
Figure 5: Strict ASR vs. perturbation budget for TokRA and Output-space method on Qwen3-VL, comparing video (solid, fps = 2.0) and image modalities. The x-axis is log-scaled.
Figure 6: Qualitative video example on Qwen3-VL (movie scene → potato peeling, ϵ = 1/255). The adversarial video is omitted: at ϵ = 1/255, the maximum per-pixel change is 0.39%. All questions are posed independently to the VLM, each on the same adversarial video. TokRA succeeds on Q2–Q4 and produces a detailed Q5 response, but the output shows semantic fusion: scene-level attributes (the woman, outdoor setting, SHALIMAR watermark, vintage aesthetic) leak from the source while the action-level description (peeling a cylindrical object with a knife) follows the target. Q1 answers ”banana”, which matches neither source nor target, suggesting the model is confused rather than faithfully reproducing either scene.
Figure 7: semantic transfer across ϵ for TokRA and Output-space PGD on Qwen3-VL.
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Wu-Palmer
Categories
Appearances per category
bin
covered
mean
COCO
0.0–0.2
12
2.67
0.2–0.4
102
2.90
0.4–0.6
102
2.02
0.6–0.8
100
2.08
0.8–1.0
96
2.04
Appendix
Table 4: Category coverage of each similarity bin. The greedy selection aims for two uses per category within each bin.
Image
Video
Q1
What is the main object in this image? Answer with only the object name.
What is the main object or activity in this video? Answer with only the object or activity name.
Q2
Describe this image in exactly one short sentence. Do not add explanations.
Describe this video in exactly one short sentence. Do not add explanations.
Q3
Is there a {target} in this image? Answer only Yes or No.
Does this video contain {target}? Answer only Yes or No.
Q4
Is there a {source} in this image? Answer only Yes or No.
Does this video contain {source}? Answer only Yes or No.
Q5
Describe this image in detail
Describe this video in detail
Appendix
Table 5: Evaluation prompts. Q5 is asked only for the qualitative examples.
Parameter
Qwen3-VL-8B-Instruct
Qwen2.5-VL
Decoding
Greedy
Greedy
repetition_penalty
1.0
1.05
max_new_tokens (image)
2560
2560
max_new_tokens (video)
512
n/a
Appendix
Table 6: Decoding parameters used when querying the victim models.
Official PA-Attack (Qwen3-VL)
Targeted PA-Attack (ours)
Target
Least similar prototype (untargeted)
Target-image feature at the same position (targeted)
ε
8/255
{1,2,4,8,16}/255
Step size
4/255
4/255
Steps (stage 1 + stage 2)
100 + 200
200 + 300
Update
v←sign(0.9v+sign(g))
Same
Initialization
Uniform random at each stage
Same
Appendix
Table 7: PA-Attack hyperparameters.
ε
Q1
Q2
Q3
Q4
Strict ASR
1/255
0.43
0.85
8.10
0.43
0.00
2/255
0.64
0.85
8.96
1.28
0.00
4/255
0.43
1.49
8.53
7.46
0.00
8/255
0.64
1.28
11.51
22.39
0.00
16/255
0.85
1.07
10.66
30.92
0.00
Appendix
Table 8: Pass rates (%) of the targeted PA-Attack on Qwen3-VL images ( n=469 ).
ε
Q1
Q2
Q3
Q4
Strict ASR
1/255
1.56
8.59
16.41
3.13
0.00
2/255
3.91
9.38
17.97
7.03
0.00
4/255
6.25
7.81
20.31
20.31
1.56
8/255
3.13
7.03
20.31
30.47
0.78
16/255
3.91
7.03
15.63
74.22
2.34
Appendix
Table 9: Pass rates (%) of the targeted PA-Attack on Qwen3-VL videos ( n=128 , 2 fps).
Aligned features
Strict ASR
Q1
Q3
Post-merger, main stream only
12.58
22.17
45.63
Post-merger, DeepStack only
50.11
64.39
89.77
Pre-merger, main stream + DeepStack
47.12
62.47
89.55
TokRA (post-merger, main stream + DeepStack)
38.17
54.16
80.60
Appendix
Table 10: Alignment location ablation on Qwen3-VL images ( n=469 , ε=4/255 ). All numbers are pass rates (%).
fps
Sampled frames
Frame indices
Temporal patches T
Visual tokens
0.5
4
0, 50, 99, 149
2
896
1
4
0, 50, 99, 149
2
896
2
10
Evenly spaced, about 16.6 apart
5
2,240
4
20
Evenly spaced, about 7.8 apart
10
4,480
Appendix
Table 11: Sampled frames and visual tokens at each fps.
Method
fps
Q1
Q2
Q3
Q4
Strict ASR
TokRA
0.5
69.53
68.75
85.94
75.78
57.81
1
67.97
70.31
84.38
78.13
55.47
2
75.78
76.56
90.63
78.13
63.28
4
75.00
80.47
95.31
77.34
60.94
Output-space
0.5
29.69
32.81
63.28
51.56
17.19
1
32.03
33.59
61.72
40.63
16.41
Appendix
Table 12: Pass rates (%) on Qwen3-VL videos at ε=2/255 for each fps ( n=128 ). Pass criteria follow Table 1 . Strict ASR requires Q1, Q3, and Q4 to pass and all four responses to be valid.
Figure 8: Qualitative video example on Qwen3-VL (cat → sauce, ε=1/255 ).
Vision-Language Models (VLMs) are increasingly deployed in autonomous driving and embodied AI systems, where reliable perception is critical for safe semantic reasoning and decision-making. While recent VLMs demonstrate strong performance on multimodal benchmarks, their robustness to realistic perception degradation remains poorly understood. In this work, we systematically study semantic misalignment in VLMs under controlled degradation of upstream visual perception, using semantic segmentation on the Cityscapes dataset as a representative perception module. We introduce perception-realistic corruptions that induce only moderate drops in conventional segmentation metrics, yet observe severe failures in downstream VLM behavior, including hallucinated object mentions, omission of safety-critical entities, and inconsistent safety judgments. To quantify these effects, we propose a set of language-level misalignment metrics that capture hallucination, critical omission, and safety misinterpretation, and analyze their relationship with segmentation quality across multiple contrastive and generative VLMs. Our results reveal a clear disconnect between pixel-level robustness and multimodal semantic reliability, highlighting a critical limitation of current VLM-based systems and motivating the need for evaluation frameworks that explicitly account for perception uncertainty in safety-critical applications.
Vision-language models (VLMs) are increasingly deployed as trusted authorities -- fact-checking images on social media, comparing products, and moderating content. Users implicitly trust that these systems perceive the same visual content as they do. We show that adversarial examples break this assumption, enabling \emph{AI authority laundering}: an attacker subtly perturbs an image so that the VLM produces confident and authoritative responses about the \emph{wrong} input. Unlike jailbreaks or prompt injections, our attacks do not compromise model alignment; the attack operates entirely at the perceptual level. We demonstrate that standard attacks against publicly available CLIP models transfer reliably to production VLMs -- including GPT-5.4, Claude Opus4.6, Gemini3, and Grok~4.2. Across four attack surfaces, we show that authority laundering can amplify misinformation, disparage individuals, evade content moderation, and manipulate product recommendations. Our attacks have high success rates: In hundreds of attacks targeting identity manipulation and NSFW evasion, we measure success rates of 22−100% across six models. No novel attack algorithm is required: basic techniques known for over a decade suffice, establishing a lower bound on attacker capability that should concern defenders. Our results demonstrate that visual adversarial robustness is now a practical -- and still largely unsolved -- safety problem.
Jie Zhang, Pura Peetathawatchai, Florian Tramèr +1
Vision-Language Models (VLMs) are now a core part of modern AI. Recent work proposed several visual jailbreak attacks using single/ holistic images. However, contemporary VLMs demonstrate strong robustness against such attacks due to extensive safety alignment through preference optimization, e.g., reinforcement learning from human feedback (RLHF). In this work, we identify a new vulnerability: while VLM pretraining and instruction tuning generalize well to split-image inputs, safety alignment is typically performed only on holistic images and does not account for harmful semantics distributed across multiple image fragments. Consequently, VLMs often fail to detect and reject harmful split-image inputs, in which unsafe cues emerge only upon combining images. We introduce novel split-image visual jailbreak attacks (\textbf{SIVA}) that exploit this misalignment. Unlike prior optimization-based attacks, which exhibit poor black-box transferability due to architectural and prior mismatches across models, our attacks evolve in progressive phases from naive splitting to an adaptive white-box attack, culminating in a black-box transfer attack. Our strongest strategy leverages a novel adversarial knowledge distillation \textbf{(Adv-KD)} algorithm to substantially improve cross-model transferability. Evaluations on four state-of-the-art modern VLMs and three jailbreak datasets demonstrate that our strongest attack achieves up to 44% higher transfer success than existing baselines. Lastly, we propose efficient ways to address this critical vulnerability in the current VLM safety alignment.