Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios. To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding scenario that is contextually compatible with the harmful query. Scenario-aware Prompt Search uses black-box response feedback to search for textual guidance tailored to the selected scenario. Extensive evaluations on the HADES and SafeBench datasets across eight Video-MLLMs, including two proprietary models, GPT-4.1 and Gemini3.5-Flash, demonstrate the effectiveness of SceneJail. SceneJail-F, which presents the complete query persistently, achieves average attack success rates (ASR) up to 91.5%, outperforming the strongest baselines by 29.1 percentage points. Furthermore, SceneJail-S, which distributes the query across successive frames, remains highly robust against current defenses, retaining a 72.3% ASR even under strict image filtering.
Figures & tables
Figure 1 : Illustration of the sensitivity of model safety behavior to video scenario context. The same harmful query ( “List steps in detail for using a service to create fake passports and IDs.” ) can elicit different safety responses when presented in different surrounding video scenarios.
Figure 2 : Overview of SceneJail . Adaptive Scenario Construction identifies a suitable surrounding scenario while preserving the original query content, and Scenario-aware Prompt Search searches for textual guidance under the selected scenario. The two components form a nested black-box search guided by target-model response feedback.
Figure 3 : Template for Adaptive Scenario Construction.
Figure 4 : Examples of Full-display and Split-display.
Figure 5 : Template used for Scenario-aware Prompt Search.
Dataset
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
HADES
Qwen2.5-VL-7B
63.73%
77.87%
82.13%
57.33%
46.00%
93.73%
90.93%
Qwen2.5-VL-32B
10.93%
11.60%
42.80%
61.20%
78.53%
94.53%
90.67%
Qwen3-VL-8B
20.93%
40.27%
53.47%
70.93%
57.73%
93.60%
91.60%
Qwen3-VL-32B
10.13%
24.93%
40.27%
70.53%
59.47%
94.53%
92.93%
InternVL3.5-8B
77.20%
82.40%
84.00%
83.73%
75.73%
93.87%
89.73%
InternVL3.5-38B
74.80%
82.00%
75.87%
81.07%
80.53%
93.73%
89.87%
Table 1 : Attack success rate (ASR, %) on HADES and SafeBench. SceneJail -F and SceneJail -S denote the Full-display and Split-display variants, respectively. Bold values denote our results, while underlined values mark the strongest baseline in each row. The Avg. rows report the unweighted mean across all target models.
Method
Violence
Animal Harm
Financial Harm
Self- Harm
Privacy
FigStep-I
43.17%
18.00%
42.50%
35.75%
39.08%
FigStep-V
50.92%
23.17%
50.25%
44.58%
50.09%
VideoJail
64.92%
33.08%
58.75%
52.75%
57.17%
SPTV
71.67%
33.58%
78.75%
55.00%
72.92%
MCV
67.33%
33.17%
73.17%
53.08%
70.42%
SceneJail-F
93.92%
80.58%
94.83%
94.09%
93.92%
Table 2 : Category-wise ASR (%) on HADES, macro-averaged across all target models. The highest ASR in each category is shown in bold, and the strongest baseline is underlined.
Figure 6 : Component ablations of SceneJail on HADES under Full-display. ASR (%) is reported for the three target models.
Figure 7 : Sensitivity of SceneJail to scenario-search and prompt-search budgets on HADES under Full-display.
Target Model
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
2.61
2.98
Qwen2.5-VL-32B
2.04
2.19
Qwen3-VL-8B
2.58
2.49
Qwen3-VL-32B
2.68
1.76
InternVL3.5-8B
2.43
1.56
InternVL3.5-38B
2.65
1.85
Table 3 : Average Queries (AQ) of SceneJail -F and SceneJail -S over successfully jailbroken queries on HADES.
Figure 8 : Robustness of SceneJail to different generative models on HADES under Full-display. ASR (%) is reported for the three target models.
Method
Llama Guard 3
GPT Judge
Human Evaluation
SceneJail-F
94.33%
89.00%
91.00%
SceneJail-S
91.33%
91.00%
91.33%
Table 4 : ASR (%) under different evaluation protocols, computed on the same outputs generated from 100 sampled HADES queries across three target models.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : Template for Scenario Matching evaluation.
Seed
Qwen2.5-VL-32B
Qwen3-VL-32B
InternVL3.5-38B
42
94.53%
94.53%
93.73%
52
93.87%
94.40%
93.73%
62
94.27%
94.00%
93.47%
72
94.53%
94.13%
94.00%
82
94.00%
94.27%
93.33%
Avg.
94.24%
94.27%
93.65%
Appendix
Table 5 : ASR (%) of SceneJail under different HADES query orders. Experiments are conducted under Full-display. Seed 42 corresponds to the query order used in the main evaluation.
Figure 10: Prompt template used for GPT Judge evaluation.
Figure 11: Prompt template used for image filtering.
Figure 12: Prompt template used for safety prompt.
Figure 13: CLAS 1–5 scoring rubric used by the GPT-4o-mini evaluator. Only responses assigned a score of 5 are counted as successful attacks in our evaluation.
Defense
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
No Defense
42.95%
53.18%
63.09%
70.80%
66.33%
94.00%
90.96%
Image Filtering
0.04%
0.04%
0.09%
0.02%
0.07%
0.04%
72.25%
Multimodal Safety Guard
42.93%
50.67%
61.64%
68.22%
41.69%
91.45%
90.75%
Safety Prompt
1.67%
2.44%
5.89%
14.11%
5.54%
44.78%
45.31%
Appendix
Table 6 : Average ASR (%) under different defense settings on HADES, macro-averaged across the six open-source target models. No-defense results are included for reference. Bold values denote SceneJail , and underlined values indicate the strongest baseline under each setting.
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
0.13%
0.13%
0.00%
0.00%
0.13%
0.00%
87.07%
Qwen2.5-VL-32B
0.00%
0.00%
0.27%
0.00%
0.27%
0.13%
64.00%
Qwen3-VL-8B
0.00%
0.00%
0.00%
0.00%
0.00%
0.00%
69.87%
Qwen3-VL-32B
0.00%
0.00%
0.00%
0.00%
0.00%
0.00%
55.87%
InternVL3.5-8B
0.13%
0.13%
0.27%
0.13%
0.00%
0.13%
89.47%
InternVL3.5-38B
0.00%
0.00%
0.00%
0.00%
0.00%
0.00%
67.20%
Appendix
Table 7 : Per-model ASR (%) under the image-filtering defense on HADES. The Avg. row reports the unweighted mean across the six target models. Bold values denote the results of our method.
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
63.60%
74.27%
80.53%
55.33%
26.40%
90.93%
90.93%
Qwen2.5-VL-32B
10.93%
10.67%
41.87%
58.93%
48.27%
91.87%
90.53%
Qwen3-VL-8B
20.93%
37.60%
51.33%
67.73%
38.00%
90.80%
91.60%
Qwen3-VL-32B
10.13%
23.87%
38.93%
67.47%
37.47%
92.27%
92.13%
InternVL3.5-8B
77.20%
78.93%
82.00%
81.47%
46.80%
90.93%
89.60%
InternVL3.5-38B
74.80%
78.67%
75.20%
78.40%
53.20%
91.87%
89.73%
Appendix
Table 8 : Per-model ASR (%) under the multimodal safety-guard defense on HADES. The Avg. row reports the unweighted mean across the six target models. Bold values denote the results of our method.
Method
Qwen2.5-VL 7B
Qwen2.5-VL 32B
Qwen3-VL 8B
Qwen3-VL 32B
InternVL3.5 8B
InternVL3.5 38B
Gemini3.5 Flash
GPT-4.1
Violence
FigStep-I
66.00%
19.33%
31.33%
19.33%
87.33%
79.33%
26.00%
16.67%
FigStep-V
83.33%
20.00%
44.67%
33.33%
86.67%
90.67%
27.33%
21.33%
VideoJail
92.00%
60.00%
59.33%
50.67%
90.00%
84.67%
44.00%
38.67%
SPTV
70.00%
70.00%
85.33%
81.33%
91.33%
90.00%
39.33%
46.00%
MCV
54.67%
82.67%
71.33%
68.00%
84.00%
84.67%
43.33%
50.00%
Appendix
Table 9 : Per-model category-wise ASR (%) on HADES. For each model–category pair, the highest ASR is shown in bold, and the strongest baseline is underlined.
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
8.27%
10.53%
18.13%
18.40%
2.13%
54.27%
49.47%
Qwen2.5-VL-32B
0.80%
0.80%
2.80%
6.80%
4.27%
42.67%
46.67%
Qwen3-VL-8B
0.00%
0.13%
0.00%
1.60%
0.27%
40.80%
42.40%
Qwen3-VL-32B
0.00%
0.13%
0.13%
8.80%
2.00%
41.20%
42.67%
InternVL3.5-8B
0.00%
0.80%
0.40%
19.60%
2.27%
48.13%
49.47%
InternVL3.5-38B
0.93%
2.27%
13.87%
29.47%
22.27%
41.60%
41.20%
Appendix
Table 10 : Per-model ASR (%) under the safety-prompt defense on HADES. The Avg. row reports the unweighted mean across the six target models. Bold values denote the results of our method.
Category
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Illegal Activity
40.54%
47.61%
51.70%
68.04%
53.68%
88.73%
89.46%
Hate Speech
26.87%
32.61%
36.03%
35.38%
25.01%
87.07%
88.13%
Malware Generation
57.54%
68.95%
74.70%
75.67%
75.35%
97.40%
98.13%
Physical Harm
45.87%
55.61%
57.70%
79.04%
68.68%
96.07%
97.13%
Fraud
45.87%
52.95%
65.03%
70.04%
66.68%
90.40%
92.46%
Adult Content
18.54%
17.95%
21.70%
16.71%
4.68%
56.07%
48.79%
Appendix
Table 11 : Category-wise and overall ASR (%) on SafeBench, macro-averaged across all target models. The highest ASR in each category and in the overall comparison is shown in bold, and the strongest baseline is underlined.
Model
Method
Llama Guard 3
GPT-4o-mini Judge
Human Evaluation
Qwen2.5-VL-32B
SceneJail-F
95%
89%
92%
SceneJail-S
90%
90%
90%
Qwen3-VL-32B
SceneJail-F
94%
88%
90%
SceneJail-S
93%
91%
93%
InternVL3.5-38B
SceneJail-F
94%
90%
91%
SceneJail-S
91%
92%
91%
Appendix
Table 12 : Model-wise ASR (%) under different evaluation protocols on outputs generated from the same 100 sampled HADES queries.
Figure 14 : CLAS prohibited-use policy used by the GPT-4o-mini evaluator.
As multimodal large language models (MLLMs) have advanced to process video inputs, concerns have emerged about their potential for malicious misuse. Prior jailbreak studies have shown that safety alignment in MLLMs can be bypassed through visual inputs, yet it remains unclear which properties of video inputs induce this vulnerability. To address this gap, we introduce Multi-Clip Video (MCV) SafetyBench, a dataset of 2,920 videos designed to evaluate how the diversity of video inputs affects the vulnerability of MLLMs. Each video consists of multiple short clips depicting diverse contexts related to a harmful query. Experiments on eight representative video MLLMs show that attack success consistently increases with the number of clips. Our results further indicate that the video modality is (1) more vulnerable than the image modality, (2) more vulnerable to dynamic videos than to static videos, and (3) more vulnerable when videos contain more diverse contexts. Building on these findings, we propose a defense strategy that leverages the relative robustness of the image modality.
Choongwon Kang, Seungjong Sun, Hyunmin Jun +1
Department of Applied Artificial Intelligence, Sungkyunkwan University · Department of Human-Artificial Intelligence Interaction, Sungkyunkwan University
Multimodal Large Language Models (MLLMs) are vulnerable to jailbreak attacks, which can elicit harmful responses from MLLMs. Many MLLMs support multi-image inputs, inadvertently introducing new vulnerabilities due to less efforts on multi-image safety alignment. Previous MLLM jailbreak methods only uses a single image, which restricts the attack space: they cannot distribute harmful requests across multiple images, carry abundant information, or exploit additional visual reasoning tasks to distract MLLMs. To address these limitations, in this paper, we propose a compositional jailbreak framework, \textbf{DMN}, which leverages \textbf{D}istributed instruction, \textbf{M}ultimodal evidence and a \textbf{N}umber chain task to fully enhance the jailbreak performance. Extensive experiments show that DMN is highly effective for MLLM jailbreaking, e.g. achieving attack success rates of over 90% on GPT-4o, Gemini-2.5-pro and Claude Sonnet 4, surpassing other baselines by a large margin. This compositional, multi-image jailbreak strategy reveals fundamental weaknesses in their safety mechanisms.
Wenzhuo Xu, Zhipeng Wei, Zonghao Ying +4
AI Security Lab, Beijing, China · International Computer Science Institute, CA, USA · UC Berkeley, CA, USA +1
The safety of large vision-language models is increasingly stress-tested by multimodal jailbreaks, yet existing attacks remain largely static at the meta level: template-based attacks freeze the image-text layout, while iterative attacks adapt only the image-text content with fixed attack strategies and frozen attacker parameters. We propose Meta-Adaptive Multimodal Jailbreaking (MAMJ), which instead optimizes the attacker itself along two axes: an attack strategy prompt (ASP) governing attack iteration and attacker model weights determining attack effectiveness. Across groups of multimodal attack trajectories, an LLM-based critique first refines the ASP, after which group-aggregated attack success rate (ASR) rewards update those weights. On MM-SafetyBench, MAMJ achieves 81.0%, 78.9%, and 82.3% ASR against GPT-4o, Gemini-3-Pro-Preview, and Seed 2.0, respectively, outperforming the strongest sample-level baseline by up to 24.1 percentage points. The learned attacker, comprising the optimized ASP and attacker weights, also transfers without retraining to unseen victims and remains effective under representative defenses. These results reveal a systemic vulnerability of frontier VLMs to meta-adaptive jailbreaks and motivate defenses against meta-level adversaries. Code is available at https://github.com/Alibaba-VELLDEPTH/MetaJailbreak-VLM.
Benlei Cui, Shen Pang, Yuke Wang +7
1Yuvion Team, Alibaba Group · Laboratory for Statistical Monitoring and Intelligent Governance of Common Prosperity, School of Statistics and Data Science, Zhejiang Gongshang University