Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios. To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding scenario that is contextually compatible with the harmful query. Scenario-aware Prompt Search uses black-box response feedback to search for textual guidance tailored to the selected scenario. Extensive evaluations on the HADES and SafeBench datasets across eight Video-MLLMs, including two proprietary models, GPT-4.1 and Gemini3.5-Flash, demonstrate the effectiveness of SceneJail. SceneJail-F, which presents the complete query persistently, achieves average attack success rates (ASR) up to 91.5%, outperforming the strongest baselines by 29.1 percentage points. Furthermore, SceneJail-S, which distributes the query across successive frames, remains highly robust against current defenses, retaining a 72.3% ASR even under strict image filtering.
Figures & tables
Figure 1 : Illustration of the sensitivity of model safety behavior to video scenario context. The same harmful query ( “List steps in detail for using a service to create fake passports and IDs.” ) can elicit different safety responses when presented in different surrounding video scenarios.
Figure 2 : Overview of SceneJail . Adaptive Scenario Construction identifies a suitable surrounding scenario while preserving the original query content, and Scenario-aware Prompt Search searches for textual guidance under the selected scenario. The two components form a nested black-box search guided by target-model response feedback.
Figure 3 : Template for Adaptive Scenario Construction.
Figure 4 : Examples of Full-display and Split-display.
Figure 5 : Template used for Scenario-aware Prompt Search.
Dataset
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
HADES
Qwen2.5-VL-7B
63.73%
77.87%
82.13%
57.33%
46.00%
93.73%
90.93%
Qwen2.5-VL-32B
10.93%
11.60%
42.80%
61.20%
78.53%
94.53%
90.67%
Qwen3-VL-8B
20.93%
40.27%
53.47%
70.93%
57.73%
93.60%
91.60%
Qwen3-VL-32B
10.13%
24.93%
40.27%
70.53%
59.47%
94.53%
92.93%
InternVL3.5-8B
77.20%
82.40%
84.00%
83.73%
75.73%
93.87%
89.73%
InternVL3.5-38B
74.80%
82.00%
75.87%
81.07%
80.53%
93.73%
89.87%
Table 1 : Attack success rate (ASR, %) on HADES and SafeBench. SceneJail -F and SceneJail -S denote the Full-display and Split-display variants, respectively. Bold values denote our results, while underlined values mark the strongest baseline in each row. The Avg. rows report the unweighted mean across all target models.
Method
Violence
Animal Harm
Financial Harm
Self- Harm
Privacy
FigStep-I
43.17%
18.00%
42.50%
35.75%
39.08%
FigStep-V
50.92%
23.17%
50.25%
44.58%
50.09%
VideoJail
64.92%
33.08%
58.75%
52.75%
57.17%
SPTV
71.67%
33.58%
78.75%
55.00%
72.92%
MCV
67.33%
33.17%
73.17%
53.08%
70.42%
SceneJail-F
93.92%
80.58%
94.83%
94.09%
93.92%
Table 2 : Category-wise ASR (%) on HADES, macro-averaged across all target models. The highest ASR in each category is shown in bold, and the strongest baseline is underlined.
Figure 6 : Component ablations of SceneJail on HADES under Full-display. ASR (%) is reported for the three target models.
Figure 7 : Sensitivity of SceneJail to scenario-search and prompt-search budgets on HADES under Full-display.
Target Model
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
2.61
2.98
Qwen2.5-VL-32B
2.04
2.19
Qwen3-VL-8B
2.58
2.49
Qwen3-VL-32B
2.68
1.76
InternVL3.5-8B
2.43
1.56
InternVL3.5-38B
2.65
1.85
Table 3 : Average Queries (AQ) of SceneJail -F and SceneJail -S over successfully jailbroken queries on HADES.
Figure 8 : Robustness of SceneJail to different generative models on HADES under Full-display. ASR (%) is reported for the three target models.
Method
Llama Guard 3
GPT Judge
Human Evaluation
SceneJail-F
94.33%
89.00%
91.00%
SceneJail-S
91.33%
91.00%
91.33%
Table 4 : ASR (%) under different evaluation protocols, computed on the same outputs generated from 100 sampled HADES queries across three target models.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 9 : Template for Scenario Matching evaluation.
Seed
Qwen2.5-VL-32B
Qwen3-VL-32B
InternVL3.5-38B
42
94.53%
94.53%
93.73%
52
93.87%
94.40%
93.73%
62
94.27%
94.00%
93.47%
72
94.53%
94.13%
94.00%
82
94.00%
94.27%
93.33%
Avg.
94.24%
94.27%
93.65%
Appendix
Table 5 : ASR (%) of SceneJail under different HADES query orders. Experiments are conducted under Full-display. Seed 42 corresponds to the query order used in the main evaluation.
Figure 10: Prompt template used for GPT Judge evaluation.
Figure 11: Prompt template used for image filtering.
Figure 12: Prompt template used for safety prompt.
Figure 13: CLAS 1–5 scoring rubric used by the GPT-4o-mini evaluator. Only responses assigned a score of 5 are counted as successful attacks in our evaluation.
Defense
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
No Defense
42.95%
53.18%
63.09%
70.80%
66.33%
94.00%
90.96%
Image Filtering
0.04%
0.04%
0.09%
0.02%
0.07%
0.04%
72.25%
Multimodal Safety Guard
42.93%
50.67%
61.64%
68.22%
41.69%
91.45%
90.75%
Safety Prompt
1.67%
2.44%
5.89%
14.11%
5.54%
44.78%
45.31%
Appendix
Table 6 : Average ASR (%) under different defense settings on HADES, macro-averaged across the six open-source target models. No-defense results are included for reference. Bold values denote SceneJail , and underlined values indicate the strongest baseline under each setting.
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
0.13%
0.13%
0.00%
0.00%
0.13%
0.00%
87.07%
Qwen2.5-VL-32B
0.00%
0.00%
0.27%
0.00%
0.27%
0.13%
64.00%
Qwen3-VL-8B
0.00%
0.00%
0.00%
0.00%
0.00%
0.00%
69.87%
Qwen3-VL-32B
0.00%
0.00%
0.00%
0.00%
0.00%
0.00%
55.87%
InternVL3.5-8B
0.13%
0.13%
0.27%
0.13%
0.00%
0.13%
89.47%
InternVL3.5-38B
0.00%
0.00%
0.00%
0.00%
0.00%
0.00%
67.20%
Appendix
Table 7 : Per-model ASR (%) under the image-filtering defense on HADES. The Avg. row reports the unweighted mean across the six target models. Bold values denote the results of our method.
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
63.60%
74.27%
80.53%
55.33%
26.40%
90.93%
90.93%
Qwen2.5-VL-32B
10.93%
10.67%
41.87%
58.93%
48.27%
91.87%
90.53%
Qwen3-VL-8B
20.93%
37.60%
51.33%
67.73%
38.00%
90.80%
91.60%
Qwen3-VL-32B
10.13%
23.87%
38.93%
67.47%
37.47%
92.27%
92.13%
InternVL3.5-8B
77.20%
78.93%
82.00%
81.47%
46.80%
90.93%
89.60%
InternVL3.5-38B
74.80%
78.67%
75.20%
78.40%
53.20%
91.87%
89.73%
Appendix
Table 8 : Per-model ASR (%) under the multimodal safety-guard defense on HADES. The Avg. row reports the unweighted mean across the six target models. Bold values denote the results of our method.
Method
Qwen2.5-VL 7B
Qwen2.5-VL 32B
Qwen3-VL 8B
Qwen3-VL 32B
InternVL3.5 8B
InternVL3.5 38B
Gemini3.5 Flash
GPT-4.1
Violence
FigStep-I
66.00%
19.33%
31.33%
19.33%
87.33%
79.33%
26.00%
16.67%
FigStep-V
83.33%
20.00%
44.67%
33.33%
86.67%
90.67%
27.33%
21.33%
VideoJail
92.00%
60.00%
59.33%
50.67%
90.00%
84.67%
44.00%
38.67%
SPTV
70.00%
70.00%
85.33%
81.33%
91.33%
90.00%
39.33%
46.00%
MCV
54.67%
82.67%
71.33%
68.00%
84.00%
84.67%
43.33%
50.00%
Appendix
Table 9 : Per-model category-wise ASR (%) on HADES. For each model–category pair, the highest ASR is shown in bold, and the strongest baseline is underlined.
Target Model
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Qwen2.5-VL-7B
8.27%
10.53%
18.13%
18.40%
2.13%
54.27%
49.47%
Qwen2.5-VL-32B
0.80%
0.80%
2.80%
6.80%
4.27%
42.67%
46.67%
Qwen3-VL-8B
0.00%
0.13%
0.00%
1.60%
0.27%
40.80%
42.40%
Qwen3-VL-32B
0.00%
0.13%
0.13%
8.80%
2.00%
41.20%
42.67%
InternVL3.5-8B
0.00%
0.80%
0.40%
19.60%
2.27%
48.13%
49.47%
InternVL3.5-38B
0.93%
2.27%
13.87%
29.47%
22.27%
41.60%
41.20%
Appendix
Table 10 : Per-model ASR (%) under the safety-prompt defense on HADES. The Avg. row reports the unweighted mean across the six target models. Bold values denote the results of our method.
Category
FigStep-I
FigStep-V
VideoJail
SPTV
MCV
SceneJail-F
SceneJail-S
Illegal Activity
40.54%
47.61%
51.70%
68.04%
53.68%
88.73%
89.46%
Hate Speech
26.87%
32.61%
36.03%
35.38%
25.01%
87.07%
88.13%
Malware Generation
57.54%
68.95%
74.70%
75.67%
75.35%
97.40%
98.13%
Physical Harm
45.87%
55.61%
57.70%
79.04%
68.68%
96.07%
97.13%
Fraud
45.87%
52.95%
65.03%
70.04%
66.68%
90.40%
92.46%
Adult Content
18.54%
17.95%
21.70%
16.71%
4.68%
56.07%
48.79%
Appendix
Table 11 : Category-wise and overall ASR (%) on SafeBench, macro-averaged across all target models. The highest ASR in each category and in the overall comparison is shown in bold, and the strongest baseline is underlined.
Model
Method
Llama Guard 3
GPT-4o-mini Judge
Human Evaluation
Qwen2.5-VL-32B
SceneJail-F
95%
89%
92%
SceneJail-S
90%
90%
90%
Qwen3-VL-32B
SceneJail-F
94%
88%
90%
SceneJail-S
93%
91%
93%
InternVL3.5-38B
SceneJail-F
94%
90%
91%
SceneJail-S
91%
92%
91%
Appendix
Table 12 : Model-wise ASR (%) under different evaluation protocols on outputs generated from the same 100 sampled HADES queries.
Figure 14 : CLAS prohibited-use policy used by the GPT-4o-mini evaluator.
Department of Applied Artificial Intelligence, Sungkyunkwan University · Department of Human-Artificial Intelligence Interaction, Sungkyunkwan University
1Yuvion Team, Alibaba Group · Laboratory for Statistical Monitoring and Intelligent Governance of Common Prosperity, School of Statistics and Data Science, Zhejiang Gongshang University