cs.CRSep 30, 2026

SceneJail: Exploiting Video Scenario Context to Jailbreak Multimodal LLMs

Authors: Wenyu Chen, Li Wang, Chuanchao Zang, Xiangtao Meng, Xinyu Gao, Jianing Wang, Zheng Li, Shanqing Guo

Organizations: Shandong University

Abstract

Video Multimodal Large Language Models (Video-MLLMs) support reasoning over video inputs, yet remain vulnerable to jailbreak attacks that elicit policy-violating responses. Existing video jailbreaks primarily manipulate how harmful queries are visually presented, thereby treating video merely as a carrier. Consequently, the surrounding video scenario remains unexplored as a contextual attack surface. In this paper, we show that the same harmful query can elicit different safety responses when placed in different video scenarios. To systematically exploit this vulnerability, we propose SceneJail, an adaptive black-box jailbreak framework with two coordinated components. Adaptive Scenario Construction dynamically searches for a surrounding scenario that is contextually compatible with the harmful query. Scenario-aware Prompt Search uses black-box response feedback to search for textual guidance tailored to the selected scenario. Extensive evaluations on the HADES and SafeBench datasets across eight Video-MLLMs, including two proprietary models, GPT-4.1 and Gemini3.5-Flash, demonstrate the effectiveness of SceneJail. SceneJail-F, which presents the complete query persistently, achieves average attack success rates (ASR) up to 91.5%, outperforming the strongest baselines by 29.1 percentage points. Furthermore, SceneJail-S, which distributes the query across successive frames, remains highly robust against current defenses, retaining a 72.3% ASR even under strict image filtering.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Jailbreaking Multimodal Large Language Models using Multi-Clip Video

    Jun 1, 2026Choongwon Kang, Seungjong Sun, Hyunmin Jun +1Large Language Model JailbreaksVideo Multimodal Large Language Models

  2. DMN: A Compositional Framework for Jailbreaking Multimodal LLMs with Multi-Image Inputs

    May 18, 2026Wenzhuo Xu, Zhipeng Wei, Zonghao Ying +4Large Language Model JailbreaksJailbreak Attacks

  3. Fully Unleashing the Multimodal Attacker: Meta-Adaptive Jailbreaking of Vision-Language Models

    Aug 27, 2026Benlei Cui, Shen Pang, Yuke Wang +7Jailbreak Attacks