cs.CROct 4, 2026

Which Image Property Carries the Jailbreak? A Controlled Dissection of Image-to-Text Jailbreaks

Authors: Boyuan Chen, Yehia Dawoud, Hailemariam Mersha, Minghao Shao, Siddharth Garg, Ramesh Karri, Muhammad Shafique

Abstract

Image-to-text jailbreaks place harmful intent in text, image content, or the relationship between them. We examine image-side factors across four published attack families on a 313-prompt StrongREJECT slice, using five multimodal models and an additional appendix evaluation of InternVL3.5-8B. The harmful instruction is held constant across conditions; the baseline matrix uses one draw per prompt, and paired ablations use three draws with an automated rubric judge. A bare harmful query, with or without a benign unrelated image, produces little attack success on most victims, while attack images substantially increase it. First, per-tile entropy and JPEG size do not reliably distinguish attack tiles from size-matched benign distractors, limiting density-only screening. Second, earlier tile-count ladders were confounded by payload visibility. A corrected region-count test found no detectable effect, so the role of tile-count structure remains unresolved. Third, on Qwen3-VL-8B, the E4 manipulation that removes query-specific relatedness lowers ASR by about 0.12. This supports a bounded attribution to the relatedness manipulation, although image-text congruence remains unmeasured. A within-category control reproduces the direction on overlapping stimuli. Five tests survive the global statistical correction, but only E4 supports attribution to one measured descriptor; the within-category result is a robustness check, not a separate attribution. These conclusions remain conditional on the rubric judge.

Explore similar work

CardsList
  1. When Think-with-Image Meets Safety: What Determines Multimodal Jailbreak Robustness?

    May 27, 2026Yuan Tian, Bing Hu, Fang Wu +3VLM RobustnessTool Use in VLMs

  2. Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate

    Sep 1, 2026Kaiyan Wen, Shijie Zhang, Lu Yu +1Adversarial Attacks on VLMsMulti-Agent Debate

  3. When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models

    Feb 10, 2026Jiacheng Hou, Yining Sun, Ruochong Jin +4Jailbreak AttacksMultimodal Jailbreak Attacks