cs.CVSep 27, 2026

One Attack to Fool Them All: Highly Transferable Black-Box Adversarial Attacks on Frontier MLLMs

Authors: Sen Nie, Jie Zhang, Zhongqi Wang, Shiguang Shan, Xilin Chen

Organizations: State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences, and the University of Chinese Academy of Sciences.

Abstract

Adversarial attacks have long posed a fundamental threat to machine learning systems. As multimodal large language models (MLLMs) rapidly evolve and become widely deployed, assessing their vulnerability to such attacks is essential for their safe use. In this work, we investigate whether a single adversarial image can consistently mislead diverse frontier MLLMs in black-box settings. We propose O-Attack, a highly transferable black-box attack framework. This framework builds on our insight that surrogate models contain a broad, high-level, cross-modally aligned semantic space. This space extends beyond final-layer outputs and provides multiple semantically consistent representations that remain underexploited by existing attacks. Within this space, O-Attack anchors aligned representations, progressively broadens semantic conditions, and optimizes perturbations through semantic consensus to promote consistent target alignment. By fully exploiting this space with the same surrogate models as M-Attack, O-Attack raises attack success rates on GPT-5.4 (29.1% to 77.2%), Claude-4.6 (42.8% to 81.6%), and Gemini-3.1 (38.2% to 80.9%). Extensive experiments across 24 MLLMs show that O-Attack outperforms six state-of-the-art methods in black-box transferability, with consistent effectiveness across prompts and improved efficiency and imperceptibility. This work exposes the practical safety risks posed by black-box adversarial attacks against frontier MLLMs, underscoring the need for more rigorous robustness evaluation and more effective defenses.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment

    May 11, 2026Haobo Wang, Xiaorong Ma, Weiqi Luo +2Multimodal Large Language ModelsLatent Representation Alignment

  2. Frequency-Domain Regularized Adversarial Alignment for Transferable Attacks against Closed-Source MLLMs

    May 20, 2026Leitao Yuan, Qinghua Mao, Daizong Liu +5Large Language Model AdaptationAdversarial Training

  3. DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models

    May 15, 2026Ye Sun, Xin Wang, Jiaming Zhang +7Black-Box Adversarial AttacksWhite-Box Spectral-Subspace-Guided Attack