cs.CVSep 30, 2026

Typographic Attack Against VLM-based AI-generated Image Detection

Authors: Eunmin Lee, Jungwoo Kim, Jong-Seok Lee

Organizations: School of Integrated Technology, Yonsei University

Abstract

Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.

Figures & tables

Explore similar work

CardsList
  1. TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images

    May 31, 2026Yuning Zhang, Changtao Miao, Mingyu Liao +5Ai-Generated Image DetectionForgeries

  2. QuISE: Defense against Typographic Attacks on VLMs via Query-Irrelevant Semantic Editing

    Aug 13, 2026Shubin Lu, Jiaqi Yin, Yihao HuangLLM Defense Mechanisms

  3. TAP into the Patch Tokens: Leveraging Vision Foundation Model Features for AI-Generated Image Detection

    Apr 29, 2026Ahmed Abdullah, Nikolas Ebert, Oliver WasenmüllerAi-Generated Image DetectionRecent Vision Foundation Models