Vision-language models (VLMs) are increasingly used for AI-generated image (AIGI) detection, providing natural-language explanations for authenticity judgments. However, their ability to interpret text within images may also expose these judgments to misleading semantic cues. We systematically evaluate typographic attack strategies across detection-oriented, open-weight, and commercial VLMs, considering both real-to-fake and fake-to-real attacks. Our results show that reasoning modes generally exhibit greater vulnerability than direct modes and that attack effectiveness exhibits pronounced directional asymmetry. Moreover, larger models tend to exhibit higher clean detection accuracy but also higher attack success rates. We further examine attack robustness under image and text transformations and investigate whether overlays indicating the correct class can aid error correction. Together, these analyses characterize how typographic attacks influence authenticity judgments and expose limitations of current VLM-based AIGI detection systems.
Figures & tables
Figure 1 : Visual Examples of Typographic Attacks. The top row shows real-to-fake attacks, and the bottom row shows fake-to-real attacks.
Detection-oriented VLMs
Open-weight VLMs
Commercial VLMs
Ivy
Veritas++
BusterX++ D
BusterX++ R
Qwen3.8 D
Qwen3.8 R
GLM4.6V-F D
GLM4.6V-F R
GPT-5.4
Sonnet 5
R
F
R
F
R
F
R
F
R
F
R
F
R
F
R
F
R
F
R
F
Clean Acc. (%)
85.68
80.88
99.28
40.40
93.92
41.92
84.80
55.20
93.04
51.52
77.36
53.04
93.36
39.44
95.76
38.96
97.20
48.72
97.92
41.76
Attack
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
Random Text
0.19
21.46
1.77
25.35
9.20
17.56
19.06
17.68
15.82
4.66
24.61
14.63
15.77
2.43
15.37
2.26
6.67
3.94
8.91
6.90
Class Label
17.18
34.12
26.11
35.05
75.81
24.81
91.89
42.75
42.39
43.32
87.07
68.78
81.92
5.88
77.36
5.95
32.92
45.16
60.95
37.16
Table 1 : Clean AIGI Detection Accuracy and Attack Success Rate (%) on Ivy-Fake [ 10 ] . For clean accuracy, R and F denote performance on real and fake images, respectively. R → F denotes real-to-fake attack, while F → R denotes fake-to-real attack. Model subscripts D and R denote direct and reasoning modes, respectively. Bold and underlined values denote the highest and second-highest ASR within each column, respectively.
Detection-oriented VLMs
Open-weight VLMs
Commercial VLMs
Ivy
Veritas++
BusterX++ D
BusterX++ R
Qwen3.8 D
Qwen3.8 R
GLM4.6V-F D
GLM4.6V-F R
GPT-5.4
Sonnet 5
R
F
R
F
R
F
R
F
R
F
R
F
R
F
R
F
R
F
R
F
Clean Acc. (%)
83.28
99.80
98.13
84.00
92.30
71.00
79.43
93.20
93.18
65.50
87.68
76.60
93.62
9.50
94.61
10.00
97.69
53.80
98.79
32.90
Attack
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
R → F
F → R
Random Text
0.93
4.01
0.79
16.55
9.77
5.77
13.71
2.15
14.64
7.63
18.44
6.53
8.81
8.42
6.28
17.00
2.03
4.28
0.67
6.08
Class Label
44.65
6.01
9.98
25.24
77.59
8.59
92.94
12.12
41.91
44.89
77.67
34.99
82.02
25.26
65.35
38.00
18.58
51.12
29.62
29.79
Table 2 : Clean AIGI Detection Accuracy and Attack Success Rate (%) on GenImage [ 29 ] . For clean accuracy, R and F denote performance on real and fake images, respectively. R → F denotes real-to-fake attack, while F → R denotes fake-to-real attack. Model subscripts D and R denote direct and reasoning modes, respectively. Bold and underlined values denote the highest and second-highest ASR within each column, respectively.
Figure 2 : Qualitative Examples of Model Responses. Examples illustrating the effects of reasoning mode and attack direction on model outputs. Blue-highlighted texts denote the failed attacks, and red-highlighted texts denote the successful attacks.
4B
9B
27B
R
F
R
F
R
F
Clean Acc. (%)
90.00
50.00
83.00
66.00
80.00
88.00
Attack
R → F
F → R
R → F
F → R
R → F
F → R
Random Text
14.44
16.00
18.07
10.61
20.00
7.95
Class Label
56.67
60.00
90.36
65.15
88.75
65.91
Instruction
74.44
74.00
93.98
90.91
97.50
86.36
Table 3 : Model Scale Analysis of Qwen3.5 R on Ivy-Fake Subset. ASR (%) is computed over the subset of images correctly classified by each model under clean conditions.
Figure 3 : Robustness of Typographic Attacks. ASR under image- and text-level perturbations.
Ivy
Veritas++
Condition
R ↪ F
F ↪ R
Avg.
R ↪ F
F ↪ R
Avg.
Random Text
2.0
74.0
38.0
9.9
66.7
15.0
Class Label
44.0
84.0
64.0
59.3
77.8
61.0
Instruction
48.0
100.0
74.0
54.9
100.0
59.0
File Path
28.0
94.0
61.0
58.2
66.7
59.0
Logo
2.0
96.0
49.0
11.0
88.9
18.0
Table 4 : Recovery Rate (%) by Amicable Aid. F ↪ R denotes correction of a real image initially misclassified as fake, while R ↪ F denotes correction of a fake image initially misclassified as real. Ivy and Veritas++ use 50/50 and 9/91 real/fake images, respectively.
(a) Position
TL
TC †
TR
BL
BC
BR
Ivy
31.88
32.50
32.50
36.88
41.25
37.88
Veritas++
38.88
36.13
39.50
36.13
34.63
33.38
Overall
35.38
34.31
36.00
36.50
37.94
35.63
Table 5 : Rendering Ablation. ASR (%) over four attacks and both directions, excluding Random Text. † denotes reference settings (Ours). In (a), T/B denote top/bottom, and L/C/R denote left/center/right, respectively. In (b), the font size is obtained by multiplying each fraction by min(W,H) , where W and H denote the image width and height.
Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated screenshots, documents, and news pages prevalent in misinformation, remains untested. We introduce TextFake, a 20,000-image benchmark for text-rich AIGI detection spanning 28 languages, 4 topic categories, and 2 scene modalities. Fake images are synthesized via a four-stage pipeline that annotates real images along three controlled dimensions and generates counterparts through distribution-aligned structured prompting, ruling out covariate shortcuts. Zero-shot evaluation of 14 specialized detectors and 3 frontier VLM APIs reveals a large systematic gap: no method exceeds 80% accuracy, with some dropping over 60% from natural-image benchmarks. Diagnostic evaluations identify three failure modes: the Text Density Curse, where dense glyphs overwhelm low-level detectors; Cloaking via Rendering Fidelity, where stronger text rendering suppresses enerative artifacts; and Threshold Collapse, where routine perturbations drive detectors toward chance-level performance.
Yuning Zhang, Changtao Miao, Mingyu Liao +5
School of Cyber Science and Technology, University of Science and Technology of China · Anhui Province Key Laboratory of Digital Security · Individual Researcher
Typographic attacks pose a critical threat to vision-language models (VLMs) by injecting misleading text into images and causing models to rely on adversarial textual cues rather than visual evidence. Existing defenses often require model-specific modifications, additional training, or access to internal model components, limiting their applicability to modern closed-source VLMs. In this paper, we propose QuISE, a model-agnostic, training-free black-box defense based on query-irrelevant semantic editing. QuISE first identifies text regions likely to affect the current query through influence-aware text localization. QuISE then replaces these regions with two semantically distinct replacement texts that are irrelevant to both the query and the image. The final answer is determined by answer consistency across the edited images. Extensive experiments on three typographic-attack benchmarks, four attack settings, and four VLMs show that QuISE consistently improves defended accuracy. QuISE achieves a recovery rate of 67.9-75.0% with a harm rate of 0.5-1.1%.
Shubin Lu, Jiaqi Yin, Yihao Huang
School of Software, Northwestern Polytechnical University, Xi’an, China · Software Engineering Institute, East China Normal University, Shanghai, China
Recent methods demonstrate that large-scale pretrained models, such as CLIP vision transformers, effectively detect AI-generated images (AIGIs) from unseen generative models when used as feature extractors. Many state-of-the-art methods for AI-generated image detection build upon the original CLIP-ViT to enhance this generalization. Since CLIP's release, numerous vision foundation models (VFMs) have emerged, incorporating architectural improvements and different training paradigms. Despite these advances, their potential for AIGI detection and AI image forensics remains largely unexplored. In this work, we present a comprehensive benchmark across multiple VFM families, covering diverse pretraining objectives, input resolutions, and model scales. We systematically evaluate their out-of-the-box performance for detecting fully-generated AI-images and AI-inpainted images, and discover that the best model outperforms the original CLIP by more than 12% in accuracy, beating established approaches in the process. To fully leverage the features of a modern VFM, we propose a simple redesign of the classifier head by utilizing tunable attention pooling (TAP), which aggregates output tokens into a refined global representation. Integrating TAP with the latest VFMs yields substantial performance gains across several AIGI detection benchmarks, establishing a new state-of-the-art on two challenging benchmarks for in-the-wild detection of AI-generated and -inpainted images.