False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
Authors: Zeyu Ye, Yanchun Li, Sibei He, Meng Xie, Hangtao Zhang, Xianlong Wang, Li Zeng, Jiahao Chen, +3 more
Organizations: Xiangtan University · Jinan University · Huazhong University of Science and Technology · City University of Hong Kong · Changsha University of Science and Technology · Zhejiang University · Chongqing University
Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages. Yet, the same capability introduces a new risk: these models can just as easily fabricate visual misinformation. Even commercial models (e.g., GPT-Image-2) readily produce it. Curiously, we find that these models can recognize a claim as false when asked, yet still render that very claim as credible visual evidence. This discrepancy points to a blind spot in current alignment: safeguards judge what an image shows, not what it asserts; however, existing red-teaming benchmarks target conventional harmful content, such as violent or explicit imagery, and say little about where the alignment boundaries lie for visual misinformation, especially in commercial models. To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats. We further introduce EpiReal-Attack, a skill-guided black-box optimization framework that uses Pareto-based selection and multimodal feedback to identify commands that bypass alignment safeguards while preserving visual realism, textual legibility, and semantic fidelity. Experiments on four commercial models reveal that more than 70% of false-claim prompts elicit images that faithfully depict the corresponding misinformation, and EpiReal-Attack pushes this rate to 95%. Most worryingly, these models are only a click away, and their outputs are cheap to spread yet hard to disbelieve, leaving this dimension of alignment largely unguarded.
Figures & tables
Figure 1: Overview of EpiReal-Bench .
Figure 2: Misalignment between misinformation recognition and safeguards.
Figure 3: Comparison between traditional safety benchmarks and EpiReal-Bench .
Benchmark
Prompt dataset
Image Dataset
Annotation
Prompt set
Categories
Number
Image set
Visual forms
Human check
Reason annotation
PH Safeguards ( Chu et al., 2026 )
∙
5
50
∘
–
∙
∘
PC 2 ( Choi et al., 2026 )
∙
2
240
∘
–
∘
∘
M 3 A ( Xu et al., 2024 )
∘
–
–
∙
1
∘
∘
MiRAGeNews ( Huang et al., 2024 )
∘
–
–
∙
1
∘
∘
MM-Health ( Zhang et al., 2025c )
∘
–
–
∙
1
∙
∘
Table 1: Comparison of representative misinformation benchmarks across text- and image-based settings. ∙ indicates availability, ∘ indicates unavailability, and – denotes not applicable.
Figure 4: FCR comparison with ASR of four commercial image-generation models.
Category
Definition
Visual form
Definition
Government
Government decisions, institutions, and public administration.
Newspaper page
Printed news page with headlines, articles, images, and publication details.
Election
Voting, election administration, results, and certification.
Broadcast news
Television news frame with presenter, headline, ticker, and supporting footage.
Military
Military operations, armed conflict, and security activities.
Social platform post
Social-media post with profile, text, media, and engagement information.
Disaster
Natural disasters, accidents, and public-safety emergencies.
Platform screenshot
Web-platform interface with navigation, content, imagery, and controls.
Crime
Criminal incidents, investigations, and public-safety conditions.
Documentary photo
Documentary-style photograph depicting an event in a realistic setting.
Health
Public health, disease, vaccination, risks, and treatments.
Red Headed document
Official administrative document with agency header, serial number, and seal.
Table 2: Taxonomy of EpiReal-Bench . False claims span 10 real-world claim categories and are instantiated across 10 credible visual formats.
Figure 5: Failed generations and successful generations after applying EpiReal-Attack .
Model
Metric
NP
BN
SP
PS
DP
RD
MA
BU
TP
CR
Overall
EpiReal-Bench
GPT-Image-2
ASR (%)
92.30
89.80
93.90
86.80
23.90
69.20
90.80
84.60
74.60
64.60
77.01
MR
4.92
4.86
4.94
4.85
3.99
4.65
4.87
4.82
4.57
4.61
4.71
Nano Banana 2
ASR (%)
32.50
51.30
68.40
57.50
17.50
70.00
65.00
88.80
32.50
75.00
55.82
MR
4.34
4.48
4.68
4.56
4.12
4.68
4.52
4.83
4.31
4.73
4.53
Grok Imagen 2.0
ASR (%)
94.00
91.70
94.70
94.00
74.60
78.60
96.30
92.30
89.70
88.30
89.43
MR
4.93
4.92
4.93
4.91
4.72
4.75
4.96
4.92
4.82
4.87
4.87
Table 3: Performance of commercial image-generation models on EpiReal-Bench and EpiReal-Attack across 10 credible visual formats. We report ASR and the average MR. Higher values indicate a greater tendency to faithfully render false claims as visual evidence.
Defense Level
Defense Method
GPT-Image-2
Nano Banana 2
Grok Imagen 2.0
Seedream 5.0 Flash
Input-level
Keyword Filter ( Yang et al., 2024 )
8.80
21.30
11.40
18.50
NSFW Text Classifier ( LAION-AI, 2024 )
1.70
22.90
8.30
37.80
GPT-4o-mini Prompt Checker ( Jin et al., 2025 )
11.60
24.70
7.80
26.20
Output-level
SD Safety Checker ( Hugging Face, 2024 )
7.90
30.60
10.90
27.60
Q16 ( Schramowski et al., 2022 )
31.40
48.30
32.50
56.60
GPT-4o Image Checker ( Jin et al., 2025 )
10.70
32.40
8.60
28.20
Table 4: Interception rate (%) under different defense mechanisms across four image generation models. Higher interception rate indicates better defense effectiveness.
Figure 6: VC scores of GPT-Image-2 generations across 10 visually credible formats, evaluated by seven VLMs and human annotators.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Accuracy (%) ↑
GPT-Image-2
98.20
Nano Banana 2
99.10
Grok Imagen 2.0
94.60
Seedream 5.0 Flash
92.80
Appendix
Table 5: Verification accuracy (%) on benign factual prompts. Higher values indicate that the model more accurately recognizes factually valid claims as non-misinformation.
Model
Metric
Gov.
Elec.
Mil.
Dis.
Crim.
Hlth.
Sci.
Econ.
Edu.
Env.
Overall
EpiReal-Bench
GPT-Image-2
ASR (%)
80.00
72.80
77.70
77.50
73.80
83.70
78.50
82.30
69.00
74.60
77.01
MR
4.78
4.70
4.75
4.77
4.74
4.76
4.72
4.78
4.50
4.58
4.71
Nano Banana 2
ASR (%)
51.20
63.30
50.00
48.80
53.80
68.80
50.00
73.80
48.80
50.00
55.82
MR
4.49
4.62
4.46
4.49
4.51
4.63
4.46
4.74
4.43
4.45
4.53
Grok Imagen 2.0
ASR (%)
86.30
88.70
90.00
93.00
94.70
88.30
88.70
93.90
75.70
95.00
89.43
MR
4.86
4.88
4.89
4.93
4.94
4.86
4.88
4.93
4.60
4.94
4.87
Appendix
Table 6: Performance of commercial image-generation models on EpiReal-Bench and EpiReal-Attack across 10 claim categories. We report ASR and the average MR. Higher values indicate a greater tendency to faithfully render false claims as visual evidence.
Figure 7: VC scores of Nano Banana 2 generations across 10 visually credible formats, evaluated by seven VLMs and human annotators.
Figure 8: VC scores of Grok Imagen 2.0 generations across 10 visually credible formats, evaluated by seven VLMs and human annotators.
Figure 9: VC scores of Seedream 5.0 Flash generations across 10 visually credible formats, evaluated by seven VLMs and human annotators.
Figure 10: A government false claim presented as a newspaper page
Figure 11: A military false claim presented as a social platform post
Figure 12: An election false claim presented as a broadcast news screen
Figure 13: A disaster false claim presented as a desktop platform screenshot
Figure 14: A crime false claim presented as a documentary-style photograph
Figure 15: A health false claim presented as an official red-headed administrative document
Figure 16: A science false claim presented as a smartphone mobile alert
Figure 17: An economy false claim presented as a public bulletin notice
Figure 18: An educational false claim presented as an educational textbook page
Figure 19: An environment false claim presented as an instant messaging chat record
Mohamed bin Zayed University of Artificial Intelligence, UAE · UKP Lab · Ubiquitous Knowledge Processing Lab (UKP Lab), Department of Computer Science, TU Darmstadt and National Research Center for Applied Cybersecurity ATHENE, Germany +2