False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators
Organizations: Xiangtan University · Jinan University · Huazhong University of Science and Technology · City University of Hong Kong · Changsha University of Science and Technology · Zhejiang University · Chongqing University
Abstract
Image-generation models can now produce text-rich, natural-looking visual artifacts that are hard to distinguish from real-world evidence, such as news reports and textbook pages. Yet, the same capability introduces a new risk: these models can just as easily fabricate visual misinformation. Even commercial models (e.g., GPT-Image-2) readily produce it. Curiously, we find that these models can recognize a claim as false when asked, yet still render that very claim as credible visual evidence. This discrepancy points to a blind spot in current alignment: safeguards judge what an image shows, not what it asserts; however, existing red-teaming benchmarks target conventional harmful content, such as violent or explicit imagery, and say little about where the alignment boundaries lie for visual misinformation, especially in commercial models. To fill this gap, we introduce EpiReal-Bench, the first systematic benchmark for evaluating visual misinformation risks in commercial image generators, comprising 10k false-claim prompts and 10k corresponding generated images that span 10 real-world claim categories and 10 credible visual formats. We further introduce EpiReal-Attack, a skill-guided black-box optimization framework that uses Pareto-based selection and multimodal feedback to identify commands that bypass alignment safeguards while preserving visual realism, textual legibility, and semantic fidelity. Experiments on four commercial models reveal that more than 70% of false-claim prompts elicit images that faithfully depict the corresponding misinformation, and EpiReal-Attack pushes this rate to 95%. Most worryingly, these models are only a click away, and their outputs are cheap to spread yet hard to disbelieve, leaving this dimension of alignment largely unguarded.
Figures & tables
| Benchmark | Prompt dataset | Image Dataset | Annotation | ||||
| Prompt set | Categories | Number | Image set | Visual forms | Human check | Reason annotation | |
| PH Safeguards ( Chu et al., 2026 ) | 5 | 50 | – | ||||
| PC 2 ( Choi et al., 2026 ) | 2 | 240 | – | ||||
| M 3 A ( Xu et al., 2024 ) | – | – | 1 | ||||
| MiRAGeNews ( Huang et al., 2024 ) | – | – | 1 | ||||
| MM-Health ( Zhang et al., 2025c ) | – | – | 1 | ||||
| Category | Definition | Visual form | Definition |
| Government | Government decisions, institutions, and public administration. | Newspaper page | Printed news page with headlines, articles, images, and publication details. |
| Election | Voting, election administration, results, and certification. | Broadcast news | Television news frame with presenter, headline, ticker, and supporting footage. |
| Military | Military operations, armed conflict, and security activities. | Social platform post | Social-media post with profile, text, media, and engagement information. |
| Disaster | Natural disasters, accidents, and public-safety emergencies. | Platform screenshot | Web-platform interface with navigation, content, imagery, and controls. |
| Crime | Criminal incidents, investigations, and public-safety conditions. | Documentary photo | Documentary-style photograph depicting an event in a realistic setting. |
| Health | Public health, disease, vaccination, risks, and treatments. | Red Headed document | Official administrative document with agency header, serial number, and seal. |
| Model | Metric | NP | BN | SP | PS | DP | RD | MA | BU | TP | CR | Overall | |
| EpiReal-Bench | GPT-Image-2 | ASR (%) | 92.30 | 89.80 | 93.90 | 86.80 | 23.90 | 69.20 | 90.80 | 84.60 | 74.60 | 64.60 | 77.01 |
| MR | 4.92 | 4.86 | 4.94 | 4.85 | 3.99 | 4.65 | 4.87 | 4.82 | 4.57 | 4.61 | 4.71 | ||
| Nano Banana 2 | ASR (%) | 32.50 | 51.30 | 68.40 | 57.50 | 17.50 | 70.00 | 65.00 | 88.80 | 32.50 | 75.00 | 55.82 | |
| MR | 4.34 | 4.48 | 4.68 | 4.56 | 4.12 | 4.68 | 4.52 | 4.83 | 4.31 | 4.73 | 4.53 | ||
| Grok Imagen 2.0 | ASR (%) | 94.00 | 91.70 | 94.70 | 94.00 | 74.60 | 78.60 | 96.30 | 92.30 | 89.70 | 88.30 | 89.43 | |
| MR | 4.93 | 4.92 | 4.93 | 4.91 | 4.72 | 4.75 | 4.96 | 4.92 | 4.82 | 4.87 | 4.87 |
| Defense Level | Defense Method | GPT-Image-2 | Nano Banana 2 | Grok Imagen 2.0 | Seedream 5.0 Flash |
| Input-level | Keyword Filter ( Yang et al., 2024 ) | 8.80 | 21.30 | 11.40 | 18.50 |
| NSFW Text Classifier ( LAION-AI, 2024 ) | 1.70 | 22.90 | 8.30 | 37.80 | |
| GPT-4o-mini Prompt Checker ( Jin et al., 2025 ) | 11.60 | 24.70 | 7.80 | 26.20 | |
| Output-level | SD Safety Checker ( Hugging Face, 2024 ) | 7.90 | 30.60 | 10.90 | 27.60 |
| Q16 ( Schramowski et al., 2022 ) | 31.40 | 48.30 | 32.50 | 56.60 | |
| GPT-4o Image Checker ( Jin et al., 2025 ) | 10.70 | 32.40 | 8.60 | 28.20 |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Accuracy (%) |
| GPT-Image-2 | 98.20 |
| Nano Banana 2 | 99.10 |
| Grok Imagen 2.0 | 94.60 |
| Seedream 5.0 Flash | 92.80 |
| Model | Metric | Gov. | Elec. | Mil. | Dis. | Crim. | Hlth. | Sci. | Econ. | Edu. | Env. | Overall | |
| EpiReal-Bench | GPT-Image-2 | ASR (%) | 80.00 | 72.80 | 77.70 | 77.50 | 73.80 | 83.70 | 78.50 | 82.30 | 69.00 | 74.60 | 77.01 |
| MR | 4.78 | 4.70 | 4.75 | 4.77 | 4.74 | 4.76 | 4.72 | 4.78 | 4.50 | 4.58 | 4.71 | ||
| Nano Banana 2 | ASR (%) | 51.20 | 63.30 | 50.00 | 48.80 | 53.80 | 68.80 | 50.00 | 73.80 | 48.80 | 50.00 | 55.82 | |
| MR | 4.49 | 4.62 | 4.46 | 4.49 | 4.51 | 4.63 | 4.46 | 4.74 | 4.43 | 4.45 | 4.53 | ||
| Grok Imagen 2.0 | ASR (%) | 86.30 | 88.70 | 90.00 | 93.00 | 94.70 | 88.30 | 88.70 | 93.90 | 75.70 | 95.00 | 89.43 | |
| MR | 4.86 | 4.88 | 4.89 | 4.93 | 4.94 | 4.86 | 4.88 | 4.93 | 4.60 | 4.94 | 4.87 |