cs.CVOct 8, 2026

WorldFact-Bench: Beyond Image-Internal Plausibility to Image-World Consistency

Authors: Zhuohong Chen, Zhengxian Wu, Yunyao Yu, Hangrui Xu, Zijian Yu, Hao Tan, Zhifang Liu, Peng Jiao, +2 more

Organizations: Ant Group · Tsinghua University

Abstract

Advances in image generation have made visual authenticity increasingly difficult to assess. Although image forensics now examines both generation artifacts and higher-level visual inconsistencies, a plausible image can still contradict real-world facts or rules. We introduce WorldFact-Bench to evaluate image-world consistency from a single image, without a predefined claim or verification target. The benchmark contains 1,274 source-aligned real-fake pairs across four verification regimes and ten semantic domains. Each pair introduces a specific, evidence-supported factual conflict while seeking to preserve non-target content and visual plausibility. Images are evaluated independently, and pair accuracy requires both members of a pair to be classified correctly. We further propose PERSIST-Agent, which organizes iterative verification around a persistent state linking candidate facts, visual observations, evidence, and verification statuses. This state guides subsequent inspection and retrieval while retaining unresolved candidates. With backbone weights fixed, harness self-optimization refines the agent's prompts and execution rules through validation feedback. Experiments reveal strong label biases in several detectors and uneven gains from retrieval. On the evaluated 8B backbones, PERSIST-Agent improves pair accuracy over both direct judgment and retrieval-augmented baselines, while ablations support the role of persistent verification state. These findings highlight the value of state-guided verification and the remaining gap between visual plausibility and factual correctness.

Explore similar work

CardsList
  1. FAGER: Factually Grounded Evaluation and Refinement of Text-to-Image Models

    May 18, 2026Youngsun Lim, Cusuh Ham, Pin-Yu Chen +1Factual Consistency EvaluationAutomated Evaluation

  2. MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation

    Sep 27, 2026Ruihong Zeng, Jonathan Tonglet, Preslav Nakov +1Factual Consistency EvaluationAutomated Fact-Checking

  3. False Claims, Credible Images: A Red-Teaming Benchmark for Commercial Image Generators

    Oct 8, 2026Zeyu Ye, Yanchun Li, Sibei He +8Safe T2I GenerationT2I Generation Evaluation