cs.CVOct 1, 2026

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

Authors: Xianda Du, Max Ku, Weiming Ren, Zhi Rui Tam, Chunlin Ren, Ping Nie, Min-Hung Chen, Wenhu Chen

Organizations: University of Waterloo · National Taiwan University · Nanyang Technological University · NVIDIA

Abstract

Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables directly verifiable post-training objectives. We train on 38K examples spanning score-only, localization-only, and joint supervision across generation and editing tasks. Starting from supervised fine-tuning, we further apply GRPO to improve defect localization using rewards that combine cell-level Dice overlap, score accuracy, and output-format validity. A parameter-free parser converts the structured predictions into readable explanations. On the primary suite, VIEScore2 achieves an overall-score SRCC of 0.601, compared with 0.491 for Gemini-3-Flash, the strongest zero-shot general-purpose VLM baseline under matched inputs. For defect localization, VIEScore2 outperforms both general-purpose VLMs and specialized spatial evaluators on three of six benchmarks in per-image grid IoU and ranks among the top three on five, including datasets beyond its training sources.

Figures & tables

Appendix figures & tables26 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Where, What, Why, and Importance: Structured Defect Grounding for Text-to-Image Feedback

    Jun 4, 2026Huaisong Zhang, Hao Yu, Yuxuan Zhang +7Text-To-ImageDiffusion Alignment

  2. Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models

    Apr 23, 2026Mohammed Safi Ur Rahman Khan, Sanjay Suryanarayanan, Tushar Anand +1Large Language Model EvaluationBlind Spots

  3. SalArt-VQA: Diagnosing Whether VLMs Understand Salient Artifacts in Generated Images

    Jun 10, 2026Xiaoxiao Sun, Ruotian Zhang, Junzhe Huang +2Ai-Generated Image DetectionTextvqa