cs.CVSep 28, 2026

When Does an Image Determine the Answer? Benchmarking Visual Answerability across Charts and Scenes

Authors: Sungguk Cha, Mintae Kim, Youngsub Han, Byoung-Ki Jeon, Sangyeob Lee

Organizations: LG Uplus, Seoul, South Korea

Abstract

Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.

Figures & tables

Appendix figures & tables9 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models

    May 26, 2026Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell +1ChartVisual Reasoning

  2. NoteVQA: Benchmarking VLMs on Real-Life Questions from Human Communities

    Sep 14, 2026Haonan Jiang, Guojian Zhan, Jiancong Xie +6TextvqaQuestion

  3. VISTAQA: Benchmarking Joint Visual Question Answering and Pixel-Level Evidence

    May 20, 2026Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu +5Weak Visual GroundingVisual Evidence