Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.
Figures & tables
Supported edits
Unsupported
Evaluation
Benchmark / cohort
Answer preserved
Answer changed
Abstention target
Joint success
Causal VQA ( 2020 )
✓
✓
×
×
CertainlyUncertain ( 2025 )
×∗
×
✓
×
MM-AQA ( 2026 )
×
×
✓
×
VISREAS ( 2024 )
×
×
✓
×
HallusionBench ( 2024 )
✓
✓
׆
✓
Table 1: Components of controlled answering and abstention. ✓ : included in the evaluated protocol; × : not included under the definitions below. Our cohorts combine supported controls and abstention targets with complete-question scoring.
Figure 1: An executable reason to abstain. Two admissible complete charts answer the same question differently, but the fixed visibility operation produces identical observed pixels. The model receives the observation and question. The diagram redraws an archived proof for readability; verification uses the original renderer’s decoded RGBA output.
Figure 2: Keep the question fixed; change the evidence and required response. CLEVR illustrates the original, answer-preserving, answer-changing, and missing-evidence views. GQA contrasts masks outside and on program dependencies, including a residual location cue. Every view is presented independently; group success requires all supported answers and all required abstentions. Scene targets follow source programs and edits. Outlines and crops explain saved inputs; examples are selected independently of outcomes.
Source
Evaluated states
Evidence and coverage
PlotQA
5,000 views; F,S,C,M,I
Depth-one numerical lookup. M : admissible complete witnesses with identical pixels. F,S,C : execution and evidence visibility. I : selector precondition.
CLEVR
4,000 views; F,S,C,M
Native programs with 3–23 recorded steps; object-attribute interventions and official Blender images. Source-derived answerability labels.
GQA
3,000 views; F,S,M
Scene-graph programs with 2–4 recorded steps; dependency/nondependency box masks. Source-derived labels with residual-cue analysis.
Table 2: Evaluation cohorts and label evidence. Each source contributes 1,000 question groups. Chart missing-information labels have complete-world pixel witnesses; scene labels use source programs and interventions.
Configuration
J5↑
B5↑
EA↓
EM↓
EI↓
Qwen 3.8 Flash Next
57.00
83.50
1.63
11.30
2.80
Gemma 4 31B
55.60
66.70
1.77
28.50
4.20
Qwen 3.8 27B
34.10
66.50
3.70
23.90
5.00
Gemma 4 26B A4B
40.70
53.90
2.27
42.70
3.20
GLM 5.3 Flash
18.20
27.30
0.50
61.20
27.20
Molmo2-8B
6.10
19.40
0.03
53.80
56.40
Table 3: Complete chart task success and decision diagnostics (%). J5 requires all supported answers and abstentions; B5 checks decisions alone. All configurations use the same 1,000 groups. EA has 3,000 views; EM and EI each have 1,000. Invalid outputs count as failures.
Figure 3: Which obligation fails on each question? Each bar partitions 1,000 chart groups. The two left categories make every decision correctly; joint success additionally requires all supported answers. An M failure takes precedence when several decisions fail, and the final category has a correct M decision. Invalid outputs follow the same scoring rules.
PlotQA
CLEVR
GQA
System
J5↑
B5↑
E↓
J4↑
B4↑
E↓
J3↑
B3↑
E↓
Qwen 3.8 Flash Next
57.0
83.5
3.80
43.5
45.4
14.98
33.7
39.9
25.67
Gemma 4 31B
55.6
66.7
7.60
8.9
21.0
34.60
28.9
37.2
27.27
Qwen 3.8 27B
34.1
66.5
8.00
19.2
27.3
25.90
28.8
35.9
28.10
Gemma 4 26B A4B
40.7
53.9
10.54
5.4
18.9
34.63
30.0
39.2
27.73
GLM 5.3 Flash
18.2
27.3
17.98
17.0
20.5
19.93
24.5
30.5
24.70
Table 4: Complete task success across all three sources (%). Jk requires every supported answer and abstention; Bk requires every decision; E is per-view decision failure. Each source has 1,000 groups with its stated labels and states (Table 2 ). Invalid outputs count as failures.
Figure 4: State-level evaluation separates the direction of failure. Each cell contains 1,000 inputs; the shared scale is 0–100%, including final invalid outputs. Supported-state rejection and missing-state acceptance identify different response failures. Blank entries denote absent states.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Paper label
Model identifier
Qwen 3.8 Flash Next
Qwen/Qwen3.8-Flash-Next
Gemma 4 31B
google/gemma-4-31B-it
Qwen 3.8 27B
Qwen/Qwen3.8-27B-FP8
Gemma 4 26B A4B
google/gemma-4-26B-A4B-it
GLM 5.3 Flash
GLM-5.3-Flash
Molmo2-8B
allenai/Molmo2-8B
Appendix
Table 5: Model identifiers. All six configurations use vLLM ( Kwon et al., 2023 ) with temperature zero and a 512-token output cap.
Domain / system
W/V
Invalid
Failure 95% interval
PLOTQA / Qwen 3.8 Flash Next
190/5000
0
3.24–4.40
PLOTQA / Gemma 4 31B
380/5000
0
6.88–8.34
PLOTQA / Qwen 3.8 27B
400/5000
0
7.24–8.80
PLOTQA / Gemma 4 26B A4B
527/5000
0
9.74–11.38
PLOTQA / GLM 5.3 Flash
899/5000
0
17.16–18.82
PLOTQA / Molmo2-8B
1095/4992
8
21.20–22.92
Appendix
Table 6: Conditional-disagreement denominators, invalid counts, and full-denominator failure intervals. Intervals use source-question group bootstrap.
Table 8: Answer correctness and complete task success (%). A is correctness among all supported inputs, counting false rejection and invalid output as wrong. Jk requires all decisions and all supported answers in a group.
System
Loc. Fail
B3
M Fail
Other Fail
B3
M Fail
Qwen 3.8 Flash Next
24.91
37.2
52.39
26.12
41.5
38.62
Gemma 4 31B
27.48
37.2
47.87
27.14
37.2
45.35
Qwen 3.8 27B
27.39
33.8
51.06
28.53
37.2
38.62
Gemma 4 26B A4B
28.37
38.8
41.49
27.35
39.4
36.22
GLM 5.3 Flash
27.48
21.5
74.73
23.02
35.9
57.69
Molmo2-8B
26.33
24.7
70.74
32.59
27.9
49.84
Appendix
Table 9: GQA location-stratum sensitivity (%). Location: 376 groups; Other: 624 groups. M Fail uses the missing-information state. Both strata retain their source-derived labels and all sibling states.
Source
Views/group
Train groups
Validation groups
Primary groups
PlotQA
5
3000
–
1000
CLEVR
4
4000
1000
–
GQA
3
5000
1000
–
Appendix
Table 10: Resource composition. Each group contains every state specified for its source.
Metric
Qwen 3.8 Flash Next
GLM 5.3 Flash
Δ (pp)
Paired 95% interval
View failure E↓
25.67
24.70
+0.97
[−0.47,+2.37]
Answerable failure EA↓
16.60
5.00
+11.60
[+9.60,+13.60]
Unanswerable failure EU↓
43.80
64.10
−20.30
[−23.30,−17.30]
Balanced failure Ebal↓
30.20
34.55
−4.35
[−5.88,−2.85]
All-state success B3↑
39.90
30.50
+9.40
[+6.50,+12.30]
Joint answer success J3↑
33.70
24.50
+9.20
[+6.50,+12.00]
Appendix
Table 11: Paired GQA diagnostics on the same 1,000 groups. Scores are percentages and Δ is Qwen 3.8 Flash Next minus GLM 5.3 Flash in percentage points. Intervals use paired group bootstrap.
Figure 5: Class weights and error concentration explain different GQA rankings. (A) Changing the class weight changes the point ranking while predictions remain fixed. (B) Qwen 3.8 Flash Next has more total failed views but fewer affected groups. Table 11 gives intervals for score differences.
Source / system
0
1
2
3
4
5
PLOTQA / Qwen 3.8 Flash Next
835
147
11
7
0
0
PLOTQA / Gemma 4 31B
667
300
19
14
0
0
PLOTQA / Qwen 3.8 27B
665
283
39
13
0
0
PLOTQA / Gemma 4 26B A4B
539
419
22
16
4
0
PLOTQA / GLM 5.3 Flash
273
558
167
1
1
0
PLOTQA / Molmo2-8B
194
509
297
0
0
0
Appendix
Table 12: Question groups with 0–5 failed states across all 18 source/configuration cells. Each row sums to 1,000; weighting counts by failure multiplicity recovers the per-view failure numerator. A dash denotes an impossible state count.
Chart question-answering (QA) benchmarks aim to pose questions that require visual reasoning to correctly answer, but models can often reach solutions through shortcuts or prior familiarity with a chart based on their own background knowledge. To strictly evaluate visual reasoning, we propose counterfactual charts where the chart-question task remains fixed, but underlying chart and the corresponding answer are varied. We introduce Chartographer, a framework to reverse engineer charts into executable code, validate reconstruction fidelity, generate seed-controlled counterfactual variants, and derive new answers from executable QA logic. We apply this framework to existing chart QA datasets and evaluate proprietary and open-source vision-language models (VLMs), measuring variation sensitivity and generalizability. Counterfactual charts reveal failures hidden by single-chart performance: VLMs often fail to generalize after answering the original chart correctly. We find failures are most prevalent when updated charts require novel visual reasoning pathways.
Yifan Jiang, Dae Yon Hwang, Jesse C. Cresswell +1
University of Waterloo · Vector Institute · Layer 6 AI
Vision-language models (VLMs) increasingly power consumer-facing AI search, yet evaluating them on the diversity of everyday visual questions remains challenging. Existing benchmarks often target predefined capabilities, such as multi-hop retrieval or long-form synthesis, whereas users ask photo-grounded questions spanning a long tail of everyday scenarios. Despite advances in VLMs, users on Xiaohongshu, a mainstream Chinese image-sharing platform, continue to turn to other people for help with everyday visual questions. Motivated by this behaviour, we curate NoteVQA from these questions, yielding 252 items across 12 topical categories and 7 user intents. Each item includes a concise reference distilled from expert community responses and a human-audited interleaved reference answer that combines textual explanations with supporting visual evidence. We evaluate both short-answer correctness and interleaved-answer quality. To support the latter, we introduce AgenticInterleave, a single-agent ReAct framework for retrieval-supported answer generation, together with IVR-12, a 12-dimensional rubric for assessing the content, presentation, and image quality of interleaved references and model outputs. Across 10 frontier VLMs, the highest short-answer accuracy is 52.8%, while adding agentic search to Qwen3.5-397B-A17B improves accuracy by only 2.0%. For interleaved answers, the same model running AgenticInterleave scores 3.52 under IVR-12, compared with 4.65 for the human-audited references, with the largest gap in content quality. These results highlight the challenges that everyday visual questions pose for current VLMs in both answer accuracy and the quality of visually grounded explanations.
Establishing a clear link between model predictions and the visual evidence that supports them is critical for transparency and reliability in multimodal reasoning, yet current multimodal large language model (MLLM) evaluations do not explicitly enforce this alignment. Existing benchmarks assess either textual answer correctness or pixel-level localization in isolation, leaving the coupling of reasoning and grounding an open challenge. We introduce VISTAQA, a comprehensive benchmark for joint evaluation of free-form answer correctness and pixel-level evidence grounding in visual question answering. VISTAQA comprises 1,157 expert-curated samples spanning six task types and six visual domains, ranging from direct perception to compositional and relational reasoning. VISTAQA requires models to not only answer correctly, but to also provide precise segmentation masks that support their answers. It also includes hallucination-aware examples where no valid visual evidence exists. To support this enhanced evaluation, we introduce GROVE, a unified evaluation metric that enforces joint correctness by combining textual accuracy and grounding quality via a per-sample geometric mean, ensuring neither dimension can compensate for deficiencies in the other. Comprehensive experiments across grounding-aware models and hybrid pipelines with general-purpose MLLMs reveal that even the strongest systems achieve limited performance under GROVE, highlighting a substantial gap between answer accuracy and visual evidence alignment.
Mozhgan Nasr Azadani, Yimu Wang, Yongpeng Zhu +5
University of Waterloo · Stanford University · NVIDIA