Reliable visual question answering requires correct answers when evidence is sufficient and abstention when it is not. We introduce a benchmark that connects complete-question evaluation with explicit evidence for its labels across PlotQA charts, CLEVR rendered scenes, and GQA photographs. Each question groups original and edited images, presented independently; success requires every supported answer and every required abstention to be correct. For chart missing-information labels, executable witnesses establish that admissible complete charts give different answers but identical pixels after masking. Scene labels follow source programs and edits, with a residual-cue analysis for photographs. Across 72,000 responses from six model configurations, the highest observed complete task success rates are 57.0%, 43.5%, and 33.7%, respectively. On charts, the strongest configuration achieves 96.2% per-view decision accuracy, yet 265 of its 835 groups with every decision correct still contain incorrect answers. Evaluating supported answers and necessary abstentions together exposes failures that answerability decisions alone conceal.
Figures & tables
Supported edits
Unsupported
Evaluation
Benchmark / cohort
Answer preserved
Answer changed
Abstention target
Joint success
Causal VQA ( 2020 )
✓
✓
×
×
CertainlyUncertain ( 2025 )
×∗
×
✓
×
MM-AQA ( 2026 )
×
×
✓
×
VISREAS ( 2024 )
×
×
✓
×
HallusionBench ( 2024 )
✓
✓
׆
✓
Table 1: Components of controlled answering and abstention. ✓ : included in the evaluated protocol; × : not included under the definitions below. Our cohorts combine supported controls and abstention targets with complete-question scoring.
Figure 1: An executable reason to abstain. Two admissible complete charts answer the same question differently, but the fixed visibility operation produces identical observed pixels. The model receives the observation and question. The diagram redraws an archived proof for readability; verification uses the original renderer’s decoded RGBA output.
Figure 2: Keep the question fixed; change the evidence and required response. CLEVR illustrates the original, answer-preserving, answer-changing, and missing-evidence views. GQA contrasts masks outside and on program dependencies, including a residual location cue. Every view is presented independently; group success requires all supported answers and all required abstentions. Scene targets follow source programs and edits. Outlines and crops explain saved inputs; examples are selected independently of outcomes.
Source
Evaluated states
Evidence and coverage
PlotQA
5,000 views; F,S,C,M,I
Depth-one numerical lookup. M : admissible complete witnesses with identical pixels. F,S,C : execution and evidence visibility. I : selector precondition.
CLEVR
4,000 views; F,S,C,M
Native programs with 3–23 recorded steps; object-attribute interventions and official Blender images. Source-derived answerability labels.
GQA
3,000 views; F,S,M
Scene-graph programs with 2–4 recorded steps; dependency/nondependency box masks. Source-derived labels with residual-cue analysis.
Table 2: Evaluation cohorts and label evidence. Each source contributes 1,000 question groups. Chart missing-information labels have complete-world pixel witnesses; scene labels use source programs and interventions.
Configuration
J5↑
B5↑
EA↓
EM↓
EI↓
Qwen 3.8 Flash Next
57.00
83.50
1.63
11.30
2.80
Gemma 4 31B
55.60
66.70
1.77
28.50
4.20
Qwen 3.8 27B
34.10
66.50
3.70
23.90
5.00
Gemma 4 26B A4B
40.70
53.90
2.27
42.70
3.20
GLM 5.3 Flash
18.20
27.30
0.50
61.20
27.20
Molmo2-8B
6.10
19.40
0.03
53.80
56.40
Table 3: Complete chart task success and decision diagnostics (%). J5 requires all supported answers and abstentions; B5 checks decisions alone. All configurations use the same 1,000 groups. EA has 3,000 views; EM and EI each have 1,000. Invalid outputs count as failures.
Figure 3: Which obligation fails on each question? Each bar partitions 1,000 chart groups. The two left categories make every decision correctly; joint success additionally requires all supported answers. An M failure takes precedence when several decisions fail, and the final category has a correct M decision. Invalid outputs follow the same scoring rules.
PlotQA
CLEVR
GQA
System
J5↑
B5↑
E↓
J4↑
B4↑
E↓
J3↑
B3↑
E↓
Qwen 3.8 Flash Next
57.0
83.5
3.80
43.5
45.4
14.98
33.7
39.9
25.67
Gemma 4 31B
55.6
66.7
7.60
8.9
21.0
34.60
28.9
37.2
27.27
Qwen 3.8 27B
34.1
66.5
8.00
19.2
27.3
25.90
28.8
35.9
28.10
Gemma 4 26B A4B
40.7
53.9
10.54
5.4
18.9
34.63
30.0
39.2
27.73
GLM 5.3 Flash
18.2
27.3
17.98
17.0
20.5
19.93
24.5
30.5
24.70
Table 4: Complete task success across all three sources (%). Jk requires every supported answer and abstention; Bk requires every decision; E is per-view decision failure. Each source has 1,000 groups with its stated labels and states (Table 2 ). Invalid outputs count as failures.
Figure 4: State-level evaluation separates the direction of failure. Each cell contains 1,000 inputs; the shared scale is 0–100%, including final invalid outputs. Supported-state rejection and missing-state acceptance identify different response failures. Blank entries denote absent states.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Paper label
Model identifier
Qwen 3.8 Flash Next
Qwen/Qwen3.8-Flash-Next
Gemma 4 31B
google/gemma-4-31B-it
Qwen 3.8 27B
Qwen/Qwen3.8-27B-FP8
Gemma 4 26B A4B
google/gemma-4-26B-A4B-it
GLM 5.3 Flash
GLM-5.3-Flash
Molmo2-8B
allenai/Molmo2-8B
Appendix
Table 5: Model identifiers. All six configurations use vLLM ( Kwon et al., 2023 ) with temperature zero and a 512-token output cap.
Domain / system
W/V
Invalid
Failure 95% interval
PLOTQA / Qwen 3.8 Flash Next
190/5000
0
3.24–4.40
PLOTQA / Gemma 4 31B
380/5000
0
6.88–8.34
PLOTQA / Qwen 3.8 27B
400/5000
0
7.24–8.80
PLOTQA / Gemma 4 26B A4B
527/5000
0
9.74–11.38
PLOTQA / GLM 5.3 Flash
899/5000
0
17.16–18.82
PLOTQA / Molmo2-8B
1095/4992
8
21.20–22.92
Appendix
Table 6: Conditional-disagreement denominators, invalid counts, and full-denominator failure intervals. Intervals use source-question group bootstrap.
Table 8: Answer correctness and complete task success (%). A is correctness among all supported inputs, counting false rejection and invalid output as wrong. Jk requires all decisions and all supported answers in a group.
System
Loc. Fail
B3
M Fail
Other Fail
B3
M Fail
Qwen 3.8 Flash Next
24.91
37.2
52.39
26.12
41.5
38.62
Gemma 4 31B
27.48
37.2
47.87
27.14
37.2
45.35
Qwen 3.8 27B
27.39
33.8
51.06
28.53
37.2
38.62
Gemma 4 26B A4B
28.37
38.8
41.49
27.35
39.4
36.22
GLM 5.3 Flash
27.48
21.5
74.73
23.02
35.9
57.69
Molmo2-8B
26.33
24.7
70.74
32.59
27.9
49.84
Appendix
Table 9: GQA location-stratum sensitivity (%). Location: 376 groups; Other: 624 groups. M Fail uses the missing-information state. Both strata retain their source-derived labels and all sibling states.
Source
Views/group
Train groups
Validation groups
Primary groups
PlotQA
5
3000
–
1000
CLEVR
4
4000
1000
–
GQA
3
5000
1000
–
Appendix
Table 10: Resource composition. Each group contains every state specified for its source.
Metric
Qwen 3.8 Flash Next
GLM 5.3 Flash
Δ (pp)
Paired 95% interval
View failure E↓
25.67
24.70
+0.97
[−0.47,+2.37]
Answerable failure EA↓
16.60
5.00
+11.60
[+9.60,+13.60]
Unanswerable failure EU↓
43.80
64.10
−20.30
[−23.30,−17.30]
Balanced failure Ebal↓
30.20
34.55
−4.35
[−5.88,−2.85]
All-state success B3↑
39.90
30.50
+9.40
[+6.50,+12.30]
Joint answer success J3↑
33.70
24.50
+9.20
[+6.50,+12.00]
Appendix
Table 11: Paired GQA diagnostics on the same 1,000 groups. Scores are percentages and Δ is Qwen 3.8 Flash Next minus GLM 5.3 Flash in percentage points. Intervals use paired group bootstrap.
Figure 5: Class weights and error concentration explain different GQA rankings. (A) Changing the class weight changes the point ranking while predictions remain fixed. (B) Qwen 3.8 Flash Next has more total failed views but fewer affected groups. Table 11 gives intervals for score differences.
Source / system
0
1
2
3
4
5
PLOTQA / Qwen 3.8 Flash Next
835
147
11
7
0
0
PLOTQA / Gemma 4 31B
667
300
19
14
0
0
PLOTQA / Qwen 3.8 27B
665
283
39
13
0
0
PLOTQA / Gemma 4 26B A4B
539
419
22
16
4
0
PLOTQA / GLM 5.3 Flash
273
558
167
1
1
0
PLOTQA / Molmo2-8B
194
509
297
0
0
0
Appendix
Table 12: Question groups with 0–5 failed states across all 18 source/configuration cells. Each row sums to 1,000; weighting counts by failure multiplicity recovers the per-view failure numerator. A dash denotes an impossible state count.