VisionQ: VLM-as-a-Judge Taxonomy, Dataset and Benchmark for Qualitative Analysis in Computer Vision
Authors: Vu Dinh Xuan, Duc-Hai Nguyen, Minh-Dung Dao, Vu Quynh Giao, Quang Hong Nguyen, Binh-Son Hua, Barry O'Sullivan, David Murphy, +1 more
Organizations: University of Information Technology, VNU-HCM, Vietnam · University College Cork, Ireland · Hanoi University of Science and Technology, Vietnam · Trinity College Dublin, Ireland
Qualitative comparison figures are central evidence in computer vision papers, and vision-language models (VLMs) are increasingly used to judge them. Yet existing benchmarks score only scalar quality or overall preference, so a judge can be rewarded for picking the preferred image for the wrong visual reason. We introduce VisionQ, the first benchmark built from peer-reviewed CV comparison figures that grounds every judgment in a named visual criterion: each question states the criterion, and a judge is credited only when it selects the output the authors identify as best on that criterion. We call this task criterion-conditioned visual discrimination. VisionQ comprises (1) a corpus of 1,409 CVPR and ICCV papers with 1,800+ validated comparison figures and 3,911 hand-annotated data points linking method crops to author-stated visual claims; (2) a six-axis, 51-leaf taxonomy of the visual criteria behind qualitative judgment; (3) a criterion-conditioned evaluation protocol that hides method names, captions, and paper identity and reports accuracy per criterion; and (4) VisionQ-Judge, a DPO-tuned Gemma-4-E4B judge trained on symmetric evidence pairs, which reduces last-option predictions by 7.0pp and improves accuracy by 2.5pp on a held-out test set. Evaluating 20 open- and closed-source VLM judges, we find that the strongest reach only 63.1% accuracy (chance 32.2%) and that reliability varies sharply across criteria. Code: https://github.com/ReML-AI/visionq. Data: https://huggingface.co/datasets/visionq-anon-2026/VisionQ-1k.
Figures & tables
Figure 1 : Overview of VisionQ. Five-phase pipeline: corpus construction and figure selection (Section 3 ), taxonomy-guided annotation (Section 2 ), benchmark protocols (Section 4 ), VLM judge evaluation (Section 4 ), and VisionQ-Judge DPO training (Section 5 , bottom strip).
Figure 2 : VisionQ six-axis taxonomy. Each axis branches into leaf nodes to which author-stated claims are assigned. Claims that don’t fit any leaf are assigned to an out-of-scope category with an explicit reason code.
Figure 3 : How the VisionQ datasets relate: (1) the source corpus, (2) claims, taxonomy, and annotation, (3) the datasets built from the annotated pool, and (4) where each is used. Every population derives from the 1,409-paper source corpus. Claims extracted from the whole corpus define the taxonomy, which labels the annotated pool. The annotated pool supplies the in-scope data used for the coverage analysis and the two question sets: VisionQ-Bench, used to evaluate VLM judges, and VisionQ-MCQ, used to train and test VisionQ-Judge. Table 1 lists the exact sizes.
Population
Papers
Figures
Data points
Questions
Used for
Source corpus
1,409
Taxonomy; 9,228 claims (§ 2.1 )
Candidate figures
3,651
Figure selection (§ 3.1 )
Comparison figures
1,800+
Annotation (§ 3.2 )
Box-annotated papers
1,403
Coverage denominator (§ 3.4 )
Annotated pool
810
1,486
3,911
Coverage check (§ 2.3 )
In scope
790
1,426
3,773
Coverage analysis (§ 3.4 – 3.5 )
Table 1 : VisionQ populations at a glance. A data point is one annotated comparison row of a figure together with its author-stated claim; a question is one multiple-choice item.
Figure 4 : How many of their field’s 10 most commonly applied criteria the 789 selective-comparison candidates test in their qualitative comparisons. Most papers (59%) test exactly one.
Figure 5 : An example VisionQ-Bench question, exactly as the judges see it: one composite image with lettered panels and the question text, without method names, caption, or author claim. The author claim identifies (B), the proposed method’s output, as best on the criterion. Twelve of the 21 evaluated models choose (B), including VisionQ-Judge; GPT-5.5, one of the two strongest judges overall, chooses (A), and the other errors are spread over (A), (C), and (D).
Figure 6 : LLM-free DPO dataset construction pipeline. All fields are extracted directly from structured metadata. The structured VisionQ metadata yields 4,524 questions spanning 513 papers and 46 of the 51 taxonomy leaves.
Figure 7 : Symmetric evidence pair construction. The same evidence body e is used in both chosen and rejected responses; only the letter reference differs. This ensures near-zero initial KL divergence, so the objective can only be improved by choosing the correct letter, not through response length or style.
Model
Overall
2-choice
3-choice
4-choice
( n =717)
( n =376)
( n =157)
( n =184)
Base (Gemma-4-E4B-it)
0.499
0.590
0.401
0.397
DPO-tuned
0.524
0.628
0.433
0.391
Δ (pp)
+2.5
+3.7
+3.2
− 0.5
95% CI (pp)
[ − 1.9, +7.0]
Table 2 : Main results on the VisionQ-Judge test set (717 questions).
Setting
Source
A
B
C
D
4-choice
Gold
28.8%
23.4%
24.5%
23.4%
Base
15.2%
18.5%
29.9%
36.4%
Tuned
17.9%
22.8% †
31.0%
28.3% †
3-choice
Gold
30.6%
34.4%
35.0%
-
Base
20.5%
23.3%
56.2%
-
Tuned
19.2%
32.2% †
48.6%
-
Table 3 : Position-bias: predicted letter frequencies vs. gold. Bold values indicate the biased slot; † marks near-gold calibration.
Figure 8 : Predicted letter-frequency distributions for the base and DPO-tuned models vs. gold on the 717-question test set, for 2-, 3-, and 4-choice questions. The base model over-predicts the last option (B, C, and D respectively); after DPO training the predicted frequencies move toward the gold distribution.
Figure 9 : Per-leaf accuracy delta (DPO-tuned vs. base) for taxonomy leaves with n≥6 test questions. Left panel: five leaves with the largest decline. Right panel: five leaves with the largest improvement. Blur, Anatomy, and Lighting show the largest gains and Noise and Semantic Match the largest declines; Lighting, Anatomy, Attribute Binding, and Noise each come from a single test paper.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Axis
Group
Leaf
Definition
Image Appearance
Blur
Focus loss, smear, or blur visible at the panel or crop level.
Noise
Random speckle, grain, sensor-like corruption, or noisy visual artifacts.
Exposure
Over-exposure, under-exposure, saturation clipping, or visibility loss from brightness.
Contrast
Insufficient or excessive tonal separation that harms visible interpretation.
Color
Color cast, color inconsistency, or visually implausible color independent of prompt/reference criteria.
Lighting
Illumination, shadow, reflection, or lighting quality visible in the output.
Appendix
Table 4: Full 51 -leaf codebook, grouped by axis. Group indicates the internal Reference Fidelity split (Source Preservation vs. Target Fidelity); all other axes are ungrouped.
Figure 10 : Headline accuracy of 20 models on VisionQ-Bench ( n=309 ). VisionQ-Judge is the checkpoint for which VisionQ-Bench is the held-out test split; its base model is shown for reference. Unparsed responses count as incorrect.
Figure 11 : Per-axis accuracy with Wilson 95% CI (axes with more than 30 samples; models sorted by overall accuracy)
Model
Overall
Object Form
Ref. Fidelity
Img. Appearance
Relation
Prompt Match
Scene Layout
309 / 109
143 / 50
66 / 23
39 / 16
31 / 8
22 / 9
8 / 4
OpenAI GPT-5.3-codex
63 [57,69]
64 [55,73]
59 [45,71]
69 [53,83]
58 [46,78]
59 [29,88]
75 [40,100]
OpenAI GPT-5.5
63 [57,69]
68 [58,77]
61 [48,72]
64 [48,79]
45 [31,60]
64 [38,85]
62 [29,100]
OpenAI GPT-5.2
61 [55,67]
66 [57,75]
55 [43,66]
62 [46,76]
58 [48,75]
59 [29,84]
50 [20,100]
OpenAI GPT-5.4
61 [55,67]
64 [56,73]
56 [43,68]
67 [49,82]
55 [44,65]
59 [28,86]
38 [0,86]
OpenAI GPT-5.1
59 [53,66]
59 [50,69]
53 [40,66]
69 [52,85]
52 [35,68]
64 [33,89]
75 [54,100]
Appendix
Table 5 : Accuracy (%) per axis on VisionQ-Bench with 95% confidence intervals from a bootstrap that resamples source papers. Column headers give questions / source papers. Relation, Prompt Match, and Scene Layout rest on few papers, so their intervals are wide and category-level conclusions for them are indicative only.
Figure 12 : Source papers by venue and year, over the full 1,409 -paper corpus underlying VisionQ ( 1,399 with known venue). 2024 is CVPR-only because ICCV is held only in odd years.
Item
Count
Notes
Papers with ≥1 annotated figure
1,403
Of the 1,409 source papers (Section 3.1 )
Task types
20
Task-type labels assigned to each paper
Annotated figures
1,486
Qualitative-comparison figures with ≥ 1 data point
Total data points
3,911
Sample-level records (figure × group)
In-scope data points
3,773
After excluding 138 out-of-scope records (60 figures)
Papers with ≥1 data point
790
56.3% of the 1,403 papers with an annotated figure
Appendix
Table 6 : VisionQ dataset composition.
Figure 13 : Number of papers per task type in the VisionQ corpus ( N=1,403 papers with an annotated figure). The five largest fields together account for over half of them.
Figure 14 : Coverage depth per task type. Left: total data-point count. Right: mean data points per covered paper (papers with at least one in-scope data point). Novel view synthesis has the highest evaluation density, followed by face-centric and 3D-generation papers.
Figure 15 : Task-type × leaf matrix. Each cell states a leaf’s share of the task type’s total claim mass (row-normalised over all 51 leaves); only the top-20 leaves by overall count are displayed, so rows sum to less than 100%. Deep red cells indicate a field’s concentrated reliance on a narrow leaf; uniformly pale rows (e.g., segmentation_detection) indicate broad evaluation vocabulary. Reconstruction 3D papers concentrate 57.8% of claim mass in three Object Form leaves (Detail, Completeness, Surface); generative 2D papers place 30.2% on Semantic Match alone. This matrix is the source of the per-field top- K leaf rankings used in the selective comparison analysis (Section 3.5 ).
Figure 16 : Primary-axis distribution per task type (row-normalised %). generative_2d and scene_understanding are Prompt Match–dominated; reconstruction_3d and gaussian_splatting are Object Form–dominated; novel_view_synthesis shows the most balanced distribution across the six axes.
Figure 17 : Shannon entropy of leaf-count distributions per task type (bits). Lower entropy indicates that a field concentrates evaluations on a narrow leaf set. Bars are annotated with the number of active leaves.
Figure 18 : Top-5 taxonomy leaves per task type by data-point count (one panel per task type). The distinct leaf profiles across panels confirm that each field applies a recognisably different evaluation vocabulary, motivating the field-signature approach to selective comparison detection.