Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.
Figures & tables
Figure 1 : The four conditions of MIST , 50 phrases each. T2 and T3 are aligned , T1 and T4 misleading . Bold marks the target.
labels that change (%)
Judge
attaching an aligned image
attaching a misleading image
prompt edited, image fixed
of those that differ, % toward the image
GPT-5.2
10.0
9.1
7.7
28
G3.5-F
10.9
11.8
7.9
34
G3.1-FL
14.5
13.0
8.4
12
Gm3-27B
17.9
19.2
10.6
30
MiS-24B
21.2
16.9
9.1
21
Table 1 : Change rates, pooled over the four prompts, for the seven judges that pass the alt-test in at least one configuration.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2 : Constructing a T1 item: the sentence comes from the compound’s figurative row and the image from the literal row’s correct_image . T2 and T3 take both from the same row; T4 mirrors T1 in the opposite direction.
FF
WF
FL
LL
Total
T1
62
76
10
2
150
T2
49
85
10
6
150
T3
5
14
64
67
150
T4
5
18
46
81
150
Aligned
54
99
74
73
300
Misleading
67
94
56
83
300
Appendix
Table 2 : Label usage: the 600 independent annotations, by condition. Each item carries three labels from one trio.
Trio
n
% unanimous
Fleiss κ
Mean Cohen κ
Informed
100
63.0
0.65
0.66
Blind
100
56.0
0.57
0.57
Pooled (not interpretable)
200
59.5
0.62
0.62
Appendix
Table 3 : Inter-annotator agreement on independent labels.
Informed
Blind
n
Fleiss κ
n
Fleiss κ
T1 (fig. sent., lit. img.)
26
0.50
24
0.30
T2 (fig. sent., fig. img.)
26
0.52
24
0.44
T3 (lit. sent., lit. img.)
24
0.63
26
0.53
T4 (lit. sent., fig. img.)
24
0.62
26
0.41
Aligned (T2 + T3)
50
0.65
50
0.62
Appendix
Table 4 : Fleiss κ per condition and per alignment, computed separately within each trio and never pooled.
Fleiss κ
Per-item agreement
Trio
aln / mis
Δκ [95% CI]
aln / mis
Δ [95% CI]
p
Informed
0.65 / 0.65
−0.00[−0.18,+0.18]
0.740 / 0.747
−0.007[−0.14,+0.13]
0.92
Blind
0.62 / 0.52
+0.10[−0.10,+0.31]
0.727 / 0.647
+0.080[−0.06,+0.22]
0.27
Appendix
Table 5 : Aligned-versus-misleading contrasts, per trio. Left: Fleiss κ with a 10,000-resample bootstrap interval. Right: mean per-item agreement with a two-sided Welch t -test, n=50 per group.
labels that change (%)
Judge
attaching an aligned image
attaching a misleading image
prompt edited, image fixed
of those that differ, % toward the image
GPT-5.2
10.0
9.1
7.7
28
GPT-4o ∘
14.5
12.2
8.8
28
G3.5-F
10.9
11.8
7.9
34
G3.1-FL
14.5
13.0
8.4
12
Gm3-27B
17.9
19.2
10.6
30
Appendix
Table 6 : Change rates for all thirteen judges, the full version of Table 1 . Columns are as in that table. Each judge contributes the same number of cells to columns 1–3, so their means are also pooled rates; column 4’s denominator varies by judge, and pooling its cells gives 37% rather than the 34% shown. ∘ marks the six judges that pass the alt-test in no configuration and are omitted from the body table; every judge, in both groups, exceeds its own column 3 rate in both image columns.
Figurative sentences
Literal sentences
Judge
Prompt
text-only → aligned
aligned → misleading
text-only → misleading
text-only → aligned
aligned → misleading
text-only → misleading
GPT-5.2
Zero-Shot
−0.04
+0.01
−0.03
−0.07
+0.04
−0.03
Few-Shot
+0.04
−0.01
+0.03
−0.18
+0.07
−0.11
CoT
+0.00
−0.07
−0.07
−0.10
+0.05
−0.05
Few-Shot + CoT
−0.02
−0.04
−0.06
−0.09
+0.05
−0.04
GPT-4o
Zero-Shot
+0.04
−0.17
−0.13
−0.18
+0.19
+0.01
Appendix
Table 7 : Drift decomposition for all thirteen judges under each prompt. Mean ordinal shift (FF=0..LL=3); blue marks the predicted sign, positive on figurative and negative on literal sentences. Bold marks the selected cell, per judge and sentence type. Judges grouped by family.
Judge
S
F
F/S
Q2.5-7B
65
39
0.60
K-VL
86
43
0.50
GPT-4o
19
9
0.47
Q3.6-27B
32
15
0.47
IV-30B
52
24
0.46
Q2-7B
76
35
0.46
Appendix
Table 8 : Direction of label changes under the aligned image, Few-Shot + CoT. S is the number of items whose label changes from text-only; F the subset moving toward that image’s sense.
Informed trio
Blind trio
aligned
misleading
all
aligned
misleading
all
ACC
MAE
ACC
MAE
ACC
MAE
ACC
MAE
ACC
MAE
ACC
MAE
Judge
Prompt
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
GPT-5.2
Zero-Shot
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
Few-Shot
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.33
.00
.33
CoT
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.33
.00
.33
Appendix
Table 9 : Alt-test winning rate per judge and prompt on both trio, at ε∈{0.15,0.20} under two scoring functions: ACC , exact match, and MAE , negative mean absolute error on the ordinal codes FF=0 to LL=3. Green marks a pass, WR ≥0.5 . Conditions are aligned (T2+T3), misleading (T1+T4) and all (T1–T4); single T-types are omitted, as n≈25 falls below the recommended minimum per annotator.
Vision-language models (VLMs) increasingly act as judges that pick the best of several generated images, so their choices decide what users see. Such judges are usually validated by score agreement with human ratings, not by the images they return. We audit VLM judges as decision-makers: on 300 culturally situated prompts, we compare the returned image with human ratings the judge never sees and with random choice from the same candidates, and repeat every decision with the candidates reordered. A 4B-parameter judge barely beats random and falls short of a CLIP similarity baseline. It picks the first image shown in 49% of calls (chance: 28%), and reordering changes its choice on 60% of prompts. For this judge, agreement across orders is informative: decisions that survive reordering are much better than random, whereas agreement with a weaker second judge keeps the wrong ones. An 8B judge shows almost no position bias and outperforms CLIP, yet for it the same filter mostly discards good decisions. Agreement helps only when it targets the judge's failure mode, so filters must be re-audited whenever the judge changes. The 4B judge's slight rise in stereotype ratings is no longer detectable after aggregating across orders or with the larger judge.
The reliability of VLM-as-a-Judge is critical for the automatic evaluation of vision-language models (VLMs). Despite recent progress, our analysis reveals that VLM-as-a-Judge often pays limited attention to the image when making decisions. Instead, they often blindly favor the more informative answer, even when they can recognize it conflicts with the image content. We call this problem informativeness bias, which significantly undermines judge reliability. To address it, we propose BIRCH (Balanced Informativeness and CoRrectness with a Truthful AnCHor), a judging paradigm that first corrects inconsistencies with the image content in candidate answers, and then compares the answers against this corrected version. This shifts the judge's focus from informativeness to image-grounded correctness. Experiments on multiple models and benchmarks show that BIRCH reduces informativeness bias by up to 17%, resulting in performance gains of up to 9.8%. Our work reveals an overlooked but fundamental flaw in current VLM-as-a-Judge systems and highlights the need for more principled designs.
Visual inputs are often assumed to improve language understanding in multimodal models. We examine this assumption by asking whether vision-language models (VLMs) can distinguish useful visual evidence from incidental image context in lexical judgments. We use human concreteness and imagery ratings because they span words with varying expected visual relevance, from abstract and low-imagery words to concrete and high-imagery words. We find that real-image contexts do not yield consistent gains and often hurt alignment with human ratings, most sharply when visual evidence is least relevant. Through probing and canonical correlation analysis, complemented by an attribution case study, we find that real-image contexts are associated with representational shifts and greater sensitivity to spurious visual cues, coinciding with weaker recoverability of the targeted lexical properties. We further show that instructing models to focus solely on textual content at inference time can reduce this degradation, with the clearest gains on these vulnerable subsets. Our findings suggest that current instruction-tuned VLMs need better calibration of when visual context should inform lexical judgments.