Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be decided from the sentence alone, so no image should change any answer. We expected each image to pull a judge's labels toward the sense it depicts, and neither kind did. Across thirteen VLM judges, an aligned image changed 20.5% of labels and a misleading one 19.4%, close for every judge and both above the 11.6% produced by deleting the ignore-the-image instruction with the image left in place. Yet only 37% of the labels that differ between the two images moved toward the sense shown, and agreement with our human annotators is unchanged whether the image is absent, aligned or misleading. The effect is smaller in the seven judges that pass the alt-test than in the six that never do, but present in all of them: what moves a judge is that an image is there, not which of the two it is, so a substitutability verdict describes a configuration as much as a model.
Figures & tables
Figure 1 : The four conditions of MIST , 50 phrases each. T2 and T3 are aligned , T1 and T4 misleading . Bold marks the target.
labels that change (%)
Judge
attaching an aligned image
attaching a misleading image
prompt edited, image fixed
of those that differ, % toward the image
GPT-5.2
10.0
9.1
7.7
28
G3.5-F
10.9
11.8
7.9
34
G3.1-FL
14.5
13.0
8.4
12
Gm3-27B
17.9
19.2
10.6
30
MiS-24B
21.2
16.9
9.1
21
Table 1 : Change rates, pooled over the four prompts, for the seven judges that pass the alt-test in at least one configuration.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 2 : Constructing a T1 item: the sentence comes from the compound’s figurative row and the image from the literal row’s correct_image . T2 and T3 take both from the same row; T4 mirrors T1 in the opposite direction.
FF
WF
FL
LL
Total
T1
62
76
10
2
150
T2
49
85
10
6
150
T3
5
14
64
67
150
T4
5
18
46
81
150
Aligned
54
99
74
73
300
Misleading
67
94
56
83
300
Appendix
Table 2 : Label usage: the 600 independent annotations, by condition. Each item carries three labels from one trio.
Trio
n
% unanimous
Fleiss κ
Mean Cohen κ
Informed
100
63.0
0.65
0.66
Blind
100
56.0
0.57
0.57
Pooled (not interpretable)
200
59.5
0.62
0.62
Appendix
Table 3 : Inter-annotator agreement on independent labels.
Informed
Blind
n
Fleiss κ
n
Fleiss κ
T1 (fig. sent., lit. img.)
26
0.50
24
0.30
T2 (fig. sent., fig. img.)
26
0.52
24
0.44
T3 (lit. sent., lit. img.)
24
0.63
26
0.53
T4 (lit. sent., fig. img.)
24
0.62
26
0.41
Aligned (T2 + T3)
50
0.65
50
0.62
Appendix
Table 4 : Fleiss κ per condition and per alignment, computed separately within each trio and never pooled.
Fleiss κ
Per-item agreement
Trio
aln / mis
Δκ [95% CI]
aln / mis
Δ [95% CI]
p
Informed
0.65 / 0.65
−0.00[−0.18,+0.18]
0.740 / 0.747
−0.007[−0.14,+0.13]
0.92
Blind
0.62 / 0.52
+0.10[−0.10,+0.31]
0.727 / 0.647
+0.080[−0.06,+0.22]
0.27
Appendix
Table 5 : Aligned-versus-misleading contrasts, per trio. Left: Fleiss κ with a 10,000-resample bootstrap interval. Right: mean per-item agreement with a two-sided Welch t -test, n=50 per group.
labels that change (%)
Judge
attaching an aligned image
attaching a misleading image
prompt edited, image fixed
of those that differ, % toward the image
GPT-5.2
10.0
9.1
7.7
28
GPT-4o ∘
14.5
12.2
8.8
28
G3.5-F
10.9
11.8
7.9
34
G3.1-FL
14.5
13.0
8.4
12
Gm3-27B
17.9
19.2
10.6
30
Appendix
Table 6 : Change rates for all thirteen judges, the full version of Table 1 . Columns are as in that table. Each judge contributes the same number of cells to columns 1–3, so their means are also pooled rates; column 4’s denominator varies by judge, and pooling its cells gives 37% rather than the 34% shown. ∘ marks the six judges that pass the alt-test in no configuration and are omitted from the body table; every judge, in both groups, exceeds its own column 3 rate in both image columns.
Figurative sentences
Literal sentences
Judge
Prompt
text-only → aligned
aligned → misleading
text-only → misleading
text-only → aligned
aligned → misleading
text-only → misleading
GPT-5.2
Zero-Shot
−0.04
+0.01
−0.03
−0.07
+0.04
−0.03
Few-Shot
+0.04
−0.01
+0.03
−0.18
+0.07
−0.11
CoT
+0.00
−0.07
−0.07
−0.10
+0.05
−0.05
Few-Shot + CoT
−0.02
−0.04
−0.06
−0.09
+0.05
−0.04
GPT-4o
Zero-Shot
+0.04
−0.17
−0.13
−0.18
+0.19
+0.01
Appendix
Table 7 : Drift decomposition for all thirteen judges under each prompt. Mean ordinal shift (FF=0..LL=3); blue marks the predicted sign, positive on figurative and negative on literal sentences. Bold marks the selected cell, per judge and sentence type. Judges grouped by family.
Judge
S
F
F/S
Q2.5-7B
65
39
0.60
K-VL
86
43
0.50
GPT-4o
19
9
0.47
Q3.6-27B
32
15
0.47
IV-30B
52
24
0.46
Q2-7B
76
35
0.46
Appendix
Table 8 : Direction of label changes under the aligned image, Few-Shot + CoT. S is the number of items whose label changes from text-only; F the subset moving toward that image’s sense.
Informed trio
Blind trio
aligned
misleading
all
aligned
misleading
all
ACC
MAE
ACC
MAE
ACC
MAE
ACC
MAE
ACC
MAE
ACC
MAE
Judge
Prompt
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
.15
.20
GPT-5.2
Zero-Shot
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
Few-Shot
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.33
.00
.33
CoT
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.00
.33
.00
.33
Appendix
Table 9 : Alt-test winning rate per judge and prompt on both trio, at ε∈{0.15,0.20} under two scoring functions: ACC , exact match, and MAE , negative mean absolute error on the ordinal codes FF=0 to LL=3. Green marks a pass, WR ≥0.5 . Conditions are aligned (T2+T3), misleading (T1+T4) and all (T1–T4); single T-types are omitted, as n≈25 falls below the recommended minimum per annotator.