When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
Authors: Yuxing Cheng, Yuan Wu, Yi Chang
Organizations: School of Artificial Intelligence, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China · International Center of Future Science, Jilin University
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45% to 58.51%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
Figures & tables
Figure 1: SceneFaith construction pipeline. Spelling perturbation and scene generation are followed by quality checks, blind crop reading, and dataset completion.
Figure 2: Qualitative example of canonical rewriting. The printed gold is ieopard; leopard is the canonical alternative.
Model alias
Clear
Blur
L (Acc.) ↑
C (RR) ↓
O ↓
L (Acc.) ↑
C (RR) ↓
O ↓
Closed-Source MLLMs
Gemini 3.1 Flash
90.14
8.45
1.41
78.36
20.74
0.90
Gemini 3.5 Flash
89.76
8.96
1.28
80.41
18.44
1.15
Gemini 3 Flash
88.35
11.01
0.64
79.39
19.97
0.64
Claude Sonnet 4.6
85.02
12.16
2.82
71.06
22.66
6.27
Table 1: Transcription outcomes on clear and moderately blurred text. All rates are percentages. The 15 aliases use full images and the neutral prompt and are ordered by clear RR within each group.
Model alias
Full image
Gray mask
Δ C
95% CI
L
C
O
L
C
O
GPT-5.2
67.35
29.58
3.07
84.25
11.01
4.74
-18.57
[-21.64, -15.49]
Gemini 3.5 Flash
89.76
8.96
1.28
96.93
2.43
0.64
-6.53
[-8.45, -4.74]
Claude Sonnet 4.6
85.02
12.16
2.82
96.80
2.56
0.64
-9.60
[-11.91, -7.43]
Qwen3-VL-32B T.
47.50
44.43
8.07
67.86
24.33
7.81
-20.10
[-23.30, -17.03]
Table 2: Scene removal with fixed target pixels. Intervals use 20,000 paired-image bootstrap resamples.
Model alias
Full image
Crop only
L (Acc.) ↑
C (RR) ↓
O ↓
L (Acc.) ↑
C (RR) ↓
O ↓
GPT-5.2
67.35
29.58
3.07
89.88
5.63
4.48
Gemini 3.5 Flash
89.76
8.96
1.28
96.93
2.18
0.90
Claude Sonnet 4.6
85.02
12.16
2.82
93.60
3.59
2.82
Qwen3-VL-32B Thinking
47.50
44.43
8.07
86.94
7.55
5.51
Table 3: Transcription outcomes for full-image and crop-only recognition. All rates are percentages, with 781 images per entry.
Method
Overall EM ↑
Clear EM ↑
Blurred EM ↑
Pair accuracy ↑
CER ↓
Base Model
67.93
66.07
69.78
42.25
7.77
SFT
81.90±0.50
83.35±0.68
80.45±1.60
66.58±0.59
5.80±0.53
SFT+GRPO
82.44±0.33
84.08±0.64
80.79±1.22
67.65±0.77
5.64±0.34
Table 4: Results with 300 training pairs. Values are percentages, reported as mean ± sample standard deviation over three training seeds. Pair accuracy requires both predictions in a pair to be correct.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Images
Category
Images
Category
Images
Animal
109
Electronics
14
Profession
79
Building
20
Home
27
Recipe
91
City
64
Landform
5
Sport
44
Clothing
26
Medicine
37
Tool
14
Command
32
Music
42
Vehicle
34
Country
56
Plant
87
Total
781
Appendix
Table 5: Category inventory of the complete 781-image benchmark. All reported experiments use this aggregate collection.
Analysis
Words
Adjusted OR
95% CI
Primary: period-terminated gap
781
1.352
[1.197, 1.527]
No-period gap
781
1.377
[1.214, 1.562]
No-period gap, per character
781
1.341
[1.195, 1.506]
Period gap, per character
781
1.312
[1.169, 1.473]
Equal-length candidates
398
1.387
[1.190, 1.617]
Alphabetic printed strings
704
1.340
[1.183, 1.519]
Appendix
Table 6: Exploratory lexical-preference analysis and 9 sensitivity checks.
Analysis
Words
Strength A : OR [95% CI]
Gap D : OR [95% CI]
All: separate predictors
781
1.202 [1.081, 1.337]
1.352 [1.197, 1.527]
All: joint predictors
781
1.017 [0.895, 1.157]
1.339 [1.155, 1.551]
Alphabetic only
704
1.225 [1.099, 1.364]
1.340 [1.183, 1.519]
Equal length
398
1.179 [1.008, 1.379]
1.387 [1.190, 1.617]
Appendix
Table 7: DistilGPT2 strength A and preference gap D on the 781-image benchmark.
Characters in ci
Images
L (%)
C / RR (%)
O (%)
RR 95% CI
4–5
199
79.80
14.27
5.93
[11.99, 16.68]
6–7
302
69.16
24.48
6.36
[21.88, 27.17]
8–9
199
60.13
33.67
6.20
[30.05, 37.39]
≥10
81
49.79
45.02
5.19
[39.26, 50.95]
Appendix
Table 8: Longer conventional strings accompany more canonical substitutions.
Model alias
Rewriting rate (%)
Overall ↓
95% Wilson CI
Category macro ↓
Gemini 3.1 Flash
8.45
[6.70, 10.61]
8.33
Gemini 3.5 Flash
8.96
[7.16, 11.17]
7.43
Gemini 3 Flash
11.01
[9.00, 13.40]
9.34
Claude Sonnet 4.6
12.16
[10.05, 14.64]
11.27
GPT-5.5
19.46
[16.84, 22.39]
18.33
Appendix
Table 9: Supplementary statistics for clear-image rewriting.
Weighting
Clear (%)
Moderate (%)
L
C
O
L
C
O
Equal alias weighting (15 aliases)
67.55
26.36
6.09
54.94
36.95
8.11
Equal family weighting (7 families)
68.03
27.16
4.81
54.65
38.31
7.05
Excluding Qwen: equal alias weighting (9 aliases)
74.52
22.36
3.12
60.95
34.40
4.64
Excluding Qwen: equal family weighting (6 families)
69.86
26.29
3.85
56.10
37.90
6.00
Appendix
Table 10: Sensitivity to model-family weighting. Brackets show 95% paired-image bootstrap confidence intervals for changes from clear to moderate blur.
Size
Variant
Literal
Canonical
Other
8B
Instruct
486 (62.23)
152 (19.46)
143 (18.31)
8B
Thinking
379 (48.53)
296 (37.90)
106 (13.57)
32B
Instruct
476 (60.95)
215 (27.53)
90 (11.52)
32B
Thinking
371 (47.50)
347 (44.43)
63 (8.07)
235B-A22B
Instruct
478 (61.20)
253 (32.39)
50 (6.40)
235B-A22B
Thinking
487 (62.36)
252 (32.27)
42 (5.38)
Appendix
Table 11: Qwen clear-image outcome counts (percentages), with N=781 for every row.
Size
Instruct outcome
Thinking L
Thinking C
Thinking O
8B
L
304
136
46
8B
C
16
128
8
8B
O
59
32
52
32B
L
317
134
25
32B
C
23
183
9
32B
O
31
30
29
Appendix
Table 12: Matched-image outcome counts from Instruct (row) to Thinking (column).
Perturbation
n
Acc.
RR
Other
RR 95% CI
Letter-shape substitution
151
53.02
41.81
5.17
[37.31, 46.23]
Character split / merge
26
50.26
38.72
11.03
[28.72, 48.72]
Character duplication
149
61.21
33.47
5.32
[29.71, 37.23]
Digit / symbol substitution
61
66.67
27.98
5.36
[20.98, 35.30]
Character deletion
172
77.09
18.18
4.73
[15.23, 21.28]
Character insertion
36
77.04
17.59
5.37
[10.56, 25.56]
Appendix
Table 13: Clear-image outcomes by recorded character perturbation.
Model alias
Clear
Mild
Moderate
Strong
GPT-5.2
29.58
30.09
45.33
60.69
Gemini 3.5 Flash
8.96
10.50
18.44
54.67
Claude Sonnet 4.6
12.16
10.37
22.66
43.02
Qwen3-VL-32B Thinking
44.43
47.25
51.47
59.72
Appendix
Table 14: Rewriting rates (%) across blur severity.
Model alias
Clear
Moderate
Δ RR
95% CI
Gemini 3.1 Flash Lite
8.45
20.74
+12.29
[+9.99, +14.72]
Gemini 3.5 Flash
8.96
18.44
+9.48
[+7.04, +11.91]
Gemini 3 Flash Preview
11.01
19.97
+8.96
[+6.53, +11.40]
Claude Sonnet 4.6
12.16
22.66
+10.50
[+7.43, +13.57]
GPT-5.5
19.46
40.20
+20.74
[+17.41, +24.07]
Qwen3-VL-8B I.
19.46
26.76
+7.30
[+4.87, +9.86]
Appendix
Table 15: Matched moderate-blur changes on the SceneFaith benchmark.
Model alias
A: L
A: C
B: L
B: C
Δ
95% CI
GPT-5.2
80.67
17.33
89.33
7.33
+10.00
[4.00,16.00]
GPT-5.5
80.67
16.00
90.00
8.00
+8.00
[2.00,14.67]
Claude Sonnet 4.6
99.33
0.67
98.00
0.67
+0.00
[−2.00,2.00]
Gemini 3.1 Flash Lite
98.67
1.33
98.67
1.33
+0.00
[−2.00,2.00]
Gemini 3.5 Flash
96.67
2.67
96.00
2.00
+0.67
[−1.33,3.33]
Qwen3-VL-8B-I
88.00
8.00
92.67
4.00
+4.00
[0.67,8.00]
Appendix
Table 16: Fixed-target-patch comparison on 150 matched pairs per alias.
Model alias
A: O
B: O
M: L
M: C
S: C
ΔA−S
95% CI
GPT-5.2
2.00
3.33
90.00
9.33
12.00
+5.33
[−1.33,12.00]
GPT-5.5
3.33
2.00
88.67
10.00
10.00
+6.00
[−0.67,12.67]
Claude Sonnet 4.6
0.00
1.33
99.33
0.67
0.00
+0.67
[0.00,2.00]
Gemini 3.1 Flash Lite
0.00
0.00
98.00
1.33
–
–
–
Gemini 3.5 Flash
0.67
2.00
95.33
1.33
2.00
+0.67
[0.00,2.00]
Qwen3-VL-8B-I
4.00
3.33
66.67
25.33
7.33
+0.67
[−3.35,4.67]
Appendix
Table 17: Additional outcomes on the same 150 targets.
Agreement rule
Accepted n
Coverage (%)
L/C/O
Error (%)
Three models, exact
474
60.69
456/16/2
3.80
Four models, exact
309
39.56
291/16/2
5.83
Three models, normalized
493
63.12
465/26/2
5.68
Four models, normalized
327
41.87
299/26/2
8.56
Appendix
Table 18: Consensus filtering on the archived 781-image evaluation.
Model alias
Full clear
Full moderate
Gray clear
Gray moderate
L
C
O
L
C
O
L
C
O
L
C
O
GPT-5.2
67.35
29.58
3.07
50.19
45.33
4.48
84.25
11.01
4.74
68.25
20.87
10.88
Gemini 3.5 Flash
89.76
8.96
1.28
80.41
18.44
1.15
96.93
2.43
0.64
87.32
10.76
1.92
Claude Sonnet 4.6
85.02
12.16
2.82
71.06
22.66
6.27
96.80
2.56
0.64
82.71
11.40
5.89
Qwen3-VL-32B T.
47.50
44.43
8.07
35.60
51.47
12.93
67.86
24.33
7.81
58.64
25.61
15.75
Appendix
Table 19: Transcription outcomes across blur and scene-removal conditions. All rates are percentages. T. denotes Thinking.
In-the-wild Bengali scene text recognition is largely unmeasured: existing resources target handwritten documents or constrained sign-board parsing, report only aggregate edit-distance metrics, and evaluate either conventional OCR or VLMs, never both on the same in-the-wild data. To address this gap, we introduce BANGLAWILD, a benchmark of 2,535 Bengali scene text images, each paired with a verbatim gold transcription, two categorical axes, four diagnostic attributes, and an orthographically standard form where the in-image text deviates from canonical spelling. We evaluate fifteen VLMs and three conventional OCR systems under three prompting strategies, fine-tune 6 open-source models with LoRA, and complement edit-distance metrics with an LLM-as-a-Judge evaluation. Our results reveal a persistent gap in which larger models within the same family do not outperform smaller ones. Our fifteen-class error taxonomy shows that visual mis-recognition accounts for ~60% of errors in the strongest systems, while conjunct-related errors contribute under 2%, challenging a long-standing assumption in Bengali OCR research; the same visual dominant profile also holds across architectures, including the one conventional baseline that reads Bengali reliably. Prompt language mainly affects cross-script drift and LoRA reduces catastrophic failures in weak models without lifting the ceiling on already competent ones. Code and data will be publicly released.
Sadab Shiper, Tawsif Tashwar Dipto, Mir Md Inzamam +1
Islamic University of Technology, Bangladesh · BRAC University, Bangladesh
Vision-language models (VLMs) are increasingly deployed where answers must follow from what is in the image, yet they often answer from textual priors, the question's phrasing together with memorized world knowledge, rather than from the image itself, which inflates benchmark scores and yields confident but ungrounded answers. Existing benchmarks rarely isolate this behavior, since each image is usually paired with a single fixed question. To measure the reliance, we build a 540-image benchmark across six reasoning categories and generate four question variants over the same images, so that phrasing rather than image content is the controlled variable. The hardest variant is written directly from the image to minimize text leakage. We benchmark eleven VLMs spanning small open-weight models to large closed-source systems: every model degrades on the hardest variant, and open models fall furthest. Our central diagnostic is a no-image ablation, which collapses the open-weight models to their text-only floor (1 to 9 percent). Three further analyses, LLM-rated difficulty, low base-to-final textual similarity, and human re-annotation, corroborate genuine image-dependence. In-context exemplars that match how a variant was built recover the most accuracy, and GRPO post-training of a small VLM yields consistent gains across all four variants that transfer to a held-out out-of-distribution set. Textual-prior reliance is measurable and partly trainable away.
Pratham Singla, Shivank Garg, Vihan Singh +1
Lossfunk · Indian Institute of Technology Roorkee · Raeth AI
Vision-language models (VLMs) can answer image-based questions confidently, and often correctly, even when no image is provided. This mirage behavior inflates benchmark scores without reflecting visual grounding. Prior work treats this as a single failure mode. We argue it is two. Using Mirage Probes, a contrastive probing framework that pairs paraphrased question variants with matched mirage and non-mirage labels on the same image, we show that mirage behavior is linearly decodable from internal activations across residual stream, MLP, post-attention, and attention-head sites in two open-source VLMs. We demonstrate that a Naive Bayes text baseline cannot recover this signal, ruling out surface lexical confounds. Cross-benchmark separability patterns, together with a novel Prior Harnessing Index (PHI) measuring how much a model can answer from text alone, expose two distinct regimes: textual biases, where the model answers from language priors without engaging visual representations, and spurious images, where it constructs false visual content in latent space and answers as if grounded. The distinction has direct mitigation consequences: text-distribution cleaning can address the first regime but cannot reach the second, since spurious-image mirages live in the model's visual representations rather than its text. Faithful visual grounding will require interventions at the representational level.
Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao +5