When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context
Organizations: School of Artificial Intelligence, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China · International Center of Future Science, Jilin University
Abstract
Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45% to 58.51%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.
Figures & tables
| Model alias | Clear | Blur | ||||
|---|---|---|---|---|---|---|
| L (Acc.) | C (RR) | O | L (Acc.) | C (RR) | O | |
| Closed-Source MLLMs | ||||||
| Gemini 3.1 Flash | 90.14 | 8.45 | 1.41 | 78.36 | 20.74 | 0.90 |
| Gemini 3.5 Flash | 89.76 | 8.96 | 1.28 | 80.41 | 18.44 | 1.15 |
| Gemini 3 Flash | 88.35 | 11.01 | 0.64 | 79.39 | 19.97 | 0.64 |
| Claude Sonnet 4.6 | 85.02 | 12.16 | 2.82 | 71.06 | 22.66 | 6.27 |
| Model alias | Full image | Gray mask | C | 95% CI | ||||
|---|---|---|---|---|---|---|---|---|
| L | C | O | L | C | O | |||
| GPT-5.2 | 67.35 | 29.58 | 3.07 | 84.25 | 11.01 | 4.74 | -18.57 | [-21.64, -15.49] |
| Gemini 3.5 Flash | 89.76 | 8.96 | 1.28 | 96.93 | 2.43 | 0.64 | -6.53 | [-8.45, -4.74] |
| Claude Sonnet 4.6 | 85.02 | 12.16 | 2.82 | 96.80 | 2.56 | 0.64 | -9.60 | [-11.91, -7.43] |
| Qwen3-VL-32B T. | 47.50 | 44.43 | 8.07 | 67.86 | 24.33 | 7.81 | -20.10 | [-23.30, -17.03] |
| Model alias | Full image | Crop only | ||||
|---|---|---|---|---|---|---|
| L (Acc.) | C (RR) | O | L (Acc.) | C (RR) | O | |
| GPT-5.2 | 67.35 | 29.58 | 3.07 | 89.88 | 5.63 | 4.48 |
| Gemini 3.5 Flash | 89.76 | 8.96 | 1.28 | 96.93 | 2.18 | 0.90 |
| Claude Sonnet 4.6 | 85.02 | 12.16 | 2.82 | 93.60 | 3.59 | 2.82 |
| Qwen3-VL-32B Thinking | 47.50 | 44.43 | 8.07 | 86.94 | 7.55 | 5.51 |
| Method | Overall EM | Clear EM | Blurred EM | Pair accuracy | CER |
|---|---|---|---|---|---|
| Base Model | |||||
| SFT | |||||
| SFT+GRPO |
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
| Category | Images | Category | Images | Category | Images |
|---|---|---|---|---|---|
| Animal | 109 | Electronics | 14 | Profession | 79 |
| Building | 20 | Home | 27 | Recipe | 91 |
| City | 64 | Landform | 5 | Sport | 44 |
| Clothing | 26 | Medicine | 37 | Tool | 14 |
| Command | 32 | Music | 42 | Vehicle | 34 |
| Country | 56 | Plant | 87 | Total | 781 |
| Analysis | Words | Adjusted OR | 95% CI |
|---|---|---|---|
| Primary: period-terminated gap | 781 | 1.352 | [1.197, 1.527] |
| No-period gap | 781 | 1.377 | [1.214, 1.562] |
| No-period gap, per character | 781 | 1.341 | [1.195, 1.506] |
| Period gap, per character | 781 | 1.312 | [1.169, 1.473] |
| Equal-length candidates | 398 | 1.387 | [1.190, 1.617] |
| Alphabetic printed strings | 704 | 1.340 | [1.183, 1.519] |
| Analysis | Words | Strength : OR [95% CI] | Gap : OR [95% CI] |
|---|---|---|---|
| All: separate predictors | 781 | 1.202 [1.081, 1.337] | 1.352 [1.197, 1.527] |
| All: joint predictors | 781 | 1.017 [0.895, 1.157] | 1.339 [1.155, 1.551] |
| Alphabetic only | 704 | 1.225 [1.099, 1.364] | 1.340 [1.183, 1.519] |
| Equal length | 398 | 1.179 [1.008, 1.379] | 1.387 [1.190, 1.617] |
| Characters in | Images | L (%) | C / RR (%) | O (%) | RR 95% CI |
|---|---|---|---|---|---|
| 4–5 | 199 | 79.80 | 14.27 | 5.93 | [11.99, 16.68] |
| 6–7 | 302 | 69.16 | 24.48 | 6.36 | [21.88, 27.17] |
| 8–9 | 199 | 60.13 | 33.67 | 6.20 | [30.05, 37.39] |
| 81 | 49.79 | 45.02 | 5.19 | [39.26, 50.95] |
| Model alias | Rewriting rate (%) | ||
|---|---|---|---|
| Overall | 95% Wilson CI | Category macro | |
| Gemini 3.1 Flash | 8.45 | [6.70, 10.61] | 8.33 |
| Gemini 3.5 Flash | 8.96 | [7.16, 11.17] | 7.43 |
| Gemini 3 Flash | 11.01 | [9.00, 13.40] | 9.34 |
| Claude Sonnet 4.6 | 12.16 | [10.05, 14.64] | 11.27 |
| GPT-5.5 | 19.46 | [16.84, 22.39] | 18.33 |
| Weighting | Clear (%) | Moderate (%) | ||||
|---|---|---|---|---|---|---|
| L | C | O | L | C | O | |
| Equal alias weighting (15 aliases) | 67.55 | 26.36 | 6.09 | 54.94 | 36.95 | 8.11 |
| Equal family weighting (7 families) | 68.03 | 27.16 | 4.81 | 54.65 | 38.31 | 7.05 |
| Excluding Qwen: equal alias weighting (9 aliases) | 74.52 | 22.36 | 3.12 | 60.95 | 34.40 | 4.64 |
| Excluding Qwen: equal family weighting (6 families) | 69.86 | 26.29 | 3.85 | 56.10 | 37.90 | 6.00 |
| Size | Variant | Literal | Canonical | Other |
|---|---|---|---|---|
| 8B | Instruct | 486 (62.23) | 152 (19.46) | 143 (18.31) |
| 8B | Thinking | 379 (48.53) | 296 (37.90) | 106 (13.57) |
| 32B | Instruct | 476 (60.95) | 215 (27.53) | 90 (11.52) |
| 32B | Thinking | 371 (47.50) | 347 (44.43) | 63 (8.07) |
| 235B-A22B | Instruct | 478 (61.20) | 253 (32.39) | 50 (6.40) |
| 235B-A22B | Thinking | 487 (62.36) | 252 (32.27) | 42 (5.38) |
| Size | Instruct outcome | Thinking L | Thinking C | Thinking O |
|---|---|---|---|---|
| 8B | L | 304 | 136 | 46 |
| 8B | C | 16 | 128 | 8 |
| 8B | O | 59 | 32 | 52 |
| 32B | L | 317 | 134 | 25 |
| 32B | C | 23 | 183 | 9 |
| 32B | O | 31 | 30 | 29 |
| Perturbation | Acc. | RR | Other | RR 95% CI | |
|---|---|---|---|---|---|
| Letter-shape substitution | 151 | 53.02 | 41.81 | 5.17 | [37.31, 46.23] |
| Character split / merge | 26 | 50.26 | 38.72 | 11.03 | [28.72, 48.72] |
| Character duplication | 149 | 61.21 | 33.47 | 5.32 | [29.71, 37.23] |
| Digit / symbol substitution | 61 | 66.67 | 27.98 | 5.36 | [20.98, 35.30] |
| Character deletion | 172 | 77.09 | 18.18 | 4.73 | [15.23, 21.28] |
| Character insertion | 36 | 77.04 | 17.59 | 5.37 | [10.56, 25.56] |
| Model alias | Clear | Mild | Moderate | Strong |
|---|---|---|---|---|
| GPT-5.2 | 29.58 | 30.09 | 45.33 | 60.69 |
| Gemini 3.5 Flash | 8.96 | 10.50 | 18.44 | 54.67 |
| Claude Sonnet 4.6 | 12.16 | 10.37 | 22.66 | 43.02 |
| Qwen3-VL-32B Thinking | 44.43 | 47.25 | 51.47 | 59.72 |
| Model alias | Clear | Moderate | RR | 95% CI |
|---|---|---|---|---|
| Gemini 3.1 Flash Lite | 8.45 | 20.74 | +12.29 | [+9.99, +14.72] |
| Gemini 3.5 Flash | 8.96 | 18.44 | +9.48 | [+7.04, +11.91] |
| Gemini 3 Flash Preview | 11.01 | 19.97 | +8.96 | [+6.53, +11.40] |
| Claude Sonnet 4.6 | 12.16 | 22.66 | +10.50 | [+7.43, +13.57] |
| GPT-5.5 | 19.46 | 40.20 | +20.74 | [+17.41, +24.07] |
| Qwen3-VL-8B I. | 19.46 | 26.76 | +7.30 | [+4.87, +9.86] |
| Model alias | A: L | A: C | B: L | B: C | 95% CI | |
|---|---|---|---|---|---|---|
| GPT-5.2 | 80.67 | 17.33 | 89.33 | 7.33 | ||
| GPT-5.5 | 80.67 | 16.00 | 90.00 | 8.00 | ||
| Claude Sonnet 4.6 | 99.33 | 0.67 | 98.00 | 0.67 | ||
| Gemini 3.1 Flash Lite | 98.67 | 1.33 | 98.67 | 1.33 | ||
| Gemini 3.5 Flash | 96.67 | 2.67 | 96.00 | 2.00 | ||
| Qwen3-VL-8B-I | 88.00 | 8.00 | 92.67 | 4.00 |
| Model alias | A: O | B: O | M: L | M: C | S: C | 95% CI | |
|---|---|---|---|---|---|---|---|
| GPT-5.2 | 2.00 | 3.33 | 90.00 | 9.33 | 12.00 | ||
| GPT-5.5 | 3.33 | 2.00 | 88.67 | 10.00 | 10.00 | ||
| Claude Sonnet 4.6 | 0.00 | 1.33 | 99.33 | 0.67 | 0.00 | ||
| Gemini 3.1 Flash Lite | 0.00 | 0.00 | 98.00 | 1.33 | – | – | – |
| Gemini 3.5 Flash | 0.67 | 2.00 | 95.33 | 1.33 | 2.00 | ||
| Qwen3-VL-8B-I | 4.00 | 3.33 | 66.67 | 25.33 | 7.33 |
| Agreement rule | Accepted | Coverage (%) | L/C/O | Error (%) |
|---|---|---|---|---|
| Three models, exact | 474 | 60.69 | 456/16/2 | 3.80 |
| Four models, exact | 309 | 39.56 | 291/16/2 | 5.83 |
| Three models, normalized | 493 | 63.12 | 465/26/2 | 5.68 |
| Four models, normalized | 327 | 41.87 | 299/26/2 | 8.56 |
| Model alias | Full clear | Full moderate | Gray clear | Gray moderate | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| L | C | O | L | C | O | L | C | O | L | C | O | |
| GPT-5.2 | 67.35 | 29.58 | 3.07 | 50.19 | 45.33 | 4.48 | 84.25 | 11.01 | 4.74 | 68.25 | 20.87 | 10.88 |
| Gemini 3.5 Flash | 89.76 | 8.96 | 1.28 | 80.41 | 18.44 | 1.15 | 96.93 | 2.43 | 0.64 | 87.32 | 10.76 | 1.92 |
| Claude Sonnet 4.6 | 85.02 | 12.16 | 2.82 | 71.06 | 22.66 | 6.27 | 96.80 | 2.56 | 0.64 | 82.71 | 11.40 | 5.89 |
| Qwen3-VL-32B T. | 47.50 | 44.43 | 8.07 | 35.60 | 51.47 | 12.93 | 67.86 | 24.33 | 7.81 | 58.64 | 25.61 | 15.75 |