cs.CVSep 28, 2026

When VLMs Trust Context: Evaluating Scene Text Recognition under Misleading Context

Authors: Yuxing Cheng, Yuan Wu, Yi Chang

Organizations: School of Artificial Intelligence, Jilin University · Engineering Research Center of Knowledge-Driven Human-Machine Intelligence, MOE, China · International Center of Future Science, Jilin University

Abstract

Vision-language models (VLMs) can read text in natural scenes, but their predictions may be influenced by the surrounding context. When the printed text conflicts with what the scene suggests, a model may return a more plausible word instead of the shown text. We introduce SceneFaith, a benchmark of 781 generated scene images for studying this behavior. Each output is classified as Literal, Canonical, or Other, separating faithful transcription from context-consistent rewriting and ordinary recognition errors. Across 15 models from seven families, all models show rewriting on clear images, with rates ranging from 8.45% to 58.51%. Controlled experiments further show that surrounding context matters: removing surrounding scene information reduces rewriting and improves literal accuracy, while changing the scene around the same text patch can also change model outputs. Moreover, weakening the target text with blur increases rewriting. These results show that reliable scene-text recognition requires VLMs to balance visual character evidence with contextual information, preserving clear text while using context mainly when the visual evidence is uncertain.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models

    Aug 4, 2026Sadab Shiper, Tawsif Tashwar Dipto, Mir Md Inzamam +1Optical Character RecognitionScene Text Recognition

  2. Do Vision-Language Models See or Guess? Measuring and Reducing Textual-Prior Reliance with a Phrasing-Controlled Benchmark

    Jun 9, 2026Pratham Singla, Shivank Garg, Vihan Singh +1Textual PriorsQuestion

  3. Mirage Probes: How Vision Models Fake Visual Understanding

    Jun 11, 2026Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao +5Weak Visual GroundingVision Foundation Models