cs.CVOct 5, 2026

Readout Blindness: VLM Scores Miss the Spatial Direction Their Frozen Encoders Retain

Authors: Guangyuan Li, Tianming Du, Yan Jiang, Bihan Wen, Jiancheng Yang

Organizations: ELLIS Institute Finland · Aalto University · University of Oulu · Nanyang Technological University

Abstract

CLIP-like vision-language models remain a cornerstone of multimodal systems, yet their scores stay near chance on directed spatial relations, such as whether one object is left of another. We call this failure readout blindness and analyze, theoretically and empirically, why deployed scores miss the direction: when scoring rules treat the subject and object symmetrically, direction cancels regardless of encoder training. Guided by this analysis, we introduce Antisymmetric Displacement Readout (ADR), which aligns caption words with image patches in the frozen features and scores each relation by the signed displacement between matched object centroids. Notably, ADR succeeds without additional training or learned parameters, thereby demonstrating that directional information remains in the frozen encoder. However, text and world priors can inflate accuracy, so we further introduce prior deflation, which measures the benefit of the image-text pairing as the grounded gain over a null that pairs each item with an unrelated image. Extensive experiments across encoder families show that ADR substantially improves over deployed scores, which remain near chance on most direction-balanced sets even for fine-tuned encoders. Compared with more complex readouts, ADR outperforms the evaluated MLLM likelihood readouts and is competitive with their chat inference at a small fraction of the computation. These results support our claim that directional information can be recovered from frozen features by an appropriate readout. Our implementation and evaluation kit will be publicly available.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Which Way Did It Move? Diagnosing and Overcoming Directional Motion Blindness in Video-LLMs

    May 21, 2026Jongseo Lee, Hyuntak Lee, Sunghun Kim +3Video-Language ModelsTemporal Video Understanding

  2. BindCLIP: One Balanced Coupling For Compositional Vision Language Scoring

    Sep 20, 2026Liuyang Song, Yi Zhang, Zhongyi Deng +2Vision-Language ModelsVision-Language Grounding

  3. Looks the Same, Answers Differently: Flip-Direction Steering for Robust Vision-Language Reasoning

    Sep 23, 2026Yeonsung Jung, Joonhyun Jeong, Hoang Pham +4VLM RobustnessCounterfactual Evaluation of VLMs