cs.CVSep 29, 2026

Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry

Authors: Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim

Organizations: Seoul National University

Abstract

Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across factors and models. Direction-specific and visual-absence controls tie these gains to the relevant visual contrast; the gains persist after global alignment and extend to compositional retrieval and natural images. Together, these results show that visual discriminability and linguistic addressability need not coincide, and that matched visual grounding can probe and reduce the resulting access gap.

Figures & tables

Appendix figures & tables20 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Real Images, Worse Judgments: Evaluating Vision-Language Models on Concreteness and Imagery

    May 26, 2026Yifan Jiang, Ruoxi Ning, Sheng Yao +1Representational CollapseRelevance

  2. Mirage Probes: How Vision Models Fake Visual Understanding

    Jun 11, 2026Daniel Ben-Levi, Judah Goldfeder, Weiliang Zhao +5Weak Visual GroundingVision Foundation Models