cs.CVMar 23, 2026

The Dual Mechanisms of Spatial Variable Binding in Vision-Language Models

Authors: Kelly Cui, Nikhil Prakash, Shoval Messica, Ayush Raina, David Bau, Antonio Torralba, Tamar Rott Shaham

Organizations: MIT CSAIL · Northeastern University · Sony Playstation

Abstract

Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechanisms to represent spatial variable binding. In the language model backbone, intermediate layers represent content-independent spatial relations on top of visual tokens corresponding to objects. However, this mechanism plays only a secondary role in shaping model predictions. Instead, the dominant source of spatial information originates in the vision encoder, whose representations encode the layout of objects and are directly exploited by the language model backbone. Notably, this spatial signal is distributed globally across visual tokens, extending beyond object regions into surrounding background areas. We validate the generalization of our findings to complex natural images from the COCO dataset, where globally amplifying the vision-derived spatial representations across all image tokens corrects spatial variable binding failures across models of various sizes. Together, our results clarify how spatial variable binding is computed within VLMs and highlight the central role of vision encoders in enabling it.

Figures & tables

Appendix figures & tables86 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Where To Look? : Causal Tracing of Vision Encoders in VLM

    Aug 11, 2026Naren Kumar S, Tirth Bhatt, Mayank SinghRecent Vision-Language ModelsVision Encoders

  2. Why Far Looks Up: Probing Spatial Representation in Vision-Language Models

    May 28, 2026Cheolhong Min, Jaeyun Jung, Daeun Lee +5Entanglement

  3. SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    Aug 3, 2026Jing Wu, Jianhua Wu, Jiayi Guan +5Recent Vision-Language ModelsStable Spatial Understanding