cs.CVOct 6, 2026

GeoPID: Decomposing and Steering Visual Information in Vision-Language Models

Authors: Seulgi Kim, Zhixiong Zhang, Xinwei Zhang, Jie Ling, Ronn Shaw

Organizations: Georgia Institute of Technology · Work done as an intern at Amazon. · Amazon

Abstract

While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63%.

Figures & tables

Appendix figures & tables29 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GeoWorld-VLM: Geometry from World Models for Vision-Language Models

    May 15, 2026Renjie Gu, Kaichen Zhou, Yan Luo +1Visual Spatial ReasoningVLM Adaptation

  2. Geometric Encoding for Spatial Reasoning in Vision-Language Models

    Sep 28, 2026Antonio Jun, Haoshui Yu, Zhengyi Lu +2Visual Spatial ReasoningVLM Reasoning

  3. When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models

    May 7, 2026Harshvardhan Saini, Samyak Jha, Yiming Tang +1Vision-Language ModelsObject Hallucination in VLMs