GeoPID: Decomposing and Steering Visual Information in Vision-Language Models
Organizations: Georgia Institute of Technology · Work done as an intern at Amazon. · Amazon
Abstract
While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we introduce a targeted intervention technique that selectively amplifies visual representations along the vision-unique subspace during inference. As a result, visual grounding capabilities were enhanced without any additional model parameter updates, achieving an average relative accuracy gain of 7.63%.
Figures & tables
| Model | POPE | MME | CVB | Hall | Reef | NatB | RWQA | MMB | SEED | AI2D | AOK | SQA | Rap4o | GenAI |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B ( Bai et al., 2025b ) | +1.0 | +0.0 | +2.9 | +3.6 | +0.5 | +2.7 | +0.7 | +0.2 | +1.0 | +1.0 | +0.2 | +2.5 | +0.8 | +4.7 |
| Qwen3-VL-8B ( Bai et al., 2025a ) | +0.5 | +3.1 | +0.0 | +0.6 | +0.2 | +0.0 | +0.0 | +0.0 | -0.5 | +1.0 | +0.2 | +0.2 | +5.9 | +4.3 |
| InternVL3-8B ( Zhu et al., 2025 ) | +0.2 | -0.8 | +1.6 | +4.2 | +0.0 | +1.3 | +1.3 | -0.5 | +0.7 | +0.5 | -0.2 | -1.0 | +2.3 | +11.0 |
| InternVL3-2B ( Zhu et al., 2025 ) | +0.2 | +1.8 | +2.3 | +0.6 | +0.2 | +3.8 | +4.7 | +1.1 | +0.7 | -0.8 | +0.2 | -0.5 | +1.1 | -5.7 |
| LLaVA-OV-7B ( Li et al., 2024c ) | +1.4 | +0.8 | -0.8 | +3.0 | +1.2 | +0.5 | -1.3 | +0.0 | +0.7 | -0.3 | +0.9 | +0.7 | +27.1 | +2.7 |
| Idefics2-8B ( Laurençon et al., 2024 ) | +0.0 | +0.5 | +1.8 | -1.2 | +2.6 | +1.1 | +1.3 | +0.2 | +0.0 | +1.6 | +1.1 | -0.5 | -1.1 | -8.7 |
| Model | POPE | MME | CVB | Hall | Reef | NatB | RWQA | MMB | SEED | AI2D | AOK | SQA | Rap4o | GenAI | Mean |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B ( Bai et al., 2025b ) | -0.2 | +0.0 | -0.5 | +0.6 | +0.2 | +2.7 | -0.7 | +0.5 | +0.7 | +1.3 | +0.0 | +2.5 | +0.8 | +4.7 | +1.05 |
| Qwen3-VL-8B ( Bai et al., 2025a ) | +0.0 | +0.0 | +0.0 | +0.6 | +0.0 | +0.0 | +0.7 | +0.0 | -1.7 | +0.8 | +0.0 | +0.2 | +5.1 | +0.7 | +0.49 |
| InternVL3-8B ( Zhu et al., 2025 ) | +0.5 | +0.3 | +0.5 | +0.6 | +0.0 | +0.0 | +1.3 | -0.2 | +0.5 | +0.3 | +0.0 | +0.0 | -1.1 | +0.0 | +0.15 |
| InternVL3-2B ( Zhu et al., 2025 ) | +0.0 | -0.3 | +0.3 | +0.6 | +0.2 | +1.6 | +1.3 | +0.0 | +0.5 | -0.8 | -0.2 | +0.0 | +1.1 | +2.0 | +0.59 |
| LLaVA-OV-7B ( Li et al., 2024c ) | +0.7 | +0.5 | -0.8 | +1.2 | +0.7 | +0.5 | -2.7 | +0.5 | +1.0 | -0.3 | +0.5 | +0.7 | +27.4 | -6.7 | +1.46 |
| Idefics2-8B ( Laurençon et al., 2024 ) | +0.5 | +0.0 | +1.6 | +0.0 | -0.2 | +0.8 | +0.7 | -0.5 | +0.7 | +0.3 | +0.5 | -0.5 | -1.7 | -1.0 | +0.17 |
Appendix figures & tables29 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | Size | Family / parent | Model type and emphasis |
|---|---|---|---|
| General-purpose instruction-tuned VLMs | |||
| Qwen2.5-VL-7B ( Bai et al., 2025b ) | 7B | Qwen2.5-VL | Image and document understanding with dynamic-resolution visual processing. |
| Qwen3-VL-8B ( Bai et al., 2025a ) | 8B | Qwen3-VL | General multimodal instruction following; multi-level visual feature integration. |
| InternVL3-8B ( Zhu et al., 2025 ) | 8B | InternVL3 | Native multimodal pre-training with broad visual and linguistic capabilities. |
| InternVL3-2B ( Zhu et al., 2025 ) | 2B | InternVL3 | Smaller-scale member of the same model family. |
| LLaVA-OV-7B ( Li et al., 2024c ) | 7B | LLaVA-OneVision | Unified single-image, multi-image, and video understanding. |
| Benchmark | Code | Task and visual content | Evaluation format |
|---|---|---|---|
| Broad multimodal understanding | |||
| MME ( Fu et al., 2026 ) | MME | Perception and cognition, including recognition and reasoning. | Binary answer selection |
| MMBench ( Liu et al., 2024b ) | MMB | Multiple visual perception and reasoning abilities. | Multiple-choice selection |
| SEED-Bench ( Li et al., 2023a ) | SEE | Visual comprehension across several evaluation dimensions. | Multiple-choice selection |
| Visual perception and hallucination | |||
| POPE ( Li et al., 2023c ) | POP | Presence or absence of objects in natural images. | Binary answer selection |
| Category | Accuracy drop (pp) | Difference from T (pp) |
|---|---|---|
| T | 3.9 | 0.0 |
| 19.3 | 15.4 | |
| V | 35.4 | 31.5 |
| J | 28.9 | 25.0 |
| Benchmark | Items | T | V | J | |
|---|---|---|---|---|---|
| POPE ( Li et al., 2023c ) | |||||
| CV-Bench ( Tong et al., 2024 ) | |||||
| RealWorldQA ( xAI, 2024 ) | |||||
| GenAI-Bench ( Li et al., 2024a ) | |||||
| Rapidata-4o ( Rapidata, 2026 ) | |||||
| HallusionBench ( Guan et al., 2024 ) |
| Mean | vs | |
|---|---|---|
| Model | ||||
|---|---|---|---|---|
| ViGoRL-7B ( Sarch et al., 2026 ) | ( ) | ( ) | ( ) | |
| ViGoRL-3B ( Sarch et al., 2026 ) | ( ) | ( ) | ( ) | |
| ViGoRL-MCTS-3B ( Sarch et al., 2026 ) | ( ) | ( ) | ( ) | |
| OpenVLThinker-7B ( Deng et al., 2025 ) | ( ) | ( ) | ( ) | |
| Vision-R1-7B ( Huang et al., 2026 ) | ( ) | ( ) | ( ) |
| Model / Benchmark | Matched | Same category | Different category | Blank |
|---|---|---|---|---|
| InternVL3-8B ( Zhu et al., 2025 ) / MMBench ( Liu et al., 2024b ) | 0.4150 | 0.0323 | 0.0274 | 0.0149 |
| Qwen2.5-VL-7B ( Bai et al., 2025b ) / POPE ( Li et al., 2023c ) | 0.1882 | 0.0106 | 0.0055 | 0.0010 |
| ViGoRL-3B ( Sarch et al., 2026 ) / MMBench ( Liu et al., 2024b ) | 0.0098 | 0.0033 | 0.0031 | 0.0019 |
| Benchmark | |||||
|---|---|---|---|---|---|
| MMBench ( Liu et al., 2024b ) | 0.0483 | 0.0063 | 0.0595 | 0.0566 | 0.3001 |
| A-OKVQA ( Schwenk et al., 2022 ) | 0.0501 | 0.0067 | 0.0435 | 0.0408 | 0.2973 |
| SEED-Bench ( Li et al., 2023a ) | 0.0340 | 0.0052 | 0.0367 | 0.0308 | 0.1940 |
| POPE ( Li et al., 2023c ) | 0.0459 | 0.0153 | 0.0386 | 0.0385 | 0.1860 |
| Reefknot ( Zheng et al., 2025 ) | 0.0341 | 0.0034 | 0.0275 | 0.0192 | 0.1844 |
| MME ( Fu et al., 2026 ) | 0.0349 | 0.0111 | 0.0358 | 0.0355 | 0.1738 |
| Benchmark | |||||
|---|---|---|---|---|---|
| POPE Li et al. (2023c) | 0.183 | 0.036 | 0.147 | 0.064 | 0.005 |
| MME ( Fu et al., 2026 ) | 0.201 | 0.061 | 0.140 | 0.050 | 0.009 |
| MMBench ( Liu et al., 2024b ) | 0.309 | 0.044 | 0.265 | 0.047 | 0.006 |
| HallusionBench Guan et al. (2024) | 0.109 | 0.016 | 0.093 | 0.015 | 0.003 |
| SEED-Bench ( Li et al., 2023a ) | 0.166 | 0.045 | 0.121 | 0.029 | 0.006 |
| AI2D ( Kembhavi et al., 2016 ) | 0.200 | 0.051 | 0.149 | 0.043 | 0.010 |
| Benchmark / outcome | ||||||
|---|---|---|---|---|---|---|
| POPE ( Li et al., 2023c ) | +0.14 | +0.59 | +0.40 | +0.46 | +0.50 | -0.58 |
| MME ( Fu et al., 2026 ) | +0.77 | +0.31 | +0.84 | +0.57 | +0.75 | -0.56 |
| MMBench ( Liu et al., 2024b ) | +0.53 | +0.25 | +0.57 | +0.65 | +0.63 | -0.64 |
| HallusionBench ( Guan et al., 2024 ) | +0.50 | +0.37 | +0.83 | +0.80 | +0.63 | -0.80 |
| Rapidata-4o ( Rapidata, 2026 ) | +0.07 | +0.26 | +0.25 | +0.33 | +0.30 | -0.36 |
| GenAI-Bench ( Li et al., 2024a ) | +0.81 | +0.57 | +0.90 | +0.83 | +0.89 | -0.77 |
| Aggregation | Included models | Count | Spearman |
|---|---|---|---|
| Model-level | All | 22 | 0.79 |
| Model-level | Excluding ViGoRL ( Sarch et al., 2026 ) | 20 | 0.68 |
| Cell-level | All | 308 | 0.73 |
| Cell-level | Excluding ViGoRL ( Sarch et al., 2026 ) | 340 | 0.67 |
| Benchmark inclusion | Mean standardized (correct V/J correct T) | CI | Benchmarks with higher for correct V/J than correct T | |
|---|---|---|---|---|
| correct T | ||||
| All with both types |
| , , | , , , | , , Boosted | full geometry, Boosted | ||
|---|---|---|---|---|---|
| (Spearman) | (Pearson) | (Pearson) | (Pearson) | (Pearson) | (Pearson) |
| All | ||||
|---|---|---|---|---|
| Question category | True image | Deranged image | Blank image |
|---|---|---|---|
| : text decides | |||
| : text prior | |||
| : image reading | |||
| : joint evidence |
| Statistic | Share |
|---|---|
| Most frequent predicted option before intervention | |
| Most frequent predicted option after intervention | |
| Most frequent ground-truth option |
| Statistic | Value |
|---|---|
| Eligible model–item pairs | |
| Correct after intervention | |
| Expected correctness under random switching |
| Amplified target | Rank | Gain | Correction | Damage |
|---|---|---|---|---|
| Vision-unique | ||||
| Shared | — | |||
| Text-unique | ||||
| Synergy-evaluation space | ||||
| Random hidden-space basis | ||||
| Random text-complement basis |
| Decision | Rule |
|---|---|
| Edit target | Vision-unique basis on visual tokens. |
| Open-branch strength | ; or when blank-image agreement . |
| Closed-branch strength | , a weak edit rather than the unedited baseline. |
| Intervention selection rule | Mean gain (pp) |
|---|---|
| Blank-image agreement with geometry-based depth selection | |
| Blank-image agreement with for agreement | |
| GeoPID geometry-based rule (Figure 4 ) |
| Rank cap | Depth | |||
|---|---|---|---|---|
| Mismatch standard deviation | Reported SE of | Median | |
|---|---|---|---|
| Ablation | Setting | Rank correlation |
| Kernel | Cosine reference | |
| RBF | ||
| Standardized linear | ||
| Rényi order | Shannon limit | |
| (reference) |