Many multimodal tasks, such as image captioning and visual question answering, require vision-language models (VLMs) to bind objects with their properties and spatial relations. Yet it remains unclear where and how such associations are computed within VLMs. In this work, we show that VLMs rely on two concurrent mechanisms to represent spatial variable binding. In the language model backbone, intermediate layers represent content-independent spatial relations on top of visual tokens corresponding to objects. However, this mechanism plays only a secondary role in shaping model predictions. Instead, the dominant source of spatial information originates in the vision encoder, whose representations encode the layout of objects and are directly exploited by the language model backbone. Notably, this spatial signal is distributed globally across visual tokens, extending beyond object regions into surrounding background areas. We validate the generalization of our findings to complex natural images from the COCO dataset, where globally amplifying the vision-derived spatial representations across all image tokens corrects spatial variable binding failures across models of various sizes. Together, our results clarify how spatial variable binding is computed within VLMs and highlight the central role of vision encoders in enabling it.
Figures & tables
Figure 1: Experimental Settings. We study the internal mechanism responsible for spatial variable binding across three synthetic settings (Squares, Shapes, Objects) and one natural controlled settings (What’sUp). We further show that our findings generalize to complex natural scenes from the COCO dataset with complex scenes and objects, where a simple intervention corrects spatial binding failures.
Figure 2: Last-Token Position Patching. The residual stream at the final token (“is”) of the counterfactual run is patched into the clean run. Clean answer: Red . If ordering information is transferred, the output becomes Blue (the clean square at the counterfactual’s correct position); if attribute is transferred, it becomes Black (the counterfactual answer).
Figure 3: Last Token Position Patching Results : The output of the clean run remains unchanged up to layer 19. From layer 20 through layer 22, the output changes to the color of the clean square corresponding to the correct position of the counterfactual run, suggesting that these layers patch the order of the target square. Beyond layer 22, the model predicts the counterfactual answer, indicating that layers after 22 encode the final answer. Shaded areas are 95% bootstrap confidence intervals.
Figure 4: Position information in visual embeddings extends beyond object regions. We train linear probes on top of object tokens of the visual embedding to predict their order in the image. At test time, these probes generalize to background tokens, producing strips-like pattern aligned with object position, suggesting that background tokens encode order information of nearby objects. See App. C.9 for more complex arrangements.
Figure 5: Causal intervention on visual token embeddings. The visual tokens of the left and right squares are swapped (left ↔ right) from the counterfactual run into the clean run, either (a) for the square tokens only or (b) together with their background strips. If the patched tokens carry ordering information, the output switches to the square on the opposite side of the queried reference.
Figure 6: Object patching . Interchange interventions applied only to visual tokens corresponding to the squares. Patching square tokens at any layer does not reliably change the model’s final output relative to the clean run, indicating that square-localized tokens alone do not carry sufficient ordering information to determine the model’s prediction.
Figure 8: Ordering information is ablated from vision embeddings by replacing square embeddings with same-colored isolated middle squares and background embeddings with those from an empty image, preserving color while removing position cues.
Figure 9: Patching results after removing ordering information from vision encoder. The model’s output switches only in intermediate layers within the range of 11-20, indicating that the LM backbone generates ordering information even when vision-derived ordering representations are absent. Panel titles give the per-direction n , which is capped by the accuracy of the ablated model.
Table 2: Accuracy of all five models on the studied tasks.
Squares
Shapes
Objects
L
R
A
B
L
R
A
B
L
R
A
B
Qwen2-VL-7B-Instruct
1.00
1.00
1.00
0.99
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
Qwen2-VL-2B-Instruct
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
0.96
0.99
0.95
0.92
Qwen3-VL-8B-Instruct
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
1.00
Gemma-3-4b-it
1.00
1.00
0.93
0.97
1.00
1.00
0.97
0.97
1.00
1.00
0.95
1.00
Pixtral-12B
0.59
0.57
0.50
0.54
0.25
0.57
0.68
0.69
0.61
0.57
0.67
0.67
Appendix
Table 3: Full behavioral performance, according to queried directions (Left, Right, Above, Below)
Figure 10: Example images of the Square, Shapes, and Objects settings.
Model
# Image tokens
# Pixels / Token
# Tokens / Object
Object spacing
Strip width
Qwen2-VL-7B-Instruct
12×12
28×28
2×2
2
4
Qwen2-VL-2B-Instruct
12×12
28×28
2×2
2
4
Qwen3-VL-8B-Instruct
12×12
32×32
2×2
2
4
Gemma-3-4b-it
16×16
56×56
3×3
3
3
Pixtral-12B
15×15
16×16
2×2
2
2
Appendix
Table 4: Image tokenization and spatial layout parameters for each model. Qwen2-VL-2B shares the 7B vision encoder (14-px patches merged 2×2 ), and Qwen3-VL-8B uses 16-px patches merged 2×2 , so both produce the same 12×12 grid as Qwen2-VL-7B. Pixtral’s encoder has a fixed 16-px patch stride, so a 240-px canvas yields a 15×15 grid.
Figure 11: Examples for images from What’sUp control ( Kamath et al., 2023 )
Figure 12: Examples for images from COCO-spatial ( Kamath et al., 2023 ; Lin et al., 2014 )
Probe
Random
Dataset
Model
Axis
nval/ntest
α
Held-out fix
α
Held-out fix
COCO-spatial
Qwen2-VL-2B
horizontal
75/31
8
32.3% (10/31)
15
16.1% (5/31)
COCO-spatial
Qwen2-VL-2B
vertical
49/21
7
38.1% (8/21)
12
19.0% (4/21)
COCO-spatial
Gemma-3-4b
horizontal
97/41
14
22.0% (9/41)
15
17.1% (7/41)
COCO-spatial
Gemma-3-4b
vertical
50/21
15
28.6% (6/21)
10
0.0% (0/21)
What’sUp
Qwen2-VL-2B
horizontal
78/33
13
45.5% (15/33)
15
0.0% (0/33)
Appendix
Table 5: Held-out generalization of the amplification coefficient. α is selected on validation failures only (seed 0, ≈ 70/30 split) and evaluated on the held-out failures; “fix” is the fraction of held-out failures corrected. Gemma-3 COCO-spatial rows come from a run on a 394-item subset of COCO-spatial (baseline 185/394=47.0% ), so their failure counts differ from Table 11 .
Figure 13: Final token interchange intervention experiments with the Qwen2-VL-7B-Instruct model.
Figure 14: Final token interchange intervention experiments with the Gemma-3-4b-it model.
Figure 15: Final token interchange intervention experiments with the Qwen2-VL-2B-Instruct model.
Figure 16: Final token interchange intervention experiments with the Qwen3-VL-8B-Instruct model.
Figure 17: Final token interchange intervention experiments with the Pixtral-12B model ( n=50 pairs; see text).
Squares
Shapes
Objects
Qwen2-VL-7B-Instruct
0.60
0.62
0.43
Qwen2-VL-2B-Instruct
0.62
0.51
0.19
Gemma-3-4b-it
0.57
0.64
0.54
Pixtral-12B
0.41
0.78
0.42
Appendix
Table 6: Behavioral performance after removing spatial information from vision embeddings (chance is 0.33 ). Pixtral is evaluated in bf16.
Figure 18: Visual token embeddings of square tokens interchange intervention experiment on Qwen2-VL-7B-Instruct model.
Figure 19: Visual token embeddings of square tokens interchange intervention experiment on Gemma-3-4b-it model.
Figure 20: Visual token embeddings of strip tokens interchange intervention experiment on Qwen2-VL-7B-Instruct model.
Figure 21: Visual token embeddings of strip tokens interchange intervention experiment on Gemma-3-4b-it model.
Figure 24: Visual token embeddings of square (object) tokens interchange intervention experiment on the Qwen2-VL-2B-Instruct model.
Figure 25: Visual token embeddings of strip tokens interchange intervention experiment on the Qwen2-VL-2B-Instruct model.
Figure 26: Visual token embeddings of square (object) tokens interchange intervention experiment on the Qwen3-VL-8B-Instruct model.
Figure 27: Visual token embeddings of strip tokens interchange intervention experiment on the Qwen3-VL-8B-Instruct model.
Figure 28: Visual token embeddings of square (object) tokens interchange intervention experiment on the Pixtral-12B model ( n=50 ).
Figure 29: Visual token embeddings of strip tokens interchange intervention experiment on the Pixtral-12B model ( n=50 ).
Figure 30: Square-token patching after removing vision-derived ordering information for Qwen2-VL-2B-Instruct (top three rows) and Qwen3-VL-8B-Instruct (bottom three rows). Empty panels are directions for which the ablated model answers no clean–counterfactual pair correctly.
Figure 31: Square-token patching after removing vision-derived ordering information, Pixtral-12B ( n=50 runs; the per-direction n in the panel titles is capped by the accuracy of the ablated model and is very small for several cells). The flip occurs in LM layers 8–17, the same window as the strip-patching flip in Fig. 29 .
Experiment
Model
Cells (tested)
Median ∣Δ∣
Max ∣Δ∣
p<.05 / q<.05
CI overlap
Last-token patching (Fig. 3 )
Qwen2-VL-7B
14 (14)
0.009
0.018
1 / 0
14/14
Gemma-3-4b
14 (14)
0.003
0.013
0 / 0
13/14
Object patching (Fig. 7 )
Qwen2-VL-7B
14 (14)
0.044
0.190
0 / 0
14/14
Gemma-3-4b
14 (14)
0.061
0.234
2 / 0
13/14
Strip patching (Fig. 7 )
Qwen2-VL-7B
14 (14)
0.005
0.041
1 / 0
13/14
Gemma-3-4b
14 (14)
0.012
0.119
0 / 0
14/14
Appendix
Table 7: n=50 vs. n=100 peak-layer effects. ∣Δ∣ : absolute change in the peak effect; p<.05 : cells with an uncorrected permutation p<0.05 ; q<.05 : cells significant after Benjamini–Hochberg correction across all 119 tests; overlap: cells whose 95% CIs overlap.
Figure 32: Diagonal arrangements, Qwen2-VL-7B-Instruct ( n=100 pairs, 95% bootstrap CIs). Rows: last-token patching (colors as in Fig. 3 ), strip patching, square patching, and square patching after the vision-ordering ablation (colors as in Fig. 7 ). Columns: the two principal-diagonal relations followed by the two secondary-diagonal relations.
Figure 33: Position-probe predictions on the diagonal (left) and anti-diagonal (right) Squares settings for Qwen2-VL-7B-Instruct . As in Fig. 4 , probes trained only on square tokens generalize to background tokens along diagonal strips.
Figure 34: Training a single probe for both layouts learns ordinal order information. The overlapping square appears at the same absolute image location in both layouts, while its ordinal label changes from 2 in the left-shifted layout to 0 in the right-shifted layout. We plot the probabilities assigned to the two candidate labels, P(0) (solid) and P(2) (dashed). For this square only, chance is shown as a dotted line. The correct curve is dashed in the top row and solid in the bottom row. Although both squares occupy the same absolute image location, the probe correctly distinguishes their ordinal positions, demonstrating that ordinal information is linearly decodable.
Gemma 3
Qwen2-VL
Layer
Left
Right
Mean
Left
Right
Mean
0
84.4
77.3
80.8
25.8
19.5
22.6
5
91.6
93.1
92.4
81.3
81.3
81.3
10
99.6
98.9
99.3
99.9
99.9
99.9
15
99.7
99.5
99.6
99.1
99.8
99.4
20
97.8
98.2
98.0
98.7
99.7
99.2
Appendix
Table 8: Overall probe accuracy by layout. Test accuracy is computed over all visual tokens corresponding to the squares. Mean averages the two layouts.
Figure 35: Training a probe on one layout does not generalize to the other, suggesting the presence of absolute position information. The overlapping square appears at the same absolute image location in both layouts, while its ordinal label changes from 2 in the left-shifted layout to 0 in the right-shifted layout. We plot the probabilities assigned to the two candidate labels, P(0) (solid) and P(2) (dashed). For this square only, chance is shown as a dotted line. A probe trained on one layout fails to correctly classify the overlapping square in the other layout, indicating that it relies on absolute image coordinates rather than purely ordinal position.
Gemma 3
Qwen2-VL
Layer
Train
Test
Train
Test
0
100.0
0.0
51.0
17.0
5
100.0
1.0
100.0
48.0
10
100.0
2.0
100.0
49.0
15
100.0
10.0
100.0
36.0
20
100.0
8.0
100.0
59.0
Appendix
Table 9: Cross-layout probe accuracy across layers. The probe is trained on square tokens from the left-shifted layout and evaluated on square tokens from the right-shifted layout.
Gemma-3-4b-it (SigLIP, absolute pos. emb.)
Qwen2-VL-7B (2D-RoPE)
Layer
Held-out acc.
Square
Strip
Outside
Held-out acc.
Square
Strip
Outside
0
100.0
0.97
0.90
0.20
29.4
0.32
0.33
0.33
5
100.0
0.98
0.67
0.24
91.3
0.68
0.65
0.27
10
100.0
0.98
0.76
0.23
100.0
0.96
0.74
0.25
Appendix
Table 10: Layer-wise vision-encoder probing on the Squares setting (horizontal arrangement; the vertical arrangement behaves the same). A three-way probe is trained on square tokens at each vision-encoder layer. “Square”, “Strip”, and “Outside” give the mean probability the probe assigns to a label on that label’s square tokens, on the background tokens of the same column excluding the square, and on all other tokens, averaged over the three labels and the held-out images.
Figure 36: Strip emergence across vision-encoder layers (Squares, horizontal). Each row applies the probe trained at that vision-encoder layer to all tokens; columns are the three position labels. Left: Gemma-3-4b-it , where the strips are present at layer 0. Right: Qwen2-VL-7B-Instruct , where layer 0 carries no position information and the strips form over the first five to ten layers.
Figure 37:
Figure 38: Label probabilities after the intervention : Intervening on the visual token embeddings generated by the vision encoder using the mirrored counterfactual produces substantial label probabilities for both the absolute and relative expected outputs.
Figure 39: Label probabilities after intervening on the LM backbone’s ordering information : Repeating the mirrored-counterfactual intervention under the vision-ordering ablation of Sec. 5.3.2 , where only the LM backbone’s own ordering representation can be transferred, produces the relative-order answer almost exclusively, with little probability on the absolute-position answer.
Figure 40: Absolute vs. relative position, per-layer curves, Qwen2-VL-7B-Instruct .
Figure 41: Absolute vs. relative position, per-layer curves, Gemma-3-4b-it .
Figure 42: Absolute vs. relative position, per-layer curves, Qwen2-VL-2B-Instruct .
Gemma3-4B
Qwen2-VL-2B
Pixtral-12B
Direction
Intervention
Acc.
% Corr. Fail.
Acc.
% Corr. Fail.
Acc.
% Corr. Fail.
Horizontal
None
0.51
–
0.62
–
0.88
–
Random
0.54
8.0%
0.67
12.3%
0.88
0.0%
Random*
0.56
10.1%
0.67
12.3%
–
–
Probe
0.62
23.9%
0.80
47.2%
0.92
30.3%
Probe*
0.70
39.1%
0.84
58.5%
–
–
Appendix
Table 11: Full breakdown of amplification results of Sec 5.4 by spatial direction. “% Corr. Fail.” denotes the fraction of baseline failures corrected by the intervention. In Random and Probe , the best single fixed α coefficient is chosen across all examples; In Random* and Probe* , we search for a per-item optimal coefficient, serving as an upper bound on intervention performance. For Pixtral-12B ( n=439 ; 33 horizontal and 11 vertical failures) the per-axis breakdown of the per-item oracle is not available and is marked “–”.
Gemma3-4b
Qwen2-VL-2B
Pixtral-12B
Intervention
Acc.
% Corr. Fail.
Acc.
% Corr. Fail.
Acc.
% Corr. Fail.
None
0.79
–
0.82
–
0.97
–
Random
0.83
19.1%
0.84
9.9%
0.97
6.7%
Random*
0.84
23.2%
0.85
13.5%
0.97
6.7%
Probe
0.91
55.2%
0.90
43.2%
0.98
40.0%
Probe*
0.93
68.0%
0.91
47.7%
0.99
53.3%
Appendix
Table 12: Amplification results on the What’sUp dataset using the strict two-ordering evaluation protocol. “% Corr. Fail.” denotes the fraction of baseline failures corrected by the intervention. In Random and Probe , the best single fixed α coefficient is chosen across all examples; In Random* and Probe* , we search for a per-item optimal coefficient, serving as an upper bound on intervention performance. Pixtral-12B is evaluated on 484 composite queries (15 strict failures).
Figure 43: Examples of the original What’sUp images used for the on/under and front/behind relations ( Kamath et al., 2023 ) . Each image contains a reference object (table, sunglasses, spoon) and a target object placed on or under it (left two) or in front of it (right two).
Relation
Model
Baseline
Method
α
Fix rate [95% CI]
Acc.
Held-out fix
on/under
Qwen2-VL-2B
72.7%
Random
7
13.6% [4.5, 25.0]
76.4%
15.0%
Probe
8
68.2% [52.3, 81.8]
91.3%
65.0%
Gemma-3-4b
65.8%
Random
14
16.4% [7.3, 27.3]
71.4%
10.0%
Probe
12
52.7% [40.0, 65.5]
83.9%
45.0%
front/behind
Qwen2-VL-2B
73.0%
Random
15
27.3% [16.4, 40.0]
80.4%
30.0%
Probe
12
76.4% [63.6, 87.3]
93.6%
70.0%
Appendix
Table 13: Amplification on additional What’sUp relations. Fix rates are over strict-wrong images (95% bootstrap CIs over the failures); “Acc.” is strict accuracy on all images after amplification. Held-out: α selected on validation failures and evaluated on 20 held-out failures. Front/behind for Gemma-3 was not run.
Transfer
Target failures
Probe acc. (train / test / target)
Random *
Transferred
Transferred *
In-domain
(i) What’sUp → COCO horiz.
106
97.4 / 95.2 / 85.7
16.0%
48.1%
55.7%
47.2%
(ii) COCO horiz. → What’sUp
111
93.7 / 89.1 / 92.4
11.7%
32.4%
34.2%
43.2%
(iii) COCO horiz. → COCO vert.
–
98.9 / 94.1 / 60.0
1.9%
33.3%
33.3%
35.2%
(iv) Shared 4-class → all COCO
176
90.2 / 64.5 / –
22.7%
42.0%
54.0%
42.6%
Appendix
Table 14: Probe transfer on Qwen2-VL-2B-Instruct . Probe accuracy is given on the source train / source test / target tokens. Fix rates are over the target’s strict failures with the best fixed α ( * : per-item oracle). The in-domain column repeats the fixed- α result of a probe trained on the target itself (Tables 11 and 12 ; for (iii) the in-domain vertical probe from the same run).
Qwen2-VL-2B
Gemma-3-4b
Acc.
Fix
Break
Acc.
Fix
Break
Baseline (one-word prompt)
0.600
–
–
0.527
–
–
Ordering amp., α=6 , full re-evaluation
0.709
36.9%
6.4%
0.536
22.6%
18.5%
Ordering amp., per-item α (Table 1 )
0.82
54.5%
–
0.72
40.2%
–
CoT prompting, full re-evaluation
0.752
65.3%
18.2%
0.464
35.6%
44.0%
Appendix
Table 15: Chain-of-thought prompting vs. ordering amplification on COCO-spatial (strict two-ordering accuracy on all 440 items; chance =0.25 ). “Fix”: fraction of baseline failures corrected; “break”: fraction of baseline-correct items made incorrect. Qwen2-VL-2B is evaluated in the environment of Table 1 ( transformers 4.47.1).
Benchmark
Qwen2-VL-2B: α=0→α=6 ( Δ )
Gemma-3-4b: α=0→α=6 ( Δ )
COCO-spatial failures repaired (in-domain)
0 → 37.5 ( +37.5 )
0 → 23.1 ( +23.1 )
POPE
89.0 → 85.6 ( −3.4 )
86.0 → 83.2 ( −2.8 )
GQA
61.6 → 56.6 ( −5.0 )
37.0 → 38.6 ( +1.6 )
MMBench
78.2 → 77.4 ( −0.8 )
65.2 → 64.4 ( −0.8 )
OK-VQA
59.0 → 58.4 ( −0.6 )
52.0 → 50.3 ( −1.7 )
VizWiz
70.1 → 68.7 ( −1.4 )
48.2 → 44.3 ( −3.9 )
Appendix
Table 16: Accuracy (%) on non-spatial VQA benchmarks before ( α=0 ) and after ( α=6 ) both-axis ordering amplification; n=500 items per benchmark. The first row gives the in-domain repair rate of COCO-spatial baseline failures.
Figure 44: Dose–response of ordering amplification. In-domain repair of COCO-spatial failures (axis-specific probe) and accuracy on non-spatial VQA benchmarks as a function of α , for Qwen2-VL-2B (left) and Gemma-3-4b (right). The paper’s operating point is α=6 .
Figure 45: Horizontal position probe for square images, Qwen
Figure 46: Vertical position probe for square images, Qwen
Figure 47: Horizontal position probe for shape images, Qwen
Figure 48: Vertical position probe for shape images, Qwen
Figure 49: Horizontal position probe for object images, Qwen
Figure 50: Vertical position probe for object images, Qwen
Figure 51: Horizontal position probe for square images, Gemma
Figure 52: Vertical position probe for square images, Gemma
Figure 53: Horizontal position probe for shape images, Gemma
Figure 54: Vertical position probe for shape images, Gemma
Figure 55: Horizontal position probe for object images, Gemma
Figure 56: Vertical position probe for object images, Gemma
Figure 57: Position-probe predictions for Qwen2-VL-2B-Instruct on the Squares (horizontal, vertical) and Shapes (horizontal, vertical) settings.
Figure 58: Position-probe predictions for Qwen2-VL-2B-Instruct on the Objects (horizontal, vertical) and What’sUp (horizontal) settings.
Figure 59: Position-probe predictions for Qwen3-VL-8B-Instruct on the Squares (horizontal, vertical) and Shapes (horizontal, vertical) settings.
Figure 60: Position-probe predictions for Qwen3-VL-8B-Instruct on the Objects (horizontal, vertical) and What’sUp (horizontal) settings.
Figure 61: Qwen and Gemma probing results on 2x2 shapes grid.
Figure 62: Qwen and Gemma probing results on 3x3 shapes grid.
Figure 63: Preprocessing for What’sUp images: constructing bounding boxes around each object.
Figure 64: Qwen What’sUp probe
Figure 66: Horizontal position probe for square images with real background, Qwen
Figure 67: Vertical position probe for square images with real background, Qwen
Figure 68: Horizontal position probe for shapes images with real background, Qwen
Figure 69: Vertical position probe for shapes images with real background, Qwen
Figure 70: Horizontal position probe for objects images with real background, Qwen
Figure 71: Vertical position probe for objects images with real background, Qwen
Figure 72: Horizontal position probe for square images with real background, Gemma
Figure 73: Vertical position probe for square images with real background, Gemma
Figure 74: Horizontal position probe for shapes images with real background, Gemma
Figure 75: Vertical position probe for shapes images with real background, Gemma
Figure 76: Horizontal position probe for objects images with real background, Gemma
Figure 77: Vertical position probe for objects images with real background, Gemma
Figure 78: Interchange interventions applied only to visual tokens corresponding to the squares in Qwen2-VL-7B-Instruct model, with real background images.
Figure 79: Interchange interventions applied only to visual tokens corresponding to the strip in Qwen2-VL-7B-Instruct model, with real background images.
Figure 80: Interchange interventions applied only to visual tokens corresponding to the squares in Gemma-3-4b-it model, with real background images.
Figure 81: Interchange interventions applied only to visual tokens corresponding to the strip in Gemma-3-4b-it model, with real background images.
Vision-language models can describe an image with remarkable accuracy, yet a more fundamental question remains unanswered: what visual information actually drives their answers? In this work, we investigate this question through causal tracing, and we observe that highly causal vision tokens often lie outside the target region. Extending the analysis to larger vision-language models reveals a similar pattern across models and corruption settings, suggesting that strong multimodal performance does not necessarily imply spatially localized causal representations. We further investigate: can these models preserve visual structure when appearance cues are removed? and find that visual cues are exploited to understand visual structures. Together, our experiments expose a gap between seeing, using, and reasoning over visual structure, and provide a causal framework for studying how visual information is transformed, preserved, and ultimately used by modern vision-language models.
Naren Kumar S, Tirth Bhatt, Mayank Singh
LINGO Research Group, Indian Institute of Technology Gandhinagar, India
Vision-language models (VLMs) achieve strong performance on spatial reasoning benchmarks, yet it remains unclear whether this reflects structured 3D understanding or reliance on statistical shortcuts in natural images. We introduce a representation-level analysis framework that constructs minimal contrastive pairs to measure how spatial axes are organized and disentangled within VLM embeddings. Our analysis across multiple model families reveals a consistent vertical-distance entanglement: models conflate vertical image position with distance, mirroring the perspective bias of natural photographs. This bias produces a significant accuracy gap between perspective-consistent and counter-heuristic examples, and intensifies under data scaling even as overall benchmark accuracy improves. We further show that models with similar benchmark scores can exhibit different internal representations, and that these differences predict accuracy and robustness across diverse spatial reasoning benchmarks. To isolate this bias from evaluation-set skew, we introduce SpatialTunnel, a synthetic benchmark designed to expose spatial shortcut biases by removing common correlations present in natural images. Experiments confirm that the entanglement is model-intrinsic, and that models with well-separated spatial axes exhibit greater robustness, suggesting that well-structured spatial representations lead to more reliable spatial reasoning across diverse benchmarks. Code and benchmark are available on the project page: https://cheolhong0916.github.io/whyfarlooksup.github.io/.
Cheolhong Min, Jaeyun Jung, Daeun Lee +5
Seoul National University · The Ohio State University · NVIDIA
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \faGithub~spatio-lm.
Jing Wu, Jianhua Wu, Jiayi Guan +5
Xiaomi EV, Beijing, China · College of Automotive and Energy Engineering, Tongji University, Shanghai, China · Independent Researcher