High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.
Figures & tables
Figure 1: Overview of the task and mechanistic analysis. (A) A scene, given as an image or an equivalent text description (the source ), and a query about the position of a target object relative to a reference object. (B) Object-location swaps and target/reference reversals, both reversing the correct answer. (C) Matched input conditions: VLM + image, VLM + text and LLM + text. (D) Hypothesized mechanism: source representations influence query-side location information, while target/reference roles condition its use in relation prediction. We test this account through activation patching and targeted interventions, and evaluate role-direction steering on natural images.
Figure 2: Benchmarking results for the three conditions: VLM+image, VLM+text and LLM+text. Numerical results are provided in Table 6 .
Figure 3: Layerwise activation patching clean-answer recovery rate across input conditions, token sources, corruption types and models in the 3-object synthetic dataset.
Setting
Patch group
LLaVA Δm / flip
InternVL Δm / flip
Pixtral Δm / flip
VLM + image
All visual
1.718 / 62.9%
21.660 / 69.1%
3.877 / 51.2%
VLM + image
Object+strip
0.881 / 34.0%
9.691 / 11.6%
1.477 / 5.1%
VLM + text
All objects
0.493 / 17.7%
7.912 / 24.3%
0.783 / 13.7%
LLM + text
All objects
0.416 / 20.0%
3.586 / 15.6%
0.701 / 10.3%
Table 1: Effects of source patching on query-side location information and answer preferences. Cells report all-layer mean corrupt-directed score shifts and corrupt-answer flip rates after patching corrupt-run source states into the clean run ( Δm→corr / flip %).
Figure 4: Location IDs in LLaVA-1.6-Mistral-7B on three-object scenes: (A) held-out location prediction with shuffled-label controls; (B–D) prototype visualization along axes defined by endpoint-ID contrasts; and (E) query-side relation-sign prediction with orthogonal-axis controls.
Figure 5: Layer-wise query-side location-ID steering measured by belief-swap rate.
Figure 6: Joint TF/RF role-contrast geometry across layers. High held-out mean cosine and norm ratio indicate that the joint object-level contrasts share a stable direction.
Figure 7: Layer-wise joint role-direction intervention at α=1 . Curves show the mean gap ΔMrole−ΔMorth . Negative values indicate stronger disruption along the joint role direction.
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
LLaVA-1.6
VLM+image
82.2→84.0(+1.8)
70.6→73.9(+3.3)
66.3→67.5(+1.3)
LLaVA-1.6
VLM+text
66.3→65.1(−1.2)
52.2→54.9(+2.7)
39.7→39.2(−0.5)
LLaVA-1.6
LLM+text
63.8→67.1(+3.3)
38.6→43.9(+5.3)
33.5→38.2(+4.7)
InternVL3.5
VLM+image
97.3→97.4(+0.1)
96.1→96.6(+0.5)
94.2→94.2(+0.0)
InternVL3.5
VLM+text
87.1→89.7(+2.6)
78.7→80.9(+2.2)
74.1→77.8(+3.7)
InternVL3.5
LLM+text
70.8→73.2(+2.4)
46.9→50.3(+3.3)
45.3→49.7(+4.4)
Table 2: Activation steering on What’sUp-A using the joint target/reference role direction estimated on synthetic data. Bold indicates improvements with 95% confidence intervals excluding zero.
Figure 8: Example data used in our experiments
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Split
Scenes
Questions
Horizontal
Vertical
Each relation
Reciprocal pairs
Synthetic 2-object
Train
1,512
3,024
1,512
1,512
756
1,512
Validation
324
648
324
324
162
324
Test
648
1,296
648
648
324
648
Total
2,484
4,968
2,484
2,484
1,242
2,484
Synthetic 3-object
Train
504
3,024
1,512
1,512
756
1,512
Validation
108
648
324
324
162
324
Appendix
Table 3: Statistics of the synthetic datasets used for mechanistic analyses. A scene denotes a unique rendered image, and a question denotes one directed target–reference query. “Each relation” reports the number of questions for each of { left , right , above , below }. Each two-object scene contains two reciprocal questions and therefore one target–reference pair. Each three-object scene contains all six directed questions among its three objects, corresponding to three reciprocal target–reference pairs.
VLM
Arch. type
LLM backbone
Layers
LLaVA-v1.5-7B
Projector-concat
Vicuna v1.5 (7B)
32
LLaVA-v1.5-13B
Projector-concat
Vicuna v1.5 (13B)
40
LLaVA-v1.6-Mistral-7B
Projector-concat
Mistral (7B)
32
LLaVA-v1.6-Vicuna-7B
Projector-concat
Vicuna v1.5 (7B)
32
LLaVA-v1.6-Vicuna-13B
Projector-concat
Vicuna v1.5 (13B)
40
Pixtral-12B
Projector-concat
Mistral (12B)
40
Appendix
Table 4: Architectures of benchmarked models. Layers denotes the number of transformer blocks in the language backbone.
Dataset
Condition
Original images
Scene families
Original queries
Reversal pairs
Swap anchors
Synthetic 2-object
VLM+image
128
128
256
128
256
VLM+text
128
128
512
256
512
LLM+text
128
128
512
256
512
Synthetic 3-object
VLM+image
128
128
256
128
256
VLM+text
128
128
512
256
512
LLM+text
128
128
512
256
512
Appendix
Table 5: Evaluation counts for Figure 2 and Table 6 . Counts are per evaluated model; text conditions pool the available description variants.
Model
Accuracy
Target–reference reversal consistency
Location-swap consistency
VLM+ image
VLM+ text
LLM+ text
VLM+ image
VLM+ text
LLM+ text
VLM+ image
VLM+ text
LLM+ text
(a) Synthetic
LLaVA-1.5-7B
92.58
56.84
50.00
85.94
48.44
0.00
83.59
47.66
0.00
LLaVA-1.5-13B
97.85
57.13
50.00
95.70
34.57
0.00
95.90
33.50
0.00
LLaVA-1.6-Mistral-7B
99.22
63.77
44.92
98.44
46.09
24.80
96.68
43.36
18.95
LLaVA-1.6-Vicuna-7B
96.29
56.05
50.00
92.58
50.39
0.00
92.97
48.44
0.00
Appendix
Table 6: Numerical benchmarking results for VLM+image, VLM+text, and the corresponding LLM+text backbone. Synthetic includes the two- and three-object 2D sets, and What’sUp includes subsets A and B.
Stage
Legend label
Tokens
source
all visual tokens
every image token (C1 only)
source
object bbox pairs
image tokens inside the bounding boxes of t and r
source
object strip pairs
image tokens in the row (or column) strips through t and r
source
desc objects
description spans naming t and r (C2, C3)
source
desc locations
description spans stating the arrangement (C2, C3)
query
target + reference object
the query spans mentioning t and r , patched jointly
Appendix
Table 7: Token groups patched in all experiments in section 4 . t means target object and r means reference object.
Figure 9: Layer-wise activation patching continuous restore score across input conditions, token sources, corruption types and models in 3-object synthetic dataset.
Figure 10: Layer-wise activation patching clean-answer recovery rate across input conditions, token sources, corruption types and models in 2-object synthetic dataset.
Figure 11: Layer-wise activation patching continuous restore score across input conditions, token sources, corruption types and models in 2-object synthetic dataset.
Figure 12: Held-out recovery and shuffle controls for extracted location IDs.
Figure 13: Location-ID prototypes projected onto axes defined by endpoint-ID contrasts. Three-object plots additionally show the positions of middle-location prototypes relative to the endpoints.
Figure 14: Query-side location information supports relational comparison across models. For each model and input setting, we project the difference between the two query-object states onto the extracted query-side spatial axis and evaluate whether the projection sign predicts their relative direction. Extracted axes achieve high relation sign accuracy, whereas matched orthogonal control axes remain close to chance.
Figure 15: Intervention-strength curves in the selected layer bands. The y-axis uses the same direction-specific margin gap as in Figure 7 .
Data
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
2obj
LLaVA
VLM+image
99.8→100.0(+0.2)
99.6→100.0(+0.4)
99.4→100.0(+0.6)
2obj
LLaVA
VLM+text
58.6→60.4(+1.9)
42.0→46.3(+4.3)
40.8→45.4(+4.6)
2obj
LLaVA
LLM+text
55.3→59.8(+4.5)
36.5→38.1(+1.6)
31.3→34.4(+3.1)
2obj
InternVL
VLM+image
100.0→100.0(+0.0)
100.0→100.0(+0.0)
100.0→100.0(+0.0)
2obj
InternVL
VLM+text
90.1→93.3(+3.1)
85.2→88.9(+3.7)
81.1→85.5(+4.5)
2obj
InternVL
LLM+text
76.1→77.4(+1.4)
52.7→55.9(+3.1)
50.6→52.9(+2.3)
Appendix
Table 8: Activation steering on synthetic held-out data using the joint role direction r=(dTF+dRF)/2 estimated on synthetic training data.
Model
VLM+image
VLM+text
LLM+text
LLaVA-1.6
10/1.5
10/1.5
6/1.5
InternVL3.5
12/1.0
15/1.5
14/1.0
Pixtral
9/1.5
8/1.5
6/1.0
Appendix
Table 9: Intervention configurations estimated on synthetic 2-object data for What’sUp and COCO-spatial steering. Each cell represents selected layer number/ α value.
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
LLaVA-1.6
VLM+image
93.2→94.5(+1.4)
87.7→91.1(+3.4)
N/A
LLaVA-1.6
VLM+text
69.9→69.2(−0.7)
45.1→48.0(+2.8)
N/A
LLaVA-1.6
LLM+text
64.0→66.9(+3.0)
36.9→39.2(+2.3)
N/A
InternVL3.5
VLM+image
97.0→97.0(+0.0)
93.9→94.8(+0.9)
N/A
InternVL3.5
VLM+text
88.1→93.0(+4.9)
82.0→89.4(+7.4)
N/A
InternVL3.5
LLM+text
71.5→76.5(+5.0)
49.2→62.4(+13.2)
N/A
Appendix
Table 10: Activation steering results of the joint target/reference role direction on COCO-Spatial. Bold changes indicate improvements with 95% confidence intervals excluding zero. Location-swap consistency is unavailable because the evaluated subset does not contain paired original and location-swapped scenes.
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
LLaVA-1.6
VLM+image
93.6→95.8(+2.3)
90.9→93.9(+3.0)
N/A
LLaVA-1.6
VLM+text
57.6→58.0(+0.4)
49.8→49.6(−0.2)
N/A
LLaVA-1.6
LLM+text
44.5→48.5(+4.0)
13.1→19.1(+6.1)
N/A
InternVL3.5
VLM+image
98.1→98.5(+0.4)
97.7→97.7(+0.0)
N/A
InternVL3.5
VLM+text
81.6→88.1(+6.4)
71.8→83.1(+11.4)
N/A
InternVL3.5
LLM+text
72.2→77.7(+5.5)
48.3→66.1(+17.8)
N/A
Appendix
Table 11: Activation steering results of the joint target/reference role direction on the left/right subset of GQA-Spatial. Bold changes indicate improvements with 95% confidence intervals excluding zero.
Current Large Reasoning Models (LRMs) exhibit remarkable general capabilities but significantly underperform in spatial reasoning tasks. Existing approaches treat this gap as a knowledge deficit, relying on supervised fine-tuning (SFT) to ingest labeled spatial data from external vision sources or synthetic engines. In contrast, we argue that for many tasks, spatial reasoning capabilities are already present in pre-trained LRMs but require alignment through logical coherence under geometric 2D and 3D constraints. In this work, we propose a self-supervised reinforcement learning (RL) framework that targets the internal reasoning process without requiring ground-truth annotations. By formalizing the notion of consistency verifiers -- reward functions that check for geometric and semantic consistency under transformations -- we demonstrate that models can improve their spatial reasoning abilities. We use both image transformations, like flipping, and textual transformations, like swapping the order of objects in the question, and propose a new optimal transport-based RL strategy, OT-GRPO, which is a minimal-matching variant of group relative policy optimization tailored to pairwise verifiers. We show that this label-free consistency training approaches the accuracy of models trained with ground-truth supervision and achieves similar generalization across diverse tasks and data domains.
Theo Uscidda, Marta Tintore Gazulla, Maks Ovsjanikov +2
CREST, ENSAE, Institut Polytechnique de Paris · Google Zurich · Google DeepMind
Standard attribution heatmaps show where a vision-language model (VLM) focuses, but they do not reveal whether the recovered evidence is organized by the queried spatial relation or merely reflects image layout. To address this problem, we introduce CREG (Compass Relational Evidence Graph), a training-free diagnostic framework that converts token-level attribution into a reference-centered compass distribution and measures its directional alignment. CREG provides a shared directional readout across attribution methods and makes comparison with geometric controls explicit. Across three spatial-relation benchmarks, box-only geometry achieves Direction Alignment Error 28.4 to 34.4 degrees lower than the best current model-based attribution method on each dataset, leaving a substantial gap between attribution structure and simple target localization. To examine this gap, we apply a diagnostic battery including target intervention, reference-center randomization, and variance partition. Taken together, the results suggest that the directional structure recoverable from current attribution methods is limited and often mixed with image layout. We further find that higher task accuracy does not reliably coincide with better directional attribution: small-scale LoRA training and newer model generations can improve task accuracy while leaving Direction Alignment Error unchanged or worse. These findings characterize what current attribution methods reveal rather than the model's internal spatial representation. CREG provides a controlled protocol for testing whether improvements in spatial reasoning are accompanied by more directionally organized evidence.
Kaizhen Tan, Yang Feng, Heqing Du
Carnegie Mellon University Pittsburgh, PA, USA · Columbia University New York, NY, USA
Large Vision-Language Models (LVLMs) have made remarkable progress across visual perception tasks, yet spatial reasoning remains a persistent weakness, especially for questions that require reasoning over visual space. Recent spatial-reasoning methods incorporate generated grounding, where models predict bounding boxes, masks, or other localization outputs for task-relevant objects as part of their reasoning trace. However, these approaches typically optimize final-answer correctness alone, allowing correct answers to be rewarded even when the model does not reason from confidently localized task-relevant objects. We introduce SpatialCORE (Spatially COnfident REasoning), a post-training framework that turns the model's own confidence in generated grounding into a learning signal for spatial reasoning. Its central idea is to reinforce grounding that is both accurate and confident, encouraging the model to reason from confidently localized task-relevant objects. SpatialCORE realizes this through a self-regulating spatial reward that weights each predicted bounding box's matching quality by its coordinate-token confidence. An answer gate further ties grounding optimization to final-answer correctness. SpatialCORE achieves state-of-the-art results among open-source and specialized spatial reasoning models across diverse benchmarks, and transfers effectively in zero-shot settings to unseen data distributions. The source code is available at https://github.com/rafiibnsultan/SpatialCORE.
Rafi Ibn Sultan, Xiangyu Zhou, Md. Sajid Alam Chowdhury +4
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1