High instance-level accuracy can mask inconsistencies in spatial reasoning when objects exchange positions or their roles are reversed in the query. The internal representations supporting relative-position reasoning remain poorly understood. We investigate two complementary components of this process: tracking object locations in the input and representing their query roles. Across three VLMs with visual or textual inputs and their language-model backbones, activation patching reveals a staged progression from early-layer source representations through intermediate-layer query-object representations to late-layer answer states. Targeted interventions further establish causal links along this progression: manipulating source-side representations shifts location information at query-object mentions and ultimately alters relation predictions. Beyond object-location information, we also identify a stable query-side direction associated with the roles of the two objects in the comparison. Steering along directions estimated on synthetic scenes generalizes to natural-image benchmarks, improving accuracy and both forms of paired consistency in most settings without retraining. Our findings reveal complementary components of relational reasoning across visual and textual settings and show how targeted interventions can improve the consistency of models' behavior.
Figures & tables
Figure 1: Overview of the task and mechanistic analysis. (A) A scene, given as an image or an equivalent text description (the source ), and a query about the position of a target object relative to a reference object. (B) Object-location swaps and target/reference reversals, both reversing the correct answer. (C) Matched input conditions: VLM + image, VLM + text and LLM + text. (D) Hypothesized mechanism: source representations influence query-side location information, while target/reference roles condition its use in relation prediction. We test this account through activation patching and targeted interventions, and evaluate role-direction steering on natural images.
Figure 2: Benchmarking results for the three conditions: VLM+image, VLM+text and LLM+text. Numerical results are provided in Table 6 .
Figure 3: Layerwise activation patching clean-answer recovery rate across input conditions, token sources, corruption types and models in the 3-object synthetic dataset.
Setting
Patch group
LLaVA Δm / flip
InternVL Δm / flip
Pixtral Δm / flip
VLM + image
All visual
1.718 / 62.9%
21.660 / 69.1%
3.877 / 51.2%
VLM + image
Object+strip
0.881 / 34.0%
9.691 / 11.6%
1.477 / 5.1%
VLM + text
All objects
0.493 / 17.7%
7.912 / 24.3%
0.783 / 13.7%
LLM + text
All objects
0.416 / 20.0%
3.586 / 15.6%
0.701 / 10.3%
Table 1: Effects of source patching on query-side location information and answer preferences. Cells report all-layer mean corrupt-directed score shifts and corrupt-answer flip rates after patching corrupt-run source states into the clean run ( Δm→corr / flip %).
Figure 4: Location IDs in LLaVA-1.6-Mistral-7B on three-object scenes: (A) held-out location prediction with shuffled-label controls; (B–D) prototype visualization along axes defined by endpoint-ID contrasts; and (E) query-side relation-sign prediction with orthogonal-axis controls.
Figure 5: Layer-wise query-side location-ID steering measured by belief-swap rate.
Figure 6: Joint TF/RF role-contrast geometry across layers. High held-out mean cosine and norm ratio indicate that the joint object-level contrasts share a stable direction.
Figure 7: Layer-wise joint role-direction intervention at α=1 . Curves show the mean gap ΔMrole−ΔMorth . Negative values indicate stronger disruption along the joint role direction.
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
LLaVA-1.6
VLM+image
82.2→84.0(+1.8)
70.6→73.9(+3.3)
66.3→67.5(+1.3)
LLaVA-1.6
VLM+text
66.3→65.1(−1.2)
52.2→54.9(+2.7)
39.7→39.2(−0.5)
LLaVA-1.6
LLM+text
63.8→67.1(+3.3)
38.6→43.9(+5.3)
33.5→38.2(+4.7)
InternVL3.5
VLM+image
97.3→97.4(+0.1)
96.1→96.6(+0.5)
94.2→94.2(+0.0)
InternVL3.5
VLM+text
87.1→89.7(+2.6)
78.7→80.9(+2.2)
74.1→77.8(+3.7)
InternVL3.5
LLM+text
70.8→73.2(+2.4)
46.9→50.3(+3.3)
45.3→49.7(+4.4)
Table 2: Activation steering on What’sUp-A using the joint target/reference role direction estimated on synthetic data. Bold indicates improvements with 95% confidence intervals excluding zero.
Figure 8: Example data used in our experiments
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Split
Scenes
Questions
Horizontal
Vertical
Each relation
Reciprocal pairs
Synthetic 2-object
Train
1,512
3,024
1,512
1,512
756
1,512
Validation
324
648
324
324
162
324
Test
648
1,296
648
648
324
648
Total
2,484
4,968
2,484
2,484
1,242
2,484
Synthetic 3-object
Train
504
3,024
1,512
1,512
756
1,512
Validation
108
648
324
324
162
324
Appendix
Table 3: Statistics of the synthetic datasets used for mechanistic analyses. A scene denotes a unique rendered image, and a question denotes one directed target–reference query. “Each relation” reports the number of questions for each of { left , right , above , below }. Each two-object scene contains two reciprocal questions and therefore one target–reference pair. Each three-object scene contains all six directed questions among its three objects, corresponding to three reciprocal target–reference pairs.
VLM
Arch. type
LLM backbone
Layers
LLaVA-v1.5-7B
Projector-concat
Vicuna v1.5 (7B)
32
LLaVA-v1.5-13B
Projector-concat
Vicuna v1.5 (13B)
40
LLaVA-v1.6-Mistral-7B
Projector-concat
Mistral (7B)
32
LLaVA-v1.6-Vicuna-7B
Projector-concat
Vicuna v1.5 (7B)
32
LLaVA-v1.6-Vicuna-13B
Projector-concat
Vicuna v1.5 (13B)
40
Pixtral-12B
Projector-concat
Mistral (12B)
40
Appendix
Table 4: Architectures of benchmarked models. Layers denotes the number of transformer blocks in the language backbone.
Dataset
Condition
Original images
Scene families
Original queries
Reversal pairs
Swap anchors
Synthetic 2-object
VLM+image
128
128
256
128
256
VLM+text
128
128
512
256
512
LLM+text
128
128
512
256
512
Synthetic 3-object
VLM+image
128
128
256
128
256
VLM+text
128
128
512
256
512
LLM+text
128
128
512
256
512
Appendix
Table 5: Evaluation counts for Figure 2 and Table 6 . Counts are per evaluated model; text conditions pool the available description variants.
Model
Accuracy
Target–reference reversal consistency
Location-swap consistency
VLM+ image
VLM+ text
LLM+ text
VLM+ image
VLM+ text
LLM+ text
VLM+ image
VLM+ text
LLM+ text
(a) Synthetic
LLaVA-1.5-7B
92.58
56.84
50.00
85.94
48.44
0.00
83.59
47.66
0.00
LLaVA-1.5-13B
97.85
57.13
50.00
95.70
34.57
0.00
95.90
33.50
0.00
LLaVA-1.6-Mistral-7B
99.22
63.77
44.92
98.44
46.09
24.80
96.68
43.36
18.95
LLaVA-1.6-Vicuna-7B
96.29
56.05
50.00
92.58
50.39
0.00
92.97
48.44
0.00
Appendix
Table 6: Numerical benchmarking results for VLM+image, VLM+text, and the corresponding LLM+text backbone. Synthetic includes the two- and three-object 2D sets, and What’sUp includes subsets A and B.
Stage
Legend label
Tokens
source
all visual tokens
every image token (C1 only)
source
object bbox pairs
image tokens inside the bounding boxes of t and r
source
object strip pairs
image tokens in the row (or column) strips through t and r
source
desc objects
description spans naming t and r (C2, C3)
source
desc locations
description spans stating the arrangement (C2, C3)
query
target + reference object
the query spans mentioning t and r , patched jointly
Appendix
Table 7: Token groups patched in all experiments in section 4 . t means target object and r means reference object.
Figure 9: Layer-wise activation patching continuous restore score across input conditions, token sources, corruption types and models in 3-object synthetic dataset.
Figure 10: Layer-wise activation patching clean-answer recovery rate across input conditions, token sources, corruption types and models in 2-object synthetic dataset.
Figure 11: Layer-wise activation patching continuous restore score across input conditions, token sources, corruption types and models in 2-object synthetic dataset.
Figure 12: Held-out recovery and shuffle controls for extracted location IDs.
Figure 13: Location-ID prototypes projected onto axes defined by endpoint-ID contrasts. Three-object plots additionally show the positions of middle-location prototypes relative to the endpoints.
Figure 14: Query-side location information supports relational comparison across models. For each model and input setting, we project the difference between the two query-object states onto the extracted query-side spatial axis and evaluate whether the projection sign predicts their relative direction. Extracted axes achieve high relation sign accuracy, whereas matched orthogonal control axes remain close to chance.
Figure 15: Intervention-strength curves in the selected layer bands. The y-axis uses the same direction-specific margin gap as in Figure 7 .
Data
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
2obj
LLaVA
VLM+image
99.8→100.0(+0.2)
99.6→100.0(+0.4)
99.4→100.0(+0.6)
2obj
LLaVA
VLM+text
58.6→60.4(+1.9)
42.0→46.3(+4.3)
40.8→45.4(+4.6)
2obj
LLaVA
LLM+text
55.3→59.8(+4.5)
36.5→38.1(+1.6)
31.3→34.4(+3.1)
2obj
InternVL
VLM+image
100.0→100.0(+0.0)
100.0→100.0(+0.0)
100.0→100.0(+0.0)
2obj
InternVL
VLM+text
90.1→93.3(+3.1)
85.2→88.9(+3.7)
81.1→85.5(+4.5)
2obj
InternVL
LLM+text
76.1→77.4(+1.4)
52.7→55.9(+3.1)
50.6→52.9(+2.3)
Appendix
Table 8: Activation steering on synthetic held-out data using the joint role direction r=(dTF+dRF)/2 estimated on synthetic training data.
Model
VLM+image
VLM+text
LLM+text
LLaVA-1.6
10/1.5
10/1.5
6/1.5
InternVL3.5
12/1.0
15/1.5
14/1.0
Pixtral
9/1.5
8/1.5
6/1.0
Appendix
Table 9: Intervention configurations estimated on synthetic 2-object data for What’sUp and COCO-spatial steering. Each cell represents selected layer number/ α value.
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
LLaVA-1.6
VLM+image
93.2→94.5(+1.4)
87.7→91.1(+3.4)
N/A
LLaVA-1.6
VLM+text
69.9→69.2(−0.7)
45.1→48.0(+2.8)
N/A
LLaVA-1.6
LLM+text
64.0→66.9(+3.0)
36.9→39.2(+2.3)
N/A
InternVL3.5
VLM+image
97.0→97.0(+0.0)
93.9→94.8(+0.9)
N/A
InternVL3.5
VLM+text
88.1→93.0(+4.9)
82.0→89.4(+7.4)
N/A
InternVL3.5
LLM+text
71.5→76.5(+5.0)
49.2→62.4(+13.2)
N/A
Appendix
Table 10: Activation steering results of the joint target/reference role direction on COCO-Spatial. Bold changes indicate improvements with 95% confidence intervals excluding zero. Location-swap consistency is unavailable because the evaluated subset does not contain paired original and location-swapped scenes.
Model
Condition
Accuracy
Target-Reference Reversal Consistency
Location-Swap Consistency
LLaVA-1.6
VLM+image
93.6→95.8(+2.3)
90.9→93.9(+3.0)
N/A
LLaVA-1.6
VLM+text
57.6→58.0(+0.4)
49.8→49.6(−0.2)
N/A
LLaVA-1.6
LLM+text
44.5→48.5(+4.0)
13.1→19.1(+6.1)
N/A
InternVL3.5
VLM+image
98.1→98.5(+0.4)
97.7→97.7(+0.0)
N/A
InternVL3.5
VLM+text
81.6→88.1(+6.4)
71.8→83.1(+11.4)
N/A
InternVL3.5
LLM+text
72.2→77.7(+5.5)
48.3→66.1(+17.8)
N/A
Appendix
Table 11: Activation steering results of the joint target/reference role direction on the left/right subset of GQA-Spatial. Bold changes indicate improvements with 95% confidence intervals excluding zero.
Department of Computer Science, Wayne State University · Department of Radiation Oncology, Henry Ford Health · Department of Electrical and Computer Engineering, The Ohio State University +1