GaugeVLM: Structuring Spatial Supervision with Measured Geometric Interventions
Authors: Hongbo Wang, Zihan Lin, Wenkui Yang, Shiran Ge, Yuang Ai, Jie Cao, Huaibo Huang, Ran He
Organizations: NLPR & MAIS, Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · National University of Singapore · The Chinese University of Hong Kong
Vision-language models (VLMs) can contradict themselves across views of the same spatial relation and fail to respond when that relation changes. Addressing these failures requires supervision that captures error magnitude and geometric dependencies across observations, both of which remain implicit in training on individual answers or ordinal preferences. Therefore, we introduce GaugeVLM, which makes this structure explicit through controlled object and camera interventions in explicit 3D scenes, producing linked observations with measured differences between spatial relations and shared truths across views. To translate this structure into learning signals, its core objective, GaugeDPO, converts measured errors into preference margins, directly supervises correct canonical rankings across views, and links intervention-induced answer-odds contrasts to measured relation changes with view-specific scales. Our analysis bounds canonical prediction error and establishes that the cross-view and intervention constraints can be jointly satisfied. Empirically, GaugeVLM improves all 10 established spatial metrics over supervised fine-tuning across three VLM backbones, with the main 7B model gaining 15.0 and 18.9 percentage points on MSMU distance and QSpatial+, respectively. These gains also extend to autonomous driving and embodied reasoning, demonstrating the robust generalization across domains.
Figures & tables
Figure 1
Figure 1: GaugeDPO turns measured 3D interventions into supervision. Left, controlled perturbations define the margin m(r⋆,r−) . Top, it shifts response- and image-side pair losses. Bottom, direct supervision and intervention profiles enforce correct, view-consistent predictions.
Figure 2: One construction across training and evaluation. Left, Gauge-50K by pair family, with geometric displacement annotations for measured pairs. Middle, Constancy-Bench groups held-out scenes by the clock change between two images and reports how often the answer flips. Right, an intervention moves one object and re-renders the scene from up to four views.
MSMU
QSpat
SURDS
SpatialRGPT
3DSR
BLINK
Constancy-Bench
Model
#P
dist.
width
height
δ2
dist.
depth
Quan.
Qual.
acc.
acc.
Rank ρ↑
VisC. % ↑
PairAcc. % ↑
Frontier models
Claude Opus 4.8
–
2.5
1.1
2.2
76.2
11.7
14.7
41.5
55.8
66.4
70.1
.127
52.7
62.0
GPT-4o
–
2.5
0.0
5.5
47.5
2.7
0.0
41.7
74.3
55.3
56.2
.136
32.4
45.4
Gemini-2.5-Pro
–
7.5
6.7
18.7
45.5
29.5
34.3
38.1
80.9
54.2
54.1
.227
37.8
64.4
Kimi-K2.6
1T
15.0
9.0
18.7
64.4
48.5
53.1
42.0
89.3
62.8
72.8
.200
46.0
68.8
Table 1: Spatial understanding results. § marks models evaluated without their native additional visual cues for fair comparison. † marks constant outputs.
Figure 3: Predicted versus true values across all MSMU metric tasks. GaugeVLM achieves the best calibration even outperforming the Claude-Fable-5. Missing or unparseable numerical outputs are not plotted, so point counts vary across models. “Within ±25% ” denotes ∣d^−d⋆∣/d⋆≤0.25 .
Model
#P
MME P
MME R
POPE
MMStar
AI2D
SEED
SQA
HallB
MMMU
Average ‡
Qwen2.5-VL
7B
1606
622
83.7
58.1
78.8
73.0
72.7
63.3
42.8
67.5
GaugeVLM
7B
1640
629
85.9
60.3
80.1
75.0
78.9
64.5
44.4
69.9
Δvs base
+34
+7
+2.2
+2.2
+1.3
+2.0
+6.2
+1.2
+1.6
+2.4
GLM-4.1V
9B
1593
555
84.6
28.4
57.1
73.1
28.7
49.8
36.6
51.2
GaugeVLM
9B
1595
560
88.9
32.1
64.3
74.0
42.7
52.2
40.7
56.4
Δvs base
+2
+5
+4.3
+3.7
+7.2
+0.9
+14.0
+2.4
+4.1
+5.2
Table 2: General Vision-Language Understanding Ability Results. Average ‡ denotes the mean of the seven percentage metrics, excluding MME.
Figure 4: Measured margins induce graded policy separation. On held-out scenes, GaugeVLM scales its reference-relative gap with geometric error (a) and maintains positive slack (b). SFT is the reference ( hSFT=0 ); bands show 95% paired scene-bootstrap CIs.
Settings
Margin m
MSMU [-.3pt]dist.
QSpat [-1.3pt] δ2
SRGPT [-.3pt]Quan
3DSR [1.3pt]acc.
BLINK [1.3pt]acc.
Rank [-.3pt] ρ↑
PairAcc. [-.3pt]% ↑
(a)
SFT Initialization
none
47.5
45.5
33.5
48.2
50.5
.109
54.6
(b)
+ Zero-offset pair loss
m≡0
50.0
49.5
35.4
50.3
51.9
.119
56.6
(c)
+ Constant margin
m≡mˉ
52.5
52.5
37.2
51.6
53.1
.146
58.0
(d)
+ Shuffled margin
mσ(i)
52.5
51.5
36.6
51.8
52.6
.132
57.6
(e)
+ Reward margin
m^ϕ (estimated)
57.5
56.4
39.7
53.7
52.9
.142
61.0
(f)
+ GaugeDPO
measured (Eq. 4 )
62.5
64.4
43.1
56.9
56.7
.279
66.3
Table 3: Ablation of preference margin design and SFT initialization. Zero-, constant-, shuffled- and measured-margin rows retain the same direct/profile supervision; only the pair offset changes.
Table 9Table 10
Agreement
Accuracy
Invalid
Model
Total
Correct
Wrong
Group
Item
Invalid
SFT Init
60.0
40.0
20.0
50.0
65.0
5.0
λdir,λint=0
68.0
50.5
17.5
60.5
71.5
4.0
λdir=0
69.5
54.5
15.0
62.5
73.5
3.5
λint=0
72.5
60.0
12.5
67.5
77.0
3.5
GaugeVLM
74.5
63.0
11.5
69.5
79.0
3.0
Table 8: Tolerance-based cross-view audit. Ablations set the indicated weights to zero and retain the others. Agreement and correctness use separate tolerances. Invalid groups remain in all denominators. Values are percentages.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Family
Release pairs
Training-pool pairs
Metric distance
17800
10600
Cross-view
12900
7600
Graded errors
9393
5600
Clock direction
7219
4300
Vertical relation
3124
1900
Total
50436
30000
Appendix
Table 9: Corpus allocation by task family for the release and the preference-training pool.
Grade
Pairs
Min
p10
Median
p90
Max
Zero
minor
2701
0.1646
0.1718
0.1992
0.2234
0.2299
0
moderate
4021
0.4324
0.4444
0.4883
0.5259
0.5347
0
fatal
2671
0.8413
0.8615
0.9298
0.9872
1.0000
0
Appendix
Table 10: Serialized margin distribution by error grade. Distances use two decimal places; Zero counts labels with zero margin.
Check
Counting unit
Audited
Pass / dup.
Rate (%)
Wilson 95%
Template parseability
preference pair
1000
997
99.7
[99.1, 99.9]
Chosen truth matches scene state
preference pair
1000
995
99.5
[98.8, 99.8]
Anchor frame identifiable
evaluation group
200
187
93.5
[89.2, 96.2]
Scale cue sufficient for metric answer
evaluation group
200
178
89.0
[83.9, 92.6]
Complete fixed-configuration groups
evaluation group
200
194
97.0
[93.6, 98.6]
Train/test duplicate scene IDs
scene–configuration
200
0
0.0
[0.0, 1.9]
Appendix
Table 11: Data-quality and observability checks. Rows 1–2 use 1,000 candidate preference pairs; rows 3–5 use the 200 camera groups from Tab. 8 . The final two rows count train/test duplicates among the inspected units. Wilson 95% intervals refer to the pass rate in rows 1–5 and the duplicate rate in rows 6–7. Tab. 12 describes the handling of camera groups requiring fallbacks or marked incomplete.
Outcome
Check
Audited
Pass
Fallback
Invalid
Excl.
Fallback rule
Anchor frame
200
187
13
0
0
dominant wall normal
Scale cue
200
178
22
0
0
stated reference extent
Group completeness
200
194
0
6
0
none; scored invalid
Appendix
Table 12: Handling of reference-frame, scale, and completeness checks for the 200 camera groups. Each row partitions the full set into the four listed outcomes. Fallback groups are scored using the specified convention; incomplete groups remain in the denominator and are scored as invalid. No groups are excluded (Excl.).
Blocks
Min/median/max
Split
Role
Distinct
Covered
K
VB
Scenes
Strata
Train
GaugeDPO updates
1600
1447
2/3/4
2/2/2
800
16
Val
coefficient selection
200
200
2/3/4
2/2/2
100
16
Test
familiar
200
200
2/3/4
2/2/3
100
16
Test
unseen magnitudes
200
200
2/3/4
2/2/3
100
8
Test
unseen combinations
200
200
2/3/4
2/2/3
100
4
Appendix
Table 13: Intervention-block statistics. Training draws 3,752 blocks with replacement over 938 updates, visiting 1,447 of the 1,600 training blocks. The five splits use disjoint scene sets. Strata correspond to the magnitude–camera combinations in Tab. 14 .
Factor
Level
λ
drel
Split role
Magnitude
M1
1.25
0.1386
train, familiar test
Magnitude
M2
1.50
0.2519
train, familiar test
Magnitude
M3
2.20
0.4899
train, familiar test
Magnitude
M4
3.60
0.7959
train, familiar test
Magnitude
M5
4.50
0.9345
train, familiar test
Magnitude
H1
1.80
0.3652
held out; NN 0.1133
Appendix
Table 14: Distance factors and camera configurations. Distance levels are reported as the normalized component drel=min(∣logλ∣,κ)/κ , where κ=log5 and λ is the centroid-distance ratio. The joint margin mk uses Eq. 4 with the default α=0.5 ; for a distance-only change, mk=(1−α)drel=0.5drel . NN denotes the distance to the nearest training level in these normalized distance units. Camera bands specify azimuth separation. Held-out combinations retain both constituent factors in training.
Benchmark
Domain
GaugeVLM main metric
Δvs SFT
SPAR-Bench
indoor room (ScanNet)
overall 34.0
+8.7
SPAR-Bench (mv)
robotic multi-view
dist_oo 38.3
+16.5
QSpatial +
indoor/mixed
64.4
+18.9
Omni3DBench
mixed
MC 60.5
+5.8
SPBench-SI
indoor desktop
size 49.8
+5.4
Appendix
Table 15: Additional metric and spatial benchmark results. Task definitions differ across benchmarks; driving results appear in the main table.
Task (bench)
Domain
sub-metric
base
SFT
GaugeVLM
Δvs SFT
SPAR (mv)
embodied
dist_oc (obj-cam)
15.7
27.1
41.2
+14.1
SPAR (mv)
embodied
dist_oo (obj-obj)
19.8
21.8
38.3
+16.5
RoboSpatial
embodied
configuration
69.1
74.8
78.1
+3.3
RefSpatial
embodied
overall
13.0
16.0
17.0
+1.0
Appendix
Table 16: Additional embodied benchmark results.
Condition
Intervention PairAcc
Accurate agreement
Default options and state order
66.3
63.0
Permuted answer options
65.9
62.5
Reversed state presentation
65.4
62.5
Image masked
25.9
18.5
Text only
26.8
19.5
Appendix
Table 17: Input and presentation controls. Intervention PairAcc. and accurate agreement use separate evaluation sets.
Construction
MSMU
QSpat
Acc. agree.
Text-only pref. acc.
One-sided, inconsistent rationale
60.0
62.4
54.0
88.0
Reciprocal/sign-balanced only
62.5
63.4
56.5
79.0
Coherent rationale only
60.0
63.4
59.0
66.5
Balanced and coherent
62.5
64.4
63.0
54.5
Appendix
Table 18: Construction controls. Text-only preference accuracy measures discrimination of chosen/rejected pairs, not spatial answer accuracy.
Objective
MSMU distance
QSpatial +
PairAcc
GroupAcc tol
BlockAcc
Zero-offset pair loss
49.5±2.1
49.1±1.7
56.2±1.5
60.2±1.4
55.6±1.5
Constant pair offset
52.0±1.1
52.1±1.5
57.8±1.3
62.0±1.3
57.3±1.4
Shuffled pair offset
52.0±2.1
51.3±1.8
57.4±1.6
61.5±1.5
56.8±1.6
Pair + direct
57.5±1.8
61.0±2.8
62.0±1.2
67.2±1.2
59.7±1.3
+ Single-gap profile
59.0±1.4
60.8±1.1
63.3±1.1
68.0±1.1
63.4±1.2
+ EqSim-style symmetry
59.5±1.1
61.6±1.3
63.9±1.2
68.3±1.2
64.7±1.2
Appendix
Table 19: Preference-stage robustness. Mean ± sample SD computed from five accuracy scores per cell. Accuracy values are percentages.
Control
Metric
Mean Δ
95% CI seeds only
95% CI seeds + scenes
Positive seeds / 5
Primary comparisons
Shuffled pair offset
QSpatial +
+12.7
[10.7,14.7]
[8.8,16.4]
5 / 5
Pair + direct
BlockAcc
+9.9
[8.3,11.5]
[6.9,12.8]
5 / 5
Secondary comparisons
Zero-offset pair loss
QSpatial +
+14.9
[12.9,16.8]
[11.1,18.5]
5 / 5
Constant pair offset
QSpatial +
+11.9
[9.9,13.8]
[8.3,15.4]
5 / 5
Appendix
Table 20: Paired gains over controls. Full GaugeDPO minus each control, in percentage points. Primary comparisons are specified before the repeated runs. Both confidence intervals describe the paired difference.
Backbone
Training
MSMU distance
QSpatial +
PairAcc
BlockAcc
Qwen2.5-VL-7B
SFT
47.0±2.1
45.3±1.9
54.2±1.8
45.2±2.0
Zero-offset pair loss
49.5±2.1
49.1±1.8
56.3±1.7
55.4±1.9
GaugeDPO
62.0±2.1
63.8±1.5
65.8±1.4
69.4±1.5
GLM-4.1V-9B
SFT
67.0±2.1
49.1±2.1
67.0±1.9
47.8±2.1
Zero-offset pair loss
68.5±2.2
52.5±1.9
67.8±1.7
56.9±1.9
GaugeDPO
72.0±2.1
58.0±1.7
69.5±1.5
71.2±1.6
Appendix
Table 21: Robustness of the complete training pipeline. Five independent SFT–preference repetitions per backbone. Entries are mean ± sample SD in percent. SFT checkpoints are shared by the two preference objectives within each repetition.
Model
Input mode
MSMU dist.
QSpatial +
SpatialReasoner
restricted
25.0
49.5
SpatialReasoner
native cue bundle
30.0
53.5
SpatialRGPT
restricted
25.0
55.4
SpatialRGPT
native cue bundle
32.5
60.4
GaugeVLM
standard input
62.5
64.4
Appendix
Table 22: Restricted-input and native-input comparison. Each native configuration retains the model’s specified cue bundle.
Estimator
Inputs
Margin MAE
Rank ρ
Fit GPU-h
Score GPU-h
Scalar regressor
image + two relations
.087
.72
1.2
.3
Text-only regressor
two serialized relations
.104
.65
.5
.1
Direct geometric margin
scene-state labels
0.000
1.00
0.0
.02
Appendix
Table 23: Reward-estimator comparison against serialized geometric targets. The direct-margin row gives algebraic agreement with the target formula.
Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized 3D visual encoders is often inflexible and cumbersome. In this paper, we argue that genuine spatial understanding should emerge from learning fundamental geometric priors, not only from high-level VQA supervision. We propose GASP (Geometric-Aware Spatial Priors), a framework that injects these priors directly into the LLM's transformer layers. GASP employs a small correspondence head, applied as a deep supervision signal across all layers, and is trained with a dual objective leveraging ground-truth geometry from large-scale video scenes: a contrastive loss on ground-truth point correspondences enforces 2D view-invariance, while a depth consistency supervision resolves 3D geometric ambiguities. Our analysis first provides a diagnostic showing that standard VLMs' internal correspondence matching accuracy is very low (often below 5%). We then demonstrate that our training substantially improves this behavior, boosting peak layer-wise correspondence to over 70% and maintaining over 85% temporal robustness while baselines remain below 5%. These internal improvements translate to significant gains on downstream spatial benchmarks including +18.2% on All-Angles Bench and +29.0% on VSI-Bench, all without training on any 3D VQA data. Our findings indicate that learning from fundamental geometric priors is a promising and generalizable pathway towards VLMs with more reliable 3D spatial reasoning.
Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between. One cause of this failure arises before language reasoning begins: the visual pathway may compress or discard critical 3D structural cues during feature extraction, so the language model receives image representations that are already insufficient for reliable spatial judgment. We introduce GeoWorld-VLM, a VLM-side distillation framework that transfers geometric structure from frozen camera-conditioned video world models into VLMs. GeoWorld-VLM fine-tunes only the image encoder and multimodal projector, aligning post-projector image features with intermediate world-model representations while leaving the main backbone frozen. Given images, a prompt, and a sampled camera trajectory, the world-model teacher converts static visual input into a synthetic multi-view spatial signal. Training combines spatial answer supervision, teacher-student feature alignment, and a preservation anchor to the original VLM. Since the language model remains frozen, GeoWorld-VLM preserves the original model's linguistic capabilities while attributing spatial improvements to the enhanced visual pathway. To evaluate the effectiveness and generality of the proposed method, we apply GeoWorld-VLM to two distinct VLM architectures and observe consistent improvements across both backbones. GeoWorld-VLM improves performance by approximately 4 percent on both the What'sUp and VSR benchmarks, suggesting that world-model-guided visual alignment generalizes across model structures and spatial reasoning datasets.
Renjie Gu, Kaichen Zhou, Yan Luo +1
Harvard AI and Robotics Lab · Kempner Institute for the Study of Natural and Artificial Intelligence · Harvard University
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \href{https://github.com/xiaomi-research/spatio-lm}{\faGithub~spatio-lm}.
Jing Wu, Jianhua Wu, Jiayi Guan +5
Xiaomi EV, Beijing, China · College of Automotive and Energy Engineering, Tongji University, Shanghai, China · Independent Researcher