Geometric Encoding for Spatial Reasoning in Vision-Language Models
Authors: Antonio Jun, Haoshui Yu, Zhengyi Lu, Huirong Fu, Yao Qiang
Organizations: School of Arts and Sciences New York University New York City, United States · Dept. of Engineering and Computer Science Oakland University Rochester, United States
Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs' prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.
Figures & tables
Fig. 1: Geometric Code improves spatial reasoning from videos. Left: Standard VLMs reason directly from sampled video frames, whereas our framework additionally provides an explicit spatial code derived from geometric analysis of the video frames. Right: Average accuracy on VSI-Bench for four open-source VLMs. Augmenting the input with spatial code improves performance over the corresponding frames-only baselines.
Fig. 2: Overview of the Geometric Code. Monocular RGB frames are processed by the perception stage, where SAM-3 produces per-object segmentation masks and DA-3 estimates per-frame metric depth. A geometric engine back-projects these outputs into a spatial code, a JSON summary of per-object 3D positions and precomputed spatial evidence. Performance is evaluated under three VLM input conditions: frames only, spatial code only, and both.
VSI-Bench
Methods
Size
Avg.
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appear. Order
Human Level
–
79.2
94.3
47.0
60.4
45.9
94.7
95.8
95.8
100
Spatial-centric MLLMs
SpatialLadder [ 20 ]
3B
44.8
62.1
35.3
61.9
41.4
45.6
46.4
27.3
38.5
Spatial-MLLM [ 21 ]
4B
46.3
66.6
38.0
63.6
35.4
40.4
48.2
32.9
44.3
SpaceR [ 22 ]
7B
41.5
44.5
24.7
53.5
37.3
41.9
46.1
29.3
54.8
TABLE I: Video spatial reasoning results on VSI-Bench [ 7 ] . Comparison of four open-source MLLMs under the frames-only baseline, spatial code only, and combined settings. The upper blocks report results from prior work as compiled by Thinking-with-Spatial-Code [ 11 ] . Comparability is approximate due to evaluation-setup differences. Rows under Geometric Code are our own evaluations, with “Frames Only”, “Code Only”, and “Frames + Code” denoting our input configurations. Bold indicates the best result per column across all methods and underline the second best.
Model
Frames Only
Code Only
Frames+Code
Qwen3.5-2B
45.1
45.9 ( +0.8 )
54.3 ( +9.2 )
Qwen3.5-4B
57.4
56.5 ( −0.9 )
59.5 ( +2.1 )
InternVL3.5-2B
53.1
53.8 ( +0.7 )
53.4 ( +0.3 )
InternVL3.5-4B
56.9
56.8 ( −0.1 )
61.0 ( +4.1 )
Overall Avg.
53.1
53.3 ( +0.1 )
57.1 ( +3.9 )
TABLE II: Input-configuration ablation on VSI-Bench, reported as average accuracy. Values in parentheses ( Δ ) indicate changes relative to the Frames baseline for the corresponding model. bold denotes the best-performing configuration for each model. Overall Avg. is the unweighted mean across the four models.
Fig. 4: A case study sample built from an ARKitScenes scene with Qwen3.5-4B thinking. It showcases one question from each task: numeric-estimation Absolute Distance and multiple-choice Relative Direction. We provide reasoning for frames only in contrast to Geometric Code to showcase how Qwen3.5-4B reasons using both input configurations.
Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized 3D visual encoders is often inflexible and cumbersome. In this paper, we argue that genuine spatial understanding should emerge from learning fundamental geometric priors, not only from high-level VQA supervision. We propose GASP (Geometric-Aware Spatial Priors), a framework that injects these priors directly into the LLM's transformer layers. GASP employs a small correspondence head, applied as a deep supervision signal across all layers, and is trained with a dual objective leveraging ground-truth geometry from large-scale video scenes: a contrastive loss on ground-truth point correspondences enforces 2D view-invariance, while a depth consistency supervision resolves 3D geometric ambiguities. Our analysis first provides a diagnostic showing that standard VLMs' internal correspondence matching accuracy is very low (often below 5%). We then demonstrate that our training substantially improves this behavior, boosting peak layer-wise correspondence to over 70% and maintaining over 85% temporal robustness while baselines remain below 5%. These internal improvements translate to significant gains on downstream spatial benchmarks including +18.2% on All-Angles Bench and +29.0% on VSI-Bench, all without training on any 3D VQA data. Our findings indicate that learning from fundamental geometric priors is a promising and generalizable pathway towards VLMs with more reliable 3D spatial reasoning.
Vision-Language Models (VLMs) perform well on commonsense reasoning tasks but struggle with visual spatial reasoning. Most existing solutions introduce extra 3D prior inputs or external spatial encoders, which increase complexity and degrade the underlying VLMs' general-purpose capabilities after spatial fine-tuning. To this end, we propose a parameter-efficient \textit{\textbf{Spatio}-vision \textbf{L}anguage \textbf{M}odels (SpatioLM)}, that enhances spatial intelligence without extra 3D prior inputs or third-party spatial encoders. Concretely, we design a plug-and-play and non-invasive spatio-vision module that elicits the spatial knowledge inherent in VLMs. Furthermore, we innovatively leverage pseudo depth and camera information as supervision to guide the model in learning physically coherent representations. Extensive experiments show that SpatioLM achieves significant improvements in diverse tasks, including spatial perception and understanding while effectively limiting the degradation of general capabilities. Notably, the model achieves an impressive score of 71.6 on the VSI-Bench (the first model to surpass 70). In addition, it attains competitive performance when transferred to embodied manipulation tasks. Code is available at \faGithub~spatio-lm.
Jing Wu, Jianhua Wu, Jiayi Guan +5
Xiaomi EV, Beijing, China · College of Automotive and Energy Engineering, Tongji University, Shanghai, China · Independent Researcher
Modern Vision-Language Models (VLMs) achieve strong semantic recognition, yet remain brittle on elementary spatial relations such as left of, on, behind, and between. One cause of this failure arises before language reasoning begins: the visual pathway may compress or discard critical 3D structural cues during feature extraction, so the language model receives image representations that are already insufficient for reliable spatial judgment. We introduce GeoWorld-VLM, a VLM-side distillation framework that transfers geometric structure from frozen camera-conditioned video world models into VLMs. GeoWorld-VLM fine-tunes only the image encoder and multimodal projector, aligning post-projector image features with intermediate world-model representations while leaving the main backbone frozen. Given images, a prompt, and a sampled camera trajectory, the world-model teacher converts static visual input into a synthetic multi-view spatial signal. Training combines spatial answer supervision, teacher-student feature alignment, and a preservation anchor to the original VLM. Since the language model remains frozen, GeoWorld-VLM preserves the original model's linguistic capabilities while attributing spatial improvements to the enhanced visual pathway. To evaluate the effectiveness and generality of the proposed method, we apply GeoWorld-VLM to two distinct VLM architectures and observe consistent improvements across both backbones. GeoWorld-VLM improves performance by approximately 4 percent on both the What'sUp and VSR benchmarks, suggesting that world-model-guided visual alignment generalizes across model structures and spatial reasoning datasets.
Renjie Gu, Kaichen Zhou, Yan Luo +1
Harvard AI and Robotics Lab · Kempner Institute for the Study of Natural and Artificial Intelligence · Harvard University