Geometric Encoding for Spatial Reasoning in Vision-Language Models
Organizations: School of Arts and Sciences New York University New York City, United States · Dept. of Engineering and Computer Science Oakland University Rochester, United States
Abstract
Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs' prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.
Figures & tables
| VSI-Bench | ||||||||||
| Methods | Size | Avg. | Obj. Count | Abs. Dist. | Obj. Size | Room Size | Rel. Dist. | Rel. Dir. | Route Plan | Appear. Order |
| Human Level | – | 79.2 | 94.3 | 47.0 | 60.4 | 45.9 | 94.7 | 95.8 | 95.8 | 100 |
| Spatial-centric MLLMs | ||||||||||
| SpatialLadder [ 20 ] | 3B | 44.8 | 62.1 | 35.3 | 61.9 | 41.4 | 45.6 | 46.4 | 27.3 | 38.5 |
| Spatial-MLLM [ 21 ] | 4B | 46.3 | 66.6 | 38.0 | 63.6 | 35.4 | 40.4 | 48.2 | 32.9 | 44.3 |
| SpaceR [ 22 ] | 7B | 41.5 | 44.5 | 24.7 | 53.5 | 37.3 | 41.9 | 46.1 | 29.3 | 54.8 |
| Model | Frames Only | Code Only | Frames+Code |
| Qwen3.5-2B | 45.1 | 45.9 ( ) | 54.3 ( ) |
| Qwen3.5-4B | 57.4 | 56.5 ( ) | 59.5 ( ) |
| InternVL3.5-2B | 53.1 | 53.8 ( ) | 53.4 ( ) |
| InternVL3.5-4B | 56.9 | 56.8 ( ) | 61.0 ( ) |
| Overall Avg. | 53.1 | 53.3 ( ) | 57.1 ( ) |