cs.CVSep 28, 2026

Geometric Encoding for Spatial Reasoning in Vision-Language Models

Authors: Antonio Jun, Haoshui Yu, Zhengyi Lu, Huirong Fu, Yao Qiang

Organizations: School of Arts and Sciences New York University New York City, United States · Dept. of Engineering and Computer Science Oakland University Rochester, United States

Abstract

Vision-Language Models (VLMs) are far more reliable at recognizing what appears in a video than at reasoning about its spatial and temporal properties, such as metric distances, object dimensions, and consistent object identities across frames. We present Geometric Code, a perception-to-geometry pipeline that computes explicit spatial structure from video and supplies it to VLMs as context to augment reasoning. A perception layer segments and classifies objects and recovers depth, camera pose, and intrinsics from monocular RGB video. A deterministic geometric engine then back-projects, merges, and cleans these outputs into a spatial code, including per-object positions, dimensions, counts, inter-object distances, appearance order, and room geometry. The code is serialized into VLMs' prompts, either alongside the video or replacing it entirely. Specifically, there is no component trained or fine-tuned in our approach. On VSI-Bench, augmenting 2B and 4B open models with the spatial code improves average accuracy by +4.1 points over the frames-only baseline, with the largest gains on numeric estimation tasks such as absolute distance (+24.1 points). The results suggest that explicitly computed geometry, delivered through the language channel, recovers spatial competence that small VLMs cannot extract from pixels alone.

Figures & tables

Explore similar work

CardsList
  1. SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    Aug 3, 2026Jing Wu, Jianhua Wu, Jiayi Guan +5Recent Vision-Language ModelsStable Spatial Understanding

  2. GeoWorld-VLM: Geometry from World Models for Vision-Language Models

    May 15, 2026Renjie Gu, Kaichen Zhou, Yan Luo +1Spatial SupervisionVideo World Models