cs.CVSep 24, 2026

Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models

Authors: Kai Glantz, Clemens Grange

Organizations: Technical University of Munich

Abstract

Vision-Language Models (VLMs) describe scenes well but reason poorly about metric 3D structure such as absolute distances, physical sizes, or egocentric directions. We present a modular, predictor agnostic, tool-augmented framework that equips a small VLM (Qwen3.5-4B) with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing. Each object is detected in the camera frame of its own best view, and the tools use that frame's pose to lift every detection into one shared world frame. Moving metric computation out of the model's weights and into explicit solvers yields large gains on three of four ReVSI-Bench tasks: with a strong monocular detector (WildDet3D), absolute distance rises from 0.46 to 0.74 Mean Relative Accuracy (MRA), relative distance from 39.1% to 67.4%, and relative direction from a below-chance 25.9% to 73.4%. Because any detector can be swapped in behind the tool interface, comparing real detectors against ground-truth boxes separates perception error from reasoning error: orchestration costs only 0.03 MRA. Object size is bounded by the detector: the tools are near-exact on groundtruth boxes (0.97) yet the best real detector barely beats the no-tool baseline (0.61 vs. 0.58), because size reads straight off a box extent monocular detectors get wrong. Without a predefined recipe, the model already sequences the tools correctly on its own, matching a scripted pipeline on three of four tasks.

Figures & tables

Explore similar work

CardsList
  1. Geometric Encoding for Spatial Reasoning in Vision-Language Models

    Sep 28, 2026Antonio Jun, Haoshui Yu, Zhengyi Lu +2Spatial Reasoning3D Layout Generation

  2. Unlocking Dense Metric Depth Estimation in VLMs

    May 15, 2026Hanxun Yu, Xuan Qu, Yuxin Wang +2Depth Estimation3D Representation

  3. SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models

    Aug 3, 2026Jing Wu, Jianhua Wu, Jiayi Guan +5Recent Vision-Language ModelsStable Spatial Understanding