Seeing Is Not Measuring: Tool-Augmented Metric Spatial Reasoning for Vision-Language Models
Organizations: Technical University of Munich
Abstract
Vision-Language Models (VLMs) describe scenes well but reason poorly about metric 3D structure such as absolute distances, physical sizes, or egocentric directions. We present a modular, predictor agnostic, tool-augmented framework that equips a small VLM (Qwen3.5-4B) with geometric tools: 3D object detection, metric depth estimation, and deterministic solvers for distance, size and bearing. Each object is detected in the camera frame of its own best view, and the tools use that frame's pose to lift every detection into one shared world frame. Moving metric computation out of the model's weights and into explicit solvers yields large gains on three of four ReVSI-Bench tasks: with a strong monocular detector (WildDet3D), absolute distance rises from 0.46 to 0.74 Mean Relative Accuracy (MRA), relative distance from 39.1% to 67.4%, and relative direction from a below-chance 25.9% to 73.4%. Because any detector can be swapped in behind the tool interface, comparing real detectors against ground-truth boxes separates perception error from reasoning error: orchestration costs only 0.03 MRA. Object size is bounded by the detector: the tools are near-exact on groundtruth boxes (0.97) yet the best real detector barely beats the no-tool baseline (0.61 vs. 0.58), because size reads straight off a box extent monocular detectors get wrong. Without a predefined recipe, the model already sequences the tools correctly on its own, matching a scripted pipeline on three of four tasks.
Figures & tables
| No-tool baselines, absolute distance (MRA ) | |||
|---|---|---|---|
| Model | visual | blind | +intrinsics |
| Qwen3.5-0.8B | 0.00 | 0.00 | 0.02 |
| Qwen3.5-2B | 0.21 | 0.00 | 0.03 |
| Qwen3.5-4B | 0.46 | 0.29 | 0.17 |
| Method | MRA | Ceiling |
|---|---|---|
| Absolute distance (401 Q) | ||
| No tools (visual) | 0.46 | – |
| Tools, 2D + depth | 0.32 | 0.29 |
| Tools, 3D boxes (OVMono3D) | 0.26 | 0.26 |
| Tools, 3D boxes (Cube R-CNN) | 0.32 | 0.28 |
| Tools, 3D boxes (WildDet3D) | 0.74 | 0.74 |
| Method | Acc. | Parse | Ceiling |
|---|---|---|---|
| Relative distance (215 Q, chance ) | |||
| No tools (blind) | 30.2% | 100.0% | – |
| No tools (visual) | 39.1% | 100.0% | – |
| Tools, 3D boxes (WildDet3D) | 67.4% | 100.0% | 69.3% |
| Tools, 3D boxes (GT oracle) | 92.1% | 100.0% | 94.9% |
| Relative direction (290 Q, chance ) | |||
| MRA | Acc. | |||
|---|---|---|---|---|
| Method | Dist. | Size | R. dist | R. dir |
| No recipe (generic prompt) | 0.94 | 0.97 | 92.5 | 70.0 |
| + few-shot order examples | 0.89 | 1.00 | 80.0 | 75.0 |
| Scripted recipe ( Tabs. 2 and 3 ) | 0.94 | 0.97 | 92.1 | 80.0 |
| Backend | Objects found | Questions fully covered |
|---|---|---|
| GT Oracle | 100.0% (802/802) | 100.0% (401/401) |
| OVMono3D | 100.0% (802/802) | 100.0% (401/401) |
| Cube R-CNN | 91.8% (736/802) | 85.0% (341/401) |
| WildDet3D | 100.0% (802/802) | 100.0% (401/401) |
| Grounding DINO (2D) | 93.6% (751/802) | 87.8% (352/401) |
| Pipeline | task / answer_format |
|---|---|
| bbox_3d_abs_dist gdino_depthpro_abs_dist | distances between objects a single number in metres, e.g.: 1.5 |
| bbox_3d_size gdino_depthpro_size | the size (longest dimension) of objects a single number in centimetres, e.g.: 120 |
| bbox_3d_reldist | which of several candidate objects is closest or farthest to a reference object the letter (A, B, C, or D) |
| bbox_3d_reldir | the egocentric direction (left/right/back, or a front/back-left/right quadrant) of a target from a viewer the letter (A, B, C, or D) |
| autonomous_3d | spatial questions about objects (all four tasks mixed) the exact format the question asks for |
| Tool | Role (in/out) |
|---|---|
| detect_object_3d | name camera-space OBB (predictor-agnostic backend) |
| project_box_to_world | name world-space OBB (pose lift) |
| calculate_object_distance | two names surface distance (m) |
| calculate_object_size | name longest dim (cm) |
| relative_direction | three names mode 3-/4-way label |
| detect_object | name 2D bbox (Grounding DINO) |
| Task | Pipeline (Toolset) | Geometry Representation | Min. Steps | Median Steps |
|---|---|---|---|---|
| Abs. Distance | gdino_depthpro_abs_dist | 2D bbox + Depth (center-to-center) | 7 | 7 |
| Abs. Distance | bbox_3d_abs_dist | 3D OBB (surface-to-surface) | 5 | 5 |
| Object Size | gdino_depthpro_size | Apparent 2D projection | 3 | 3 |
| Object Size | bbox_3d_size | 3D OBB extent (rigid invariant) | 2 | 2 |
| Rel. Distance | bbox_3d_reldist | 3D OBB (nearest-instance surface) | 14 | 16 |
| Rel. Direction | bbox_3d_reldir | 3D OBB (egocentric bearing angle) | 7 | 8 |
| Subtype | N | GT Oracle | WildDet3D |
|---|---|---|---|
| Backward Easy (3-way) | 53 | 79.2% | 75.5% |
| Backward Hard (4-way) | 69 | 84.1% | 68.1% |
| Forward Easy (3-way) | 96 | 82.3% | 74.0% |
| Forward Hard (4-way) | 72 | 73.6% | 76.4% |
| Overall | 290 | 80.0% | 73.4% |
| Condition | MRA | Dist. | Size | Errors | Zero-tool |
|---|---|---|---|---|---|
| 1. Terse tool errors | 0.729 | 0.575 | 0.883 | 138 | 5/80 |
| 2. + actionable errors | 0.800 | 0.878 | 0.723 | 36 | 14/80 |
| 3. + few-shot examples | 0.933 | 0.938 | 0.928 | 13 | 0/80 |