Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Figures & tables
Figure 1: Panel (a) shows our baseline, which incorporates geometric priors through feature fusion and is trained on spatial-reasoning QA without QA-RP or CoT-VC. Panel (b) shows SpatialSpeak , which trains the VLM to conduct multi-view 3D reconstruction through QA-Native Reconstruction Pretraining (QA-RP) with complementary local geometry and global scene context, then to involve the learned geometry in explicit spatial reasoning through spatial Chain-of-Thought with Visual Compensation (CoT-VC), achieving leading results on ReVSI ( Zhang et al., 2026c ) . The top-right chart compares average ReVSI scores with SpatialStack-4B ( Zhang et al., 2026a ) , GeoThinker-8B ( Li et al., 2026a ) , VLM-3R-7B ( Fan et al., 2026 ) , Cambrian-S-7B ( Yang et al., 2026 ) , Omni-View-7B ( Hu et al., 2026a ) , VST-7B-SFT ( Yang et al., 2025c ) , and VG-LLM-8B ( Zheng et al., 2025a ) . SpatialSpeak-4B achieves 62.8, outperforming SpatialStack-4B (54.1) by 8.7 points.
Figure 2: Overview of the SpatialSpeak framework . In Training Stage I, QA-Native Reconstruction Pretraining (QA-RP) trains the VLM to conduct metric-scale multi-view reconstruction within the standard QA interface. Local point queries and global object-center queries jointly supervise fine-grained local geometry and global scene context, with both targets expressed in the first-frame camera coordinate system. In Training Stage II, spatial Chain-of-Thought with Visual Compensation (CoT-VC) trains the model to reason explicitly with the learned geometry, assess the reliability of geometry-derived answers, and refine answers using direct visual evidence when needed.
Figure 3: Qualitative example . The bottom-left panel shows the scene reconstruction predicted after Stage I (QA-RP), and the bottom-right panel shows the spatial CoT output after Stage II (CoT-VC). SpatialSpeak lists two blackboard instances with their estimated 3D centers and predicts a count of 2, matching the ground truth. The variant with neither QA-RP nor CoT-VC predicts 3.
Method
Avg.
SpatialSpeak (full)
62.8
w/o CoT-VC
55.9
w/o QA-RP
55.0
w/o both
52.4
Table 1: Training components . QA-RP and CoT-VC each improve spatial reasoning, with larger gains when both designs are used.
Method
Avg.
SpatialSpeak (full)
62.8
w/o CoT-VC
55.9
w/o QA-RP
55.0
w/o both
52.4
Table 1: Training components . QA-RP and CoT-VC each improve spatial reasoning, with larger gains when both designs are used.
Method
Avg.
SpatialSpeak (full)
62.8
w/o global queries
59.9
w/o local queries
59.3
w/o both
55.0
Table 2: QA-RP context . Ablating global object-center queries, local point queries, or both from multi-view reconstruction pretraining.
Method
Avg.
SpatialSpeak ( τ=0.1 )
59.6
SpatialSpeak ( τ=0.3 )
62.8
SpatialSpeak ( τ=0.5 )
59.5
w/o CoT-VC
55.9
Table 3: Reliability threshold . Varying τ for reliability labels in CoT-VC (Sec. 3.3 ). All three outperform the variant w/o CoT-VC.
Method
Avg.
SpatialSpeak (full)
62.8
w/o VC
58.5
w/o CoT-VC
55.9
Table 4: Spatial CoT . All model variants retain QA-RP. Removing VC preserves spatial CoT. Removing CoT-VC retains only direct answer supervision in Stage II.
Method
Avg.
SpatialSpeak (full)
62.8
w/o VC
58.5
w/o CoT-VC
55.9
Table 4: Spatial CoT . All model variants retain QA-RP. Removing VC preserves spatial CoT. Removing CoT-VC retains only direct answer supervision in Stage II.
Method
Acc. ↓
Comp. ↓
Acc. ∗ ↓
Comp. ∗ ↓
MapAnything
36.3
28.4
5.7
6.2
CUT3R
13.7
12.7
4.7
4.6
SpatialSpeak (Ours)
8.9
9.3
5.0
5.0
Table 5: Pointmap reconstruction on ScanNet . Acc. and Comp. are evaluated without alignment. Acc. ∗ and Comp. ∗ are evaluated after GT Sim(3) alignment. All errors are reported in cm. Lower values are better. Baselines are MapAnything ( Keetha et al., 2026 ) and CUT3R ( Wang et al., 2025b ) . Bold type marks the lowest error in each column.
Figure 4: Qualitative reconstruction comparison . Red dashed arrows mark corresponding distances in meters. SpatialSpeak estimates the distance as 1.24 m, close to the ground truth of 1.25 m.
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Method
Avg.
Numerical Answer
Multiple-Choice Answer
SpaceR-7B (SG-RLVR) ( Ouyang et al., 2025 )
30.5
30.7
34.5
52.0
18.6
22.8
34.5
20.2
Spatial-MLLM-4B-135k ( Wu et al., 2025 )
40.5
40.7
45.3
46.8
–
32.3
37.4
–
Spatial-MLLM-4B-820k ( Wu et al., 2025 )
40.9
41.5
40.0
53.1
–
30.7
39.2
–
VST-7B-SFT ( Yang et al., 2025c )
46.4
35.4
52.6
67.9
47.2
49.2
36.9
35.4
VG-LLM-8B ( Zheng et al., 2025a )
46.4
37.8
53.0
56.8
48.0
57.2
33.8
38.0
Table 6: Comparison with state-of-the-art methods on ReVSI . We evaluate SpatialSpeak and official checkpoints of SpatialStack, GeoThinker, VG-LLM, and Omni-View under the same 32-frame setting. Other baseline results are sourced from the ReVSI paper ( Zhang et al., 2026c ) .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Method
Avg.
Numerical Answer
Multiple-Choice Answer
Proprietary Models (API)
GPT-4o
34.0
46.2
5.3
43.8
38.2
37.0
41.3
31.5
28.5
Gemini-1.5-Flash
42.1
49.8
30.8
53.5
54.4
37.7
41.0
31.5
37.8
Gemini-1.5-Pro
45.4
56.2
30.9
64.1
43.6
51.3
46.3
36.0
34.6
Open-source Models
Table 7: Comparison with state-of-the-art methods on VSI-Bench (normal training setting).
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Method
Avg.
Numerical Answer
Multiple-Choice Answer
Qwen3-VL-8B ( Bai et al., 2025a )
59.8
67.5
52.6
76.2
62.3
60.6
52.5
32.5
73.8
VLM-3R-7B ( Fan et al., 2026 )
60.9
70.2
49.4
69.2
67.1
65.4
80.5
45.4
40.1
Map2Thought-7B ( Gao et al., 2026 )
61.0
70.8
55.0
70.1
69.4
56.9
69.8
38.1
57.4
VST-7B ( Yang et al., 2025c )
61.2
-
-
-
-
-
-
-
-
VG-LLM-8B ( Zheng et al., 2025a )
62.2
71.4
56.8
69.0
69.1
67.9
83.2
47.4
32.5
Table 8: Comparison with state-of-the-art methods on VSI-Bench (scaled training setting).
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Avg.
Low
Depth-OC
Depth-OC-MV
Depth-OO
Depth-OO-MV
Dist-OC
Dist-OC-MV
Dist-OO
Dist-OO-MV
Medium
PosMatch
CamMotion
ViewChgI
High
DistI-OO
DistI-OO-MV
ObjRel-OC-MV
ObjRel-OO
ObjRel-OO-MV
SpImag-OC
SpImag-OC-MV
SpImag-OO
SpImag-OO-MV
InternVL2-2B
28.1
21.7
18.1
24.8
23.2
21.0
19.5
20.0
26.8
20.6
22.8
39.7
23.0
5.8
35.4
51.2
56.0
46.0
31.6
23.8
36.0
34.3
17.6
22.4
InternVL2-4B
32.0
28.9
23.9
27.2
20.0
18.1
42.6
40.2
31.3
28.2
29.2
49.9
21.0
16.6
35.7
56.8
55.4
40.3
36.8
25.2
28.8
32.3
21.2
24.7
InternVL2.5-2B
30.1
25.8
39.7
39.7
12.1
15.0
30.9
29.6
20.2
19.0
22.9
37.9
24.3
6.6
36.4
51.5
56.9
50.3
33.8
24.1
27.2
35.2
26.5
22.4
InternVL2.5-4B
30.6
25.7
29.1
33.0
21.8
16.8
20.8
26.9
28.1
28.8
29.8
47.1
33.3
8.9
35.2
54.1
58.9
35.5
29.7
34.6
24.7
31.4
19.2
28.3
InternVL2.5-8B
36.3
29.5
25.8
29.3
23.8
18.8
46.8
42.7
22.6
25.9
31.9
61.3
28.0
6.3
43.8
59.7
56.9
51.8
44.2
41.6
36.6
41.6
22.5
39.5
LLaVA-OV-0.5B
29.5
30.1
49.2
42.7
18.0
14.9
31.5
25.7
29.0
30.1
15.9
24.4
21.8
1.5
33.4
50.9
50.0
32.0
27.8
26.0
30.9
34.0
24.5
24.7
Appendix
Table 9: Comparison with state-of-the-art models on SPAR-Bench ( Zhang et al., 2025 ) . Baselines include InternVL2 ( Chen et al., 2024b ) , InternVL2.5 ( Chen et al., 2024a ) , LLaVA-OV ( Li et al., 2025a ) , Qwen2-VL ( Wang et al., 2024a ) , Qwen2.5-VL ( Bai et al., 2025b ) , LLaVA-v1.5 ( Liu et al., 2024a ) , LLaVA-v1.6 ( Liu et al., 2024b ) , Spatial-MLLM ( Wu et al., 2025 ) , VLM-3R ( Fan et al., 2026 ) , UniUGG-3B ( Xu et al., 2025 ) , G 2 VLM-SR ( Hu et al., 2026b ) , GeoThinker ( Li et al., 2026a ) , SenseNova-SI ( Cai et al., 2026b ) and SpatialStack ( Zhang et al., 2026a ) . Baseline results are primarily sourced from the supplementary material of G 2 VLM ( Hu et al., 2026b ) .
Figure 5: Qualitative example . The bottom-left panel shows the Stage I (QA-RP) reconstruction, and the bottom-right panel shows the Stage II (CoT-VC) reasoning output. SpatialSpeak estimates the 3D centers and dimensions of the trash bin and toilet, then approximates their closest-point distance as 3.6 m, compared with 2.7 m from the variant with neither QA-RP nor CoT-VC and a ground truth of 3.5 m.
Figure 6: Qualitative example . The bottom-left panel shows the Stage I (QA-RP) reconstruction, and the bottom-right panel shows the Stage II (CoT-VC) response. SpatialSpeak enumerates four chair instances with distinct estimated 3D centers and predicts a count of 4, matching the ground truth, whereas the variant with neither QA-RP nor CoT-VC predicts 6.
Figure 7: Extended qualitative reconstruction comparisons . Rows show MapAnything ( Keetha et al., 2026 ) , CUT3R ( Wang et al., 2025b ) , SpatialSpeak after Stage I, and ground truth from top to bottom. The left column provides enlarged visualizations of the scene in Fig. 4 , while the right column shows an additional scene. Red dashed arrows mark corresponding distances in meters. SpatialSpeak’s marked distances (1.24/2.52 m, left/right) are closer to the ground-truth values (1.25/2.54 m) than those of CUT3R (1.38/2.82 m) and MapAnything (1.63/2.90 m).
Hyperparameter
Value
accelerator model
H800 GPUs
accelerator count
8
global batch size
32
tune_mm_llm
True
tune_mm_vision
False
tune_mm_mlp
False
Appendix
Table 10: Training hyperparameters in Training Stage I .
Hyperparameter
Value
accelerator model
H800 GPUs
accelerator count
8
global batch size
64
tune_mm_llm
True
tune_mm_vision
False
tune_mm_mlp
False
Appendix
Table 11: Training hyperparameters in Training Stage II .
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial queries call for fundamentally different strategies: some are best addressed through purely linguistic, step-by-step deduction, while others require explicit 3D grounding before quantitative inference. We present Dual-Path Spatial Reasoning via Reinforcement Learning for Spatial VLMs (SR-REAL), a unified framework that equips a spatial VLM with two complementary reasoning paths: Language-Only Reasoning (LOR), which performs step-by-step linguistic deduction, and Detect-Then-Reason (DTR), which detects 3D geometric cues (e.g., centers or bounding boxes) via region tokens before explicit geometric inference. SR-REAL begins with a cold-start supervised fine-tuning stage that constructs LOR and DTR chain-of-thought supervision and exposes a region-to-3D interface, followed by RL that optimizes the policy model with accuracy and format rewards; for DTR, a discrete center-based detection reward further refines geometric alignment. Across diverse spatial benchmarks, SR-REAL significantly outperforms spatial VLM baselines: (i) a single RL-trained model supports both reasoning paths, with DTR excelling in region-aware tasks through precise 3D localization and LOR enhancing general spatial reasoning; (ii) jointly training both paths fosters mutual reinforcement; (iii) high-quality, blended cold-start data is crucial for stable RL optimization; and (iv) the model generalizes across datasets and domains without per-task tuning, demonstrating positive transfer between LOR and DTR.
Yatai Ji, An-Chieh Cheng, Yang Fu +13
The University of Hong Kong · NVIDIA · University of California, San Diego
Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: https://yuqunw.github.io/Render2Reason/.
Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized 3D visual encoders is often inflexible and cumbersome. In this paper, we argue that genuine spatial understanding should emerge from learning fundamental geometric priors, not only from high-level VQA supervision. We propose GASP (Geometric-Aware Spatial Priors), a framework that injects these priors directly into the LLM's transformer layers. GASP employs a small correspondence head, applied as a deep supervision signal across all layers, and is trained with a dual objective leveraging ground-truth geometry from large-scale video scenes: a contrastive loss on ground-truth point correspondences enforces 2D view-invariance, while a depth consistency supervision resolves 3D geometric ambiguities. Our analysis first provides a diagnostic showing that standard VLMs' internal correspondence matching accuracy is very low (often below 5%). We then demonstrate that our training substantially improves this behavior, boosting peak layer-wise correspondence to over 70% and maintaining over 85% temporal robustness while baselines remain below 5%. These internal improvements translate to significant gains on downstream spatial benchmarks including +18.2% on All-Angles Bench and +29.0% on VSI-Bench, all without training on any 3D VQA data. Our findings indicate that learning from fundamental geometric priors is a promising and generalizable pathway towards VLMs with more reliable 3D spatial reasoning.