SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning
Organizations: Hong Kong University of Science and Technology · Harbin Institute of Technology · Huawei Noah’s Ark Lab
Abstract
Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Figures & tables
| Method | Avg. |
| SpatialSpeak (full) | 62.8 |
| w/o CoT-VC | 55.9 |
| w/o QA-RP | 55.0 |
| w/o both | 52.4 |
| Method | Avg. |
| SpatialSpeak (full) | 62.8 |
| w/o CoT-VC | 55.9 |
| w/o QA-RP | 55.0 |
| w/o both | 52.4 |
| Method | Avg. |
| SpatialSpeak (full) | 62.8 |
| w/o global queries | 59.9 |
| w/o local queries | 59.3 |
| w/o both | 55.0 |
| Method | Avg. |
| SpatialSpeak ( ) | 59.6 |
| SpatialSpeak ( ) | 62.8 |
| SpatialSpeak ( ) | 59.5 |
| w/o CoT-VC | 55.9 |
| Method | Avg. |
| SpatialSpeak (full) | 62.8 |
| w/o VC | 58.5 |
| w/o CoT-VC | 55.9 |
| Method | Avg. |
| SpatialSpeak (full) | 62.8 |
| w/o VC | 58.5 |
| w/o CoT-VC | 55.9 |
| Method | Acc. | Comp. | Acc. ∗ | Comp. ∗ |
| MapAnything | 36.3 | 28.4 | 5.7 | 6.2 |
| CUT3R | 13.7 | 12.7 | 4.7 | 4.6 |
| SpatialSpeak (Ours) | 8.9 | 9.3 | 5.0 | 5.0 |
| Obj. Count | Abs. Dist. | Obj. Size | Room Size | Rel. Dist. | Rel. Dir. | Route Plan | ||
| Method | Avg. | Numerical Answer | Multiple-Choice Answer | |||||
| SpaceR-7B (SG-RLVR) ( Ouyang et al., 2025 ) | 30.5 | 30.7 | 34.5 | 52.0 | 18.6 | 22.8 | 34.5 | 20.2 |
| Spatial-MLLM-4B-135k ( Wu et al., 2025 ) | 40.5 | 40.7 | 45.3 | 46.8 | – | 32.3 | 37.4 | – |
| Spatial-MLLM-4B-820k ( Wu et al., 2025 ) | 40.9 | 41.5 | 40.0 | 53.1 | – | 30.7 | 39.2 | – |
| VST-7B-SFT ( Yang et al., 2025c ) | 46.4 | 35.4 | 52.6 | 67.9 | 47.2 | 49.2 | 36.9 | 35.4 |
| VG-LLM-8B ( Zheng et al., 2025a ) | 46.4 | 37.8 | 53.0 | 56.8 | 48.0 | 57.2 | 33.8 | 38.0 |
| Obj. Count | Abs. Dist. | Obj. Size | Room Size | Rel. Dist. | Rel. Dir. | Route Plan | Appr. Order | ||
| Method | Avg. | Numerical Answer | Multiple-Choice Answer | ||||||
| Proprietary Models (API) | |||||||||
| GPT-4o | 34.0 | 46.2 | 5.3 | 43.8 | 38.2 | 37.0 | 41.3 | 31.5 | 28.5 |
| Gemini-1.5-Flash | 42.1 | 49.8 | 30.8 | 53.5 | 54.4 | 37.7 | 41.0 | 31.5 | 37.8 |
| Gemini-1.5-Pro | 45.4 | 56.2 | 30.9 | 64.1 | 43.6 | 51.3 | 46.3 | 36.0 | 34.6 |
| Open-source Models | |||||||||
| Obj. Count | Abs. Dist. | Obj. Size | Room Size | Rel. Dist. | Rel. Dir. | Route Plan | Appr. Order | ||
| Method | Avg. | Numerical Answer | Multiple-Choice Answer | ||||||
| Qwen3-VL-8B ( Bai et al., 2025a ) | 59.8 | 67.5 | 52.6 | 76.2 | 62.3 | 60.6 | 52.5 | 32.5 | 73.8 |
| VLM-3R-7B ( Fan et al., 2026 ) | 60.9 | 70.2 | 49.4 | 69.2 | 67.1 | 65.4 | 80.5 | 45.4 | 40.1 |
| Map2Thought-7B ( Gao et al., 2026 ) | 61.0 | 70.8 | 55.0 | 70.1 | 69.4 | 56.9 | 69.8 | 38.1 | 57.4 |
| VST-7B ( Yang et al., 2025c ) | 61.2 | - | - | - | - | - | - | - | - |
| VG-LLM-8B ( Zheng et al., 2025a ) | 62.2 | 71.4 | 56.8 | 69.0 | 69.1 | 67.9 | 83.2 | 47.4 | 32.5 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Method | Avg. | Low | Depth-OC | Depth-OC-MV | Depth-OO | Depth-OO-MV | Dist-OC | Dist-OC-MV | Dist-OO | Dist-OO-MV | Medium | PosMatch | CamMotion | ViewChgI | High | DistI-OO | DistI-OO-MV | ObjRel-OC-MV | ObjRel-OO | ObjRel-OO-MV | SpImag-OC | SpImag-OC-MV | SpImag-OO | SpImag-OO-MV |
| InternVL2-2B | 28.1 | 21.7 | 18.1 | 24.8 | 23.2 | 21.0 | 19.5 | 20.0 | 26.8 | 20.6 | 22.8 | 39.7 | 23.0 | 5.8 | 35.4 | 51.2 | 56.0 | 46.0 | 31.6 | 23.8 | 36.0 | 34.3 | 17.6 | 22.4 |
| InternVL2-4B | 32.0 | 28.9 | 23.9 | 27.2 | 20.0 | 18.1 | 42.6 | 40.2 | 31.3 | 28.2 | 29.2 | 49.9 | 21.0 | 16.6 | 35.7 | 56.8 | 55.4 | 40.3 | 36.8 | 25.2 | 28.8 | 32.3 | 21.2 | 24.7 |
| InternVL2.5-2B | 30.1 | 25.8 | 39.7 | 39.7 | 12.1 | 15.0 | 30.9 | 29.6 | 20.2 | 19.0 | 22.9 | 37.9 | 24.3 | 6.6 | 36.4 | 51.5 | 56.9 | 50.3 | 33.8 | 24.1 | 27.2 | 35.2 | 26.5 | 22.4 |
| InternVL2.5-4B | 30.6 | 25.7 | 29.1 | 33.0 | 21.8 | 16.8 | 20.8 | 26.9 | 28.1 | 28.8 | 29.8 | 47.1 | 33.3 | 8.9 | 35.2 | 54.1 | 58.9 | 35.5 | 29.7 | 34.6 | 24.7 | 31.4 | 19.2 | 28.3 |
| InternVL2.5-8B | 36.3 | 29.5 | 25.8 | 29.3 | 23.8 | 18.8 | 46.8 | 42.7 | 22.6 | 25.9 | 31.9 | 61.3 | 28.0 | 6.3 | 43.8 | 59.7 | 56.9 | 51.8 | 44.2 | 41.6 | 36.6 | 41.6 | 22.5 | 39.5 |
| LLaVA-OV-0.5B | 29.5 | 30.1 | 49.2 | 42.7 | 18.0 | 14.9 | 31.5 | 25.7 | 29.0 | 30.1 | 15.9 | 24.4 | 21.8 | 1.5 | 33.4 | 50.9 | 50.0 | 32.0 | 27.8 | 26.0 | 30.9 | 34.0 | 24.5 | 24.7 |
| Hyperparameter | Value |
| accelerator model | H800 GPUs |
| accelerator count | 8 |
| global batch size | 32 |
| tune_mm_llm | True |
| tune_mm_vision | False |
| tune_mm_mlp | False |
| Hyperparameter | Value |
| accelerator model | H800 GPUs |
| accelerator count | 8 |
| global batch size | 64 |
| tune_mm_llm | True |
| tune_mm_vision | False |
| tune_mm_mlp | False |