Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: https://yuqunw.github.io/Render2Reason/.
Figures & tables
Figure 1: Render to Reason. Left: We introduce novel-view semantic rendering as an auxiliary training task, which takes input images and a target camera token, and renders the semantic layout of the unseen view. Right: On ReVSI ( Zhang et al., 2026c ) and VSI-Bench ( Yang et al., 2025a ) , adding geometry features under standard QA training yields only marginal gains (+0.9, +0.1). When trained with our novel-view semantic rendering, adding geometry features yields larger gains (+2.1, +3.0), showing that the auxiliary task encourages the model to integrate geometry features more effectively.
Figure 2: Model Architecture. Our model jointly supports semantic novel-view rendering and standard QA, sharing a backbone and differing only in prompt and output head. Input views are encoded into visual and geometry features, fused via cross-attention, and passed to the LLM decoder. For novel view semantic rendering, we extract a camera token from the target view via the geometry encoder and prepend it to the prompt; the LLM produces 256 output tokens, one per target patch in raster order, each mapped to a semantic class by a 2-layer MLP. For standard QA, the prompt only contains the question, and the LLM generates the answer without using the rendering MLP.
Figure 3: Modality Ablation for Novel View Semantic Rendering. The model predicts a semantic map of an unseen viewpoint, given input frames and the target camera pose. Top row: input frames. Bottom row: target view. With both modalities, the prediction matches the target geometry. With vision only, the prediction fails to reach the target pose and resembles the last input frame , indicating that the geometry features are required to render at the correct viewpoint.
Figure 4: Visualization of P2P relative distance in 3D-Point-QA. We render arrows as prompts to make recognition easier for VLMs ( Xu et al., 2025 ) . See more visualizations in Appendix A.6 .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Methods
Backbone
#Params
#QA
Avg.
Numerical Answer
Multiple-Choice Answer
Baseline
Chance (Frequency)
–
–
–
34.0
62.1
32.0
29.9
33.1
25.1
47.9
28.4
25.2
Proprietary Models (API)
GPT-4o
–
–
–
34.0
46.2
5.3
43.8
38.2
37.0
41.3
31.5
28.5
Gemini-2.5 Pro
–
–
–
51.5
43.8
34.9
64.3
42.8
61.1
47.8
45.9
71.3
Table 1: VSI-Bench sub-task breakdown. Best results within each model group are bolded . #QA denotes the number of spatial training QAs. 9B parameters include the frozen geometry encoder.
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Methods
Backbone
#Params
#QA
Avg.
Numerical Answer
Multiple-Choice Answer
Baseline
Chance (Frequency)
–
–
–
31.4
52.2
40.1
17.4
20.9
25.8
31.9
30.2
Proprietary Models (64+ Frames)
GPT-5.2
–
–
–
50.9
56.2
41.5
73.9
63.0
48.4
34.9
38.2
Gemini 3 Pro
–
–
–
60.9
60.1
54.7
79.3
51.9
68.1
56.0
56.4
Table 2: ReVSI sub-task breakdown. Best results within each section are bolded . #QA denotes the number of geometry QAs generated with ground truth. 9B parameters include the geometry encoder.
Table 3: Ablation studies. We report results on VSI-Bench and ReVSI (16 frames) and on 3D-Point-QA. Sem. refers to semantic rendering. Best results are bolded . In (e), all models include VGGT features; we disable the geometry pathway at inference. See Appendix A.4 for full results.
Figure 5: Semantic Rendering Visualization. We visualize novel-view semantic renderings along a forward-moving camera trajectory. The first row shows the five initial RGB input frames. Each prediction uses the five RGB frames right before its target frame: the first prediction uses all five frames in row 1; the second uses frames 2–5 from row 1 and the first RGB frame in row 3; and so on.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Semantic Rendering Visualization. We visualize sequential novel-view semantic renderings in a forward-moving sequence, with one-second intervals between adjacent views. The first row shows the 5 input frames. Each subsequent prediction uses the previous 5 frames as input (sliding window): the first prediction in row 2 uses all 5 frames from row 1; the second prediction uses frames 2-5 of row 1 plus the first RGB frame in row 3; and so on.
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Methods
Avg.
Numerical Answer
Multiple-Choice Answer
Finetuned
62.9
70.6
45.5
70.2
62.6
59.3
70.8
40.7
83.3
+ VGGT
63.0
69.7
47.6
67.8
66.3
58.9
75.5
42.8
75.2
+ Sem.
61.6
69.3
41.3
69.5
63.4
61.0
66.0
40.7
81.6
+ VGGT + RGB
62.4
69.7
42.7
69.2
64.2
61.0
72.2
45.4
74.9
+ VGGT + RGB + Sem.
62.6
69.0
42.8
68.1
65.5
60.1
74.0
44.3
77.0
Appendix
Table 4: Full results of the component and rendering-target ablations on VSI-Bench. All models use InternVL3.5-4B with 16 input frames. Best results are bolded .
ReVSI-16-Frame
3D-Point-QA
Methods
Avg.
Obj. Cnt.
Abs. Dist.
Obj. Size
Rel. Dist.
Rel. Dir.
Avg.
P2C Dist.
P2P Dist.
P2P R-Dist.
Point Match
3D Mapping
Finetuned
44.9
33.6
53.4
57.7
35.0
45.0
55.3
66.0
41.2
90.1
34.8
44.7
+ VGGT
45.8
32.3
55.0
57.0
36.4
48.1
57.0
59.6
44.0
92.5
44.1
45.0
+ Sem.
45.9
34.6
51.7
58.4
36.1
48.5
60.6
69.5
54.6
91.3
38.2
49.3
+ VGGT + RGB
45.2
33.1
50.5
58.1
35.3
49.1
60.2
65.3
44.4
92.1
48.6
50.7
+ VGGT + RGB + Sem.
46.6
35.4
55.0
58.7
36.1
47.9
61.1
70.3
52.9
93.3
38.5
50.6
Appendix
Table 5: Full results of the component and rendering-target ablations on ReVSI-16-Frame and 3D-Point-QA. All models use InternVL3.5-4B. Best results are bolded .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Methods
Avg.
Numerical Answer
Multiple-Choice Answer
Finetuned
66.0
70.7
47.0
75.8
71.8
64.9
71.1
43.8
83.2
+ VGGT
67.0
68.8
51.6
73.6
71.6
63.0
79.3
43.3
85.0
+ VGGT + GeoSR
67.2
71.4
53.7
73.0
63.2
67.5
80.8
44.3
84.0
+ VGGT + Sem.
68.2
70.6
56.0
75.7
65.8
72.8
77.5
41.2
85.6
Appendix
Table 6: Full results of the component ablation and GeoSR comparison on VSI-Bench. All models use Qwen3-VL-4B with 16 input frames. Best results are bolded .
ReVSI-16-Frame
3D-Point-QA
Methods
Avg.
Obj. Cnt.
Abs. Dist.
Obj. Size
Rel. Dist.
Rel. Dir.
Avg.
P2C Dist.
P2P Dist.
P2P R-Dist.
Point Match
3D Mapping
Finetuned
50.7
38.6
59.8
68.4
37.8
49.0
72.2
81.7
63.1
92.7
64.5
58.9
+ VGGT
51.5
39.2
65.1
64.3
38.2
50.4
77.1
82.5
68.6
95.6
69.9
68.9
+ VGGT + GeoSR
50.7
39.0
65.4
63.1
39.6
46.6
84.8
90.2
78.8
97.0
75.1
83.0
+ VGGT + Sem.
52.5
39.0
72.6
62.8
39.1
48.8
80.0
85.3
69.9
96.6
72.6
75.7
Appendix
Table 7: Full results of the component ablation and GeoSR comparison on ReVSI-16-Frame and 3D-Point-QA. All models use Qwen3-VL-4B. Best results are bolded .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Grid
Avg.
Numerical Answer
Multiple-Choice Answer
8×8
64.9
69.5
47.0
68.0
64.9
64.8
75.7
48.5
80.7
16×16
64.6
70.4
44.7
69.3
64.6
61.8
78.3
47.9
79.6
32×32
64.3
70.6
46.3
69.2
61.7
63.4
77.0
45.9
80.6
Appendix
Table 8: Full results of the grid resolution ablation on VSI-Bench. All models use InternVL3.5-4B with VGGT features, semantic rendering, and 16 input frames. Best results are bolded .
ReVSI-16-Frame
3D-Point-QA
Grid
Avg.
Obj. Cnt.
Abs. Dist.
Obj. Size
Rel. Dist.
Rel. Dir.
Avg.
P2C Dist.
P2P Dist.
P2P R-Dist.
Point Match
3D Mapping
8×8
46.2
34.2
55.4
55.3
38.7
47.7
53.0
65.2
44.5
85.0
26.4
43.9
16×16
48.0
37.4
55.6
57.1
38.4
51.6
59.9
67.9
52.9
91.3
38.3
49.2
32×32
48.4
36.4
60.5
59.4
38.4
47.4
56.6
71.5
51.2
86.8
29.1
44.1
Appendix
Table 9: Full results of the grid resolution ablation on ReVSI-16-Frame and 3D-Point-QA. All models use InternVL3.5-4B with VGGT features and semantic rendering. Best results are bolded .
Figure 7: Visualization of 3D-Point-QA examples. We visualize the question and answer for each sample.
Methods
Rotation
Among
Around
Overall
Pretrained
37.5
35.8
42.0
37.6
Finetuned
30.5
39.3
48.4
39.8
+ VGGT
37.0
38.2
34.4
37.1
+ VGGT + Sem.
32.0
41.5
50.8
41.9
Appendix
Table 10: Zero-shot results on MindCube-Tiny. All models are based on InternVL3.5-4B and use the checkpoints from Tab. 3(a) without additional training. Best results are bolded .
Vision-Language Models (VLMs) often struggle with robust 3D spatial reasoning. Prevailing methods that rely on fine-tuning with 3D visual question-answering (VQA) datasets may overfit dataset-specific biases, while integrating specialized 3D visual encoders is often inflexible and cumbersome. In this paper, we argue that genuine spatial understanding should emerge from learning fundamental geometric priors, not only from high-level VQA supervision. We propose GASP (Geometric-Aware Spatial Priors), a framework that injects these priors directly into the LLM's transformer layers. GASP employs a small correspondence head, applied as a deep supervision signal across all layers, and is trained with a dual objective leveraging ground-truth geometry from large-scale video scenes: a contrastive loss on ground-truth point correspondences enforces 2D view-invariance, while a depth consistency supervision resolves 3D geometric ambiguities. Our analysis first provides a diagnostic showing that standard VLMs' internal correspondence matching accuracy is very low (often below 5%). We then demonstrate that our training substantially improves this behavior, boosting peak layer-wise correspondence to over 70% and maintaining over 85% temporal robustness while baselines remain below 5%. These internal improvements translate to significant gains on downstream spatial benchmarks including +18.2% on All-Angles Bench and +29.0% on VSI-Bench, all without training on any 3D VQA data. Our findings indicate that learning from fundamental geometric priors is a promising and generalizable pathway towards VLMs with more reliable 3D spatial reasoning.
Vision-language models (VLMs) can benefit from geometric priors for multi-view spatial reasoning, yet answer-only training does not directly supervise the intermediate geometric estimates and their use in deriving quantitative spatial answers. We hypothesize that spatial chain-of-thought (CoT) supervision becomes more effective when the VLM first jointly learns complementary local geometry and global scene context through multi-view reconstruction. We introduce SpatialSpeak, a two-stage framework that connects QA-native reconstruction pretraining with spatial CoT learning. In Stage I, QA-Native Reconstruction Pretraining (QA-RP) combines marked-point 3D queries for fine-grained local geometry with object-center queries for global scene context across views. Both tasks are formulated as text-based question answering, allowing geometric estimation and subsequent reasoning to share the same autoregressive output interface. In Stage II, spatial CoT with Visual Compensation (CoT-VC) trains the model to express question-relevant geometric estimates and use them to derive answers, with reliability assessment and visual compensation supporting answer refinement when needed. On ReVSI, QA-RP increases the gain from CoT-VC from 2.6 to 6.9 points, and ablations show that both local and global reconstruction supervision are beneficial. SpatialSpeak achieves state-of-the-art results on ReVSI, VSI-Bench, and SPAR-Bench, with a ReVSI score of 62.8 that exceeds the strongest compared baseline by 8.7 points.
Yang Cao, Jiaxin Zhang, Dave Zhenyu Chen +4
Hong Kong University of Science and Technology · Harbin Institute of Technology · Huawei Noah’s Ark Lab
Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging. Moreover, different spatial queries call for fundamentally different strategies: some are best addressed through purely linguistic, step-by-step deduction, while others require explicit 3D grounding before quantitative inference. We present Dual-Path Spatial Reasoning via Reinforcement Learning for Spatial VLMs (SR-REAL), a unified framework that equips a spatial VLM with two complementary reasoning paths: Language-Only Reasoning (LOR), which performs step-by-step linguistic deduction, and Detect-Then-Reason (DTR), which detects 3D geometric cues (e.g., centers or bounding boxes) via region tokens before explicit geometric inference. SR-REAL begins with a cold-start supervised fine-tuning stage that constructs LOR and DTR chain-of-thought supervision and exposes a region-to-3D interface, followed by RL that optimizes the policy model with accuracy and format rewards; for DTR, a discrete center-based detection reward further refines geometric alignment. Across diverse spatial benchmarks, SR-REAL significantly outperforms spatial VLM baselines: (i) a single RL-trained model supports both reasoning paths, with DTR excelling in region-aware tasks through precise 3D localization and LOR enhancing general spatial reasoning; (ii) jointly training both paths fosters mutual reinforcement; (iii) high-quality, blended cold-start data is crucial for stable RL optimization; and (iv) the model generalizes across datasets and domains without per-task tuning, demonstrating positive transfer between LOR and DTR.
Yatai Ji, An-Chieh Cheng, Yang Fu +13
The University of Hong Kong · NVIDIA · University of California, San Diego