Recent works augment Vision-Language Models with geometry features from pretrained 3D models, expecting that the geometric signal will boost spatial reasoning. However, we find that simply fusing geometry features and training on standard spatial QA yields only marginal improvements on high-level multi-hop tasks. We attribute this gap to a training-signal problem: standard spatial QA can be largely answered from visual features and language priors, so the geometry pathway receives weak gradients and fails to integrate with the visual features. To provide a training signal that requires geometry, we propose \textbf{novel-view semantic rendering} as an auxiliary training task that requires the model to predict the semantic layout of an unobserved viewpoint, inspired by humans' ability to mentally simulate novel viewpoints during spatial reasoning. This task encourages joint use of both pathways: geometry provides pose-dependent visibility, while vision provides semantic content. Our auxiliary task yields consistent improvements over the geometry-augmented baseline across all three benchmarks (up to +1.6 on VSI-Bench, +2.2 on ReVSI, +2.9 on our 3D-Point-QA dataset) and our full model surpasses prior open-source methods on VSI-Bench and on ReVSI. Project page: https://yuqunw.github.io/Render2Reason/.
Figures & tables
Figure 1: Render to Reason. Left: We introduce novel-view semantic rendering as an auxiliary training task, which takes input images and a target camera token, and renders the semantic layout of the unseen view. Right: On ReVSI ( Zhang et al., 2026c ) and VSI-Bench ( Yang et al., 2025a ) , adding geometry features under standard QA training yields only marginal gains (+0.9, +0.1). When trained with our novel-view semantic rendering, adding geometry features yields larger gains (+2.1, +3.0), showing that the auxiliary task encourages the model to integrate geometry features more effectively.
Figure 2: Model Architecture. Our model jointly supports semantic novel-view rendering and standard QA, sharing a backbone and differing only in prompt and output head. Input views are encoded into visual and geometry features, fused via cross-attention, and passed to the LLM decoder. For novel view semantic rendering, we extract a camera token from the target view via the geometry encoder and prepend it to the prompt; the LLM produces 256 output tokens, one per target patch in raster order, each mapped to a semantic class by a 2-layer MLP. For standard QA, the prompt only contains the question, and the LLM generates the answer without using the rendering MLP.
Figure 3: Modality Ablation for Novel View Semantic Rendering. The model predicts a semantic map of an unseen viewpoint, given input frames and the target camera pose. Top row: input frames. Bottom row: target view. With both modalities, the prediction matches the target geometry. With vision only, the prediction fails to reach the target pose and resembles the last input frame , indicating that the geometry features are required to render at the correct viewpoint.
Figure 4: Visualization of P2P relative distance in 3D-Point-QA. We render arrows as prompts to make recognition easier for VLMs ( Xu et al., 2025 ) . See more visualizations in Appendix A.6 .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Methods
Backbone
#Params
#QA
Avg.
Numerical Answer
Multiple-Choice Answer
Baseline
Chance (Frequency)
–
–
–
34.0
62.1
32.0
29.9
33.1
25.1
47.9
28.4
25.2
Proprietary Models (API)
GPT-4o
–
–
–
34.0
46.2
5.3
43.8
38.2
37.0
41.3
31.5
28.5
Gemini-2.5 Pro
–
–
–
51.5
43.8
34.9
64.3
42.8
61.1
47.8
45.9
71.3
Table 1: VSI-Bench sub-task breakdown. Best results within each model group are bolded . #QA denotes the number of spatial training QAs. 9B parameters include the frozen geometry encoder.
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Methods
Backbone
#Params
#QA
Avg.
Numerical Answer
Multiple-Choice Answer
Baseline
Chance (Frequency)
–
–
–
31.4
52.2
40.1
17.4
20.9
25.8
31.9
30.2
Proprietary Models (64+ Frames)
GPT-5.2
–
–
–
50.9
56.2
41.5
73.9
63.0
48.4
34.9
38.2
Gemini 3 Pro
–
–
–
60.9
60.1
54.7
79.3
51.9
68.1
56.0
56.4
Table 2: ReVSI sub-task breakdown. Best results within each section are bolded . #QA denotes the number of geometry QAs generated with ground truth. 9B parameters include the geometry encoder.
Table 3: Ablation studies. We report results on VSI-Bench and ReVSI (16 frames) and on 3D-Point-QA. Sem. refers to semantic rendering. Best results are bolded . In (e), all models include VGGT features; we disable the geometry pathway at inference. See Appendix A.4 for full results.
Figure 5: Semantic Rendering Visualization. We visualize novel-view semantic renderings along a forward-moving camera trajectory. The first row shows the five initial RGB input frames. Each prediction uses the five RGB frames right before its target frame: the first prediction uses all five frames in row 1; the second uses frames 2–5 from row 1 and the first RGB frame in row 3; and so on.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Semantic Rendering Visualization. We visualize sequential novel-view semantic renderings in a forward-moving sequence, with one-second intervals between adjacent views. The first row shows the 5 input frames. Each subsequent prediction uses the previous 5 frames as input (sliding window): the first prediction in row 2 uses all 5 frames from row 1; the second prediction uses frames 2-5 of row 1 plus the first RGB frame in row 3; and so on.
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Methods
Avg.
Numerical Answer
Multiple-Choice Answer
Finetuned
62.9
70.6
45.5
70.2
62.6
59.3
70.8
40.7
83.3
+ VGGT
63.0
69.7
47.6
67.8
66.3
58.9
75.5
42.8
75.2
+ Sem.
61.6
69.3
41.3
69.5
63.4
61.0
66.0
40.7
81.6
+ VGGT + RGB
62.4
69.7
42.7
69.2
64.2
61.0
72.2
45.4
74.9
+ VGGT + RGB + Sem.
62.6
69.0
42.8
68.1
65.5
60.1
74.0
44.3
77.0
Appendix
Table 4: Full results of the component and rendering-target ablations on VSI-Bench. All models use InternVL3.5-4B with 16 input frames. Best results are bolded .
ReVSI-16-Frame
3D-Point-QA
Methods
Avg.
Obj. Cnt.
Abs. Dist.
Obj. Size
Rel. Dist.
Rel. Dir.
Avg.
P2C Dist.
P2P Dist.
P2P R-Dist.
Point Match
3D Mapping
Finetuned
44.9
33.6
53.4
57.7
35.0
45.0
55.3
66.0
41.2
90.1
34.8
44.7
+ VGGT
45.8
32.3
55.0
57.0
36.4
48.1
57.0
59.6
44.0
92.5
44.1
45.0
+ Sem.
45.9
34.6
51.7
58.4
36.1
48.5
60.6
69.5
54.6
91.3
38.2
49.3
+ VGGT + RGB
45.2
33.1
50.5
58.1
35.3
49.1
60.2
65.3
44.4
92.1
48.6
50.7
+ VGGT + RGB + Sem.
46.6
35.4
55.0
58.7
36.1
47.9
61.1
70.3
52.9
93.3
38.5
50.6
Appendix
Table 5: Full results of the component and rendering-target ablations on ReVSI-16-Frame and 3D-Point-QA. All models use InternVL3.5-4B. Best results are bolded .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Methods
Avg.
Numerical Answer
Multiple-Choice Answer
Finetuned
66.0
70.7
47.0
75.8
71.8
64.9
71.1
43.8
83.2
+ VGGT
67.0
68.8
51.6
73.6
71.6
63.0
79.3
43.3
85.0
+ VGGT + GeoSR
67.2
71.4
53.7
73.0
63.2
67.5
80.8
44.3
84.0
+ VGGT + Sem.
68.2
70.6
56.0
75.7
65.8
72.8
77.5
41.2
85.6
Appendix
Table 6: Full results of the component ablation and GeoSR comparison on VSI-Bench. All models use Qwen3-VL-4B with 16 input frames. Best results are bolded .
ReVSI-16-Frame
3D-Point-QA
Methods
Avg.
Obj. Cnt.
Abs. Dist.
Obj. Size
Rel. Dist.
Rel. Dir.
Avg.
P2C Dist.
P2P Dist.
P2P R-Dist.
Point Match
3D Mapping
Finetuned
50.7
38.6
59.8
68.4
37.8
49.0
72.2
81.7
63.1
92.7
64.5
58.9
+ VGGT
51.5
39.2
65.1
64.3
38.2
50.4
77.1
82.5
68.6
95.6
69.9
68.9
+ VGGT + GeoSR
50.7
39.0
65.4
63.1
39.6
46.6
84.8
90.2
78.8
97.0
75.1
83.0
+ VGGT + Sem.
52.5
39.0
72.6
62.8
39.1
48.8
80.0
85.3
69.9
96.6
72.6
75.7
Appendix
Table 7: Full results of the component ablation and GeoSR comparison on ReVSI-16-Frame and 3D-Point-QA. All models use Qwen3-VL-4B. Best results are bolded .
Obj. Count
Abs. Dist.
Obj. Size
Room Size
Rel. Dist.
Rel. Dir.
Route Plan
Appr. Order
Grid
Avg.
Numerical Answer
Multiple-Choice Answer
8×8
64.9
69.5
47.0
68.0
64.9
64.8
75.7
48.5
80.7
16×16
64.6
70.4
44.7
69.3
64.6
61.8
78.3
47.9
79.6
32×32
64.3
70.6
46.3
69.2
61.7
63.4
77.0
45.9
80.6
Appendix
Table 8: Full results of the grid resolution ablation on VSI-Bench. All models use InternVL3.5-4B with VGGT features, semantic rendering, and 16 input frames. Best results are bolded .
ReVSI-16-Frame
3D-Point-QA
Grid
Avg.
Obj. Cnt.
Abs. Dist.
Obj. Size
Rel. Dist.
Rel. Dir.
Avg.
P2C Dist.
P2P Dist.
P2P R-Dist.
Point Match
3D Mapping
8×8
46.2
34.2
55.4
55.3
38.7
47.7
53.0
65.2
44.5
85.0
26.4
43.9
16×16
48.0
37.4
55.6
57.1
38.4
51.6
59.9
67.9
52.9
91.3
38.3
49.2
32×32
48.4
36.4
60.5
59.4
38.4
47.4
56.6
71.5
51.2
86.8
29.1
44.1
Appendix
Table 9: Full results of the grid resolution ablation on ReVSI-16-Frame and 3D-Point-QA. All models use InternVL3.5-4B with VGGT features and semantic rendering. Best results are bolded .
Figure 7: Visualization of 3D-Point-QA examples. We visualize the question and answer for each sample.
Methods
Rotation
Among
Around
Overall
Pretrained
37.5
35.8
42.0
37.6
Finetuned
30.5
39.3
48.4
39.8
+ VGGT
37.0
38.2
34.4
37.1
+ VGGT + Sem.
32.0
41.5
50.8
41.9
Appendix
Table 10: Zero-shot results on MindCube-Tiny. All models are based on InternVL3.5-4B and use the checkpoints from Tab. 3(a) without additional training. Best results are bolded .