Existing approaches to multi-view spatial reasoning operate largely on sparse input views. Vision-language models (VLMs) are thus restricted to understand a scene and infer spatial relations within these fixed views, leading to fragile cross-view alignment and geometry-to-language bottleneck. To address these issues, we formulate a novel Seek-and-View reasoning approach to find implicit cross-view spatial evidence by locating a question-relevant view to support the spatial reasoning. To realize this approach, we propose Vantage, a training-free model-agnostic reasoning framework that pairs a VLM with a 3D foundation model: a viewpoint-grounded reasoning stage for question analysis and view planning, followed by a geometry-grounded evidence augmentation stage to effectively synthesize and incorporate visual evidence into the final reasoning. Comprehensive experiments on six VLMs demonstrate consistent improvements on five benchmarks without fine-tuning. Overall, by revealing spatial evidence through view-grounded reasoning, Vantage can largely reduce reliance on language-based cross-view alignment and improve multi-view spatial understanding. Our code is available at https://github.com/q1xiangchen/Vantage.
Figures & tables
Figure 1: (a) Existing approaches reason over fixed views, suffering from fragile cross-view alignment and the geometry-to-language bottleneck. (b) Our novel Seek-and-View reasoning paradigm seeks a question-relevant view (see bottom right) to make the spatial evidence directly observable.
Figure 2: Overview of Vantage pipeline. Stage 1: Viewpoint-grounded reasoning employs an VLM (in blue) to (a) produce an analysis on the input question (top-right), select reference view vr , identify queried spatial relation, then (b) plan a question-relevant camera action Δv on reference view vr and generate reasoning guidance for supporting view interpretation. Stage 2: (c) Geometry-grounded evidence augmentation synthesizes the planned view vs from the 3DFM reconstruction (in green) and provides it together with the original images, viewpoint analysis cana , and reasoning guidance cguide to support (d) the final reasoning.
Model
MindCube-tiny
MMSI-Bench
BLINK
OmniSpatial
SPINBench
Avg.
% Gain
Rot.
Ard.
Amg.
All
PR.
Attr.
All
MV
Pers.
D.Rot
D.Tr
All
Training-Scaled Spatial Model
SenseNova-SI-1.5-InternVL3-8B
92.5
90.8
94.5
93.24
47.1
37.7
45.2
63.9
52.0
43.6
41.7
43.0
59.5
-
Model-Scaled Spatial Agent
GCA (Qwen3-VL-235B-A22B-Thinking) †
82.0
61.8
59.8
64.2
52.8
45.0
51.2
-
58.6
-
-
-
-
-
Vantage on base models
Table 1: Main results on five benchmarks. For each base model, the better result is shown in bold . The final two columns report the average accuracy across all five benchmarks and the relative improvement. † Results are reported as in their original papers.
Model
Question Analysis
View Planning
MindCube-tiny
MMSI-Bench
BLINK
OmniSpatial
Acc.
Δ
Acc.
Δ
Acc.
Δ
Acc.
Δ
Qwen3-VL-4B
–
–
26.0
−12.9
27.9
−5.8
41.4
−0.7
42.1
−5.1
Self
Self
38.9
–
33.7
–
42.1
–
47.2
–
Cross
Self
51.0
+12.1
35.3
+1.6
37.6
−4.5
50.6
+3.4
Cross
Cross
49.9
+11.0
36.2
+2.5
36.8
−5.3
48.1
+0.9
Qwen3-VL-8B
–
–
31.3
−0.3
30.1
0.0
44.4
−0.7
43.9
−1.2
Table 2: Cross-source composition results. Self denotes steps executed by the evaluated model, while Cross denotes steps replaced by Qwen3.6-27B. Bold indicates the highest accuracy within each model group. Δ denotes the accuracy change relative to model + Vantage.
Table 5
Model
Augmentation
Benchmark
Avg.
Abs Δ
Relative Δ
View
Analysis
Guidance
MindCube-tiny
MMSI-Bench
OmniSpatial
Qwen3-VL-4B
26.0
27.9
42.1
32.0
–
–
✓
✓
34.5
32.1
43.1
36.6
+4.6
+14.4
✓
29.7
32.7
43.5
35.3
+3.3
+10.3
✓
✓
38.2
34.0
42.4
38.2
+6.2
+19.4
✓
✓
36.7
31.4
42.1
36.7
+4.7
+14.7
Table 5: Final VQA context ablation. Checkmarks indicate the component provided to the original VQA. Bold and italic denote the best and second-best average accuracy within each model, respectively. The final two columns report the absolute and relative improvements over its baseline.
Figure 3: Major failure types of Vantage. Left : Macro-averaged failure analysis. Right : Representative failure cases by type. More detailed visualizations of these representative failure cases are provided in the Appendix.
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Category
Action
Transformation
Small
Medium
Large
Rotation
Pan left
+Δyaw
30∘
60∘
90∘
Pan right
−Δyaw
30∘
60∘
90∘
Tilt up
+Δpitch
30∘
60∘
90∘
Tilt down
−Δpitch
30∘
60∘
90∘
Turn around
Δyaw
180∘
Translation
Move right
+Δx
0.40s
0.80s
1.40s
Appendix
Table 6: Action set for view planning. Each semantic action is deterministically mapped to a geometric transformation with three magnitude levels (small, medium, and large), except for turn around , which uses a fixed 180∘ rotation. Translation magnitudes are expressed in scene unit s .
Figure 4: Effect of gravity alignment on view synthesis. Given the same input views, reference view, and planned camera action, we compare view synthesis using a gravity-aligned G3T reconstruction and a non-gravity-aligned VGGT- Ω reconstruction. Gravity alignment provides a ground-consistent coordinate frame for executing relative camera actions, while an unaligned reconstruction can cause the realized viewpoint to deviate from the intended ground-relative motion.
Model
FoV
MindCube-tiny
Δ
Qwen3.6-27B
120∘ (ours)
68.8
–
80∘
68.1
−0.7
Gemma-4-31B
120∘ (ours)
61.9
–
80∘
59.4
−2.5
Appendix
Table 7: Field-of-view ablation on MindCube.
Model
Backend
MindCube-tiny
MMSI-Bench
Avg.
Qwen3-VL-4B
G3T
38.9
33.7
36.3
VGGT- Ω + GeoCalib
37.4
30.7
34.1
Qwen3-VL-8B
G3T
31.6
30.1
30.9
VGGT- Ω + GeoCalib
31.6
29.9
30.8
Qwen3.6-27B
G3T
68.8
44.8
56.8
VGGT- Ω + GeoCalib
67.2
42.2
54.7
Appendix
Table 8: Reconstruction backends ablation.
Figure 5: Failure distribution by model. For each model, we annotate 150 incorrect predictions sampled across the three multi-view benchmarks and categorize them into five failure types. Bars indicate the proportion of annotated failures falling into each category.
Figure 6: Qualitative example of Vantage with GPT-5.4. For clarity, we simplify the original model outputs and present only the key intermediate information relevant to the reasoning process.
Figure 7: Qualitative example of Vantage with GPT-5.4.
Figure 8: Qualitative example of Vantage with GPT-5.4.
Figure 9: Qualitative example of Vantage with GPT-5.4.
Figure 10: Incorrect analysis. Qualitative failure case from Qwen3.6-27B.
Figure 11: Incorrect view plan. Qualitative failure case from Qwen3-VL-4B.
Figure 12: Unreliable synthesis. Qualitative failure case from Gemma-4-31B.
Figure 13: Reasoning failure. Qualitative failure case from Qwen3-VL-4B.
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China, China · Huawei Noah’s Ark Lab · Hefei University of Technology, China
Key Laboratory of Multimedia Trusted Perception and Efficient Computing, Ministry of Education of China, Xiamen University, 361005, P.R. China · Tencent Youtu Lab · Beijing Institute of Technology