VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience
Authors: Siqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu, Shufan Shen, Xihui Liu, Zhihua Wei, Tai Wang, +1 more
Organizations: Tongji University · The University of Hong Kong · Zhejiang University · Institute of Computing Technology, Chinese Academy of Sciences · Shanghai AI Laboratory
Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.
Figures & tables
Figure 1: VCN-Bench is a video-contextualized navigation benchmark tailored for probing closed-loop spatial reasoning over prior visual experience in MLLMs. The spatial context inferred from the prior video directly informs destination resolution and guides closed-loop action decisions.
Figure 2: Demonstration of VCN-Bench. (a) Prior video with room labels. (b) Top-down view of the prior video. (c) The reasoning process during navigation. Navigation tasks (d) are conditioned on the prior video with target room labels.
Figure 3: Instruction descriptions and primary dimensions of reasoning over prior visual experience.
Figure 4: Benchmark construction pipeline.
Figure 5: The overall framework of MV-DualVLN.
Room-to-Object
Object
Models
Metrics
Average
Count
Order
Distance
Orientation
Instance
Hierarchical methods
SPL
4.8
3.4
3.4
6.2
3.3
7.9
SR
17.6
15.6
18.4
18.4
8.8
26.8
Gemini-3+MTU3D
RSR
20.5
18.8
21.1
20.0
9.6
33.2
SPL
5.3
5.3
4.7
5.0
3.7
7.6
Table 1: Comparison of navigation performance. † : UniNavid directly outputs low-level actions. Other models utilize an oracle executor to move to the predicted waypoint.
Room-to-Object (Rid-SR)
Object (Oid-SR)
Models
Count
Order
Distance
Orientation
Avg.
Instance
Baseline
Complete Chance Level
2.0
2.4
5.6
4.0
3.5
0.8
Category Chance Level
13.2
14.4
12.8
17.2
14.4
14.0
Tiny Set
Human Performance
93.3
95.0
86.7
83.3
89.6
95.0
Table 2: Comparison of goal identification performance.
#
Models
SPL
SR
RSR
1
MV-DualVLN
19.1
27.6
33.9
2
target frame only
12.2
19.7
24.8
3
50 frames, w/o history
14.9
24.1
32.8
4
100 frames, w/o history
16.4
25.0
33.7
Table 3: Ablation results for navigation. Variants are trained under the ablation setting.
Figure 9
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 10
Figure 8: Templates and tags of instructions.
Figure 9: Example of VCN-Bench. In the top-down map, we plot the video trajectory with a green curve, which starts from the red flag and ends at the black flag. The red and yellow stars are the markers for the initialization and target positions of each episode. “FR”, “RR”, “R”, “RL”, “FL”, “F” are the abbreviations of Front-Right, Rear-Right, Rear, Rear-Left, Front-Left, Front, respectively.
Figure 13
Figure 10: Benchmark statistics, including (a) the number of prior video frames, (b) number of rooms and (c) area (the entire area of any region visited is included) covered by the prior video, (d) number of words in instructions, (e) the shortest geodesic distance, and (f) action steps of episodes.
Figure 11: Prompt for zero-shot goal identification.
Figure 12: Error analysis of goal identification.
Figure 13: Visualization of Disturbance and Position error.
Figure 14: Visualization of Cognition error.
Models
SPL
SR
RSR
MV-DualVLN
19.1
27.6
33.9
GT frame prediction
22.8
33.1
39.9
Appendix
Table 19
MV-DualVLN
Counterfactual Relation
Disordered Video
SPL
SR
RSR
SPL
SR
RSR
SPL
SR
RSR
Count
22.0
30.8
38.0
7.1
9.9
15.0
8.8
10.8
16.8
Order
23.5
34.0
40.4
4.2
6.4
10.8
7.3
10.0
15.6
Distance
14.1
20.8
28.2
6.5
9.2
11.6
7.8
11.6
17.2
Orientation
11.9
18.0
20.4
4.4
7.2
13.2
6.6
9.2
14.8
Instance
23.7
34.4
42.4
-
-
-
11.9
16.8
26.0
Appendix
Table 4: Navigation performance comparison on instruction types with interventions.
School of Information Science and Electronic Engineering & School of Integrated Circuits, Shanghai Jiao Tong University · Institute of Artificial Intelligence, China Telecom · 3Central South University +1