cs.AISep 28, 2026

VCN-Bench: A Video-Contextualized Navigation Benchmark for Spatial Reasoning over Prior Visual Experience

Authors: Siqi Zhang, Meng Wei, Chenyang Wan, Shaohao Zhu, Shufan Shen, Xihui Liu, Zhihua Wei, Tai Wang, +1 more

Organizations: Tongji University · The University of Hong Kong · Zhejiang University · Institute of Computing Technology, Chinese Academy of Sciences · Shanghai AI Laboratory

Abstract

Spatial reasoning is fundamental to embodied agents, yet it remains unclear whether spatial understanding can be carried forward to guide sequential interactions. Existing spatial-reasoning benchmarks typically terminate at offline predictions, while navigation benchmarks evaluate spatial reasoning as part of instruction following and exploration. We introduce VCN-Bench, a \textbf{V}ideo-\textbf{C}ontextualized \textbf{N}avigation benchmark for probing closed-loop spatial reasoning over prior visual experience in MLLMs. Given a prior video covering both the initial location and destination, the agent is tasked with reasoning out the instruction-specified target and navigating toward it with the inferred spatial context. Built on Matterport3D, VCN-Bench contains five instruction types, 100k training episodes, and 1,250 evaluation episodes. Navigation serves as the primary evaluation, while diagnostic goal identification helps distinguish destination-resolution errors from subsequent navigation failures. We further propose MV-DualVLN, a planning-oriented baseline that jointly leverages prior video and in-episode observations. Experiments reveal limited navigation performance, a substantial destination-resolution-to-navigation gap, and frequent navigation failures even after correct destination identification.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Beyond Spatial Benchmarks: From Spatial Reasoning to Navigation

    Sep 24, 2026Xun Huang, Shijia Zhao, Rongsheng Qu +5Spatial ReasoningSpatial Supervision

  2. SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning

    Jun 8, 2026Yucheng Deng, Pingrui Lai, Xinhai Li +5Vision-Language NavigationSpatial Memory

  3. GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

    Sep 16, 2026Kailing Li, Yu Han, Tianwen Qian +4Vision-Language NavigationGrounding