cs.AISep 29, 2026

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

Authors: Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, +3 more

Organizations: Tsinghua University · Pengcheng Laboratory · South China University of Technology · International Digital Economy Academy · Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China · Harbin Institute of Technology (Shenzhen)

Abstract

Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.

Figures & tables

Appendix figures & tables7 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

    Sep 16, 2026Kailing Li, Yu Han, Tianwen Qian +4Vision-Language NavigationGrounding

  2. ReflectVLN: Training Vision-Language Navigation Agents with Reflective Reasoning

    Jul 14, 2026Jiahang Wang, Yirong Yang, Yanqing Zhu +4Vision-Language NavigationNavigation

  3. StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation

    Oct 5, 2026Anh Dao, Quan-Dung Pham, Le Danh Vinh +6Vision-Language Navigation