Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation
Authors: Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, +3 more
Organizations: Tsinghua University · Pengcheng Laboratory · South China University of Technology · International Digital Economy Academy · Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China · Harbin Institute of Technology (Shenzhen)
Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.
Figures & tables
Figure 1: Illustration of Progress Myopia . Left: out-of-view landmarks (a) and ambiguous instructions (b) make the current observation insufficient to determine the correct decision. Right: representative navigation agents show similar decision confidence and action entropy on both successful and failed segments, indicating that unreliable decisions can remain highly confident.
Figure 2: Overview of SeekVLN . (1) The policy decides whether to navigate or seek evidence. (2) FRG reasons backward from future expert actions to determine where to seek evidence and annotate key evidence, while using historical context for progress annotation. (3) C2PO compares evidence-seeking and direct-navigation branches and jointly optimizes evidence-seeking and navigation.
Method
Observation
R2R-CE Val-Unseen
RxR-CE Val-Unseen
S.RGB
Depth
NE ↓
OSR ↑
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
nDTW ↑
BEVBert † An et al. (2022)
✓
4.57
67.0
59.0
50.0
–
–
–
–
ETPNav † An et al. (2024)
✓
4.71
65.0
57.0
49.0
5.64
54.8
44.9
61.9
ENP-ETPNav † Liu et al. (2024)
✓
4.69
65.0
58.0
50.0
5.51
55.3
45.1
63.0
Seq2Seq † Krantz et al. (2020)
✓
✓
7.77
37.0
25.0
22.0
12.10
13.9
11.9
–
CMA † Krantz et al. (2020)
✓
✓
7.37
40.0
32.0
30.0
–
–
–
–
Table 1: Comparison of different methods on the R2R-CE and RxR-CE Val-Unseen splits. Observations include a single RGB camera (S.RGB) and depth sensor (Depth). † indicates methods without using LLMs.
Figure 3: Deep analysis of evidence-seeking behavior. Left: navigation success and path efficiency under three trigger strategies, highlighting the effectiveness of adaptive evidence seeking. Center: beneficial action change rate (BACR) during reinforcement fine-tuning, measuring the proportion of beneficial action changes. Right: mean short-horizon progress gain of the evidence-seeking branch over the direct-navigation branch, quantifying the navigation benefit of acquired evidence.
Figure 4: Evidence seeking in simulated navigation. Two SEEK decisions along a representative trajectory. SeekVLN first identifies the dining area on the left; after completing that subgoal, it locates the hallway beyond the kitchen island on the right. In both cases, evidence seeking resolves uncertainty about instruction progress and grounds the next navigation action.
Figure 5: Evidence seeking in real-world navigation. At the end of the hallway, the destination is not visible ahead and the instruction leaves the turn direction unspecified. SeekVLN seeks task-relevant evidence, identifies the chair to its left, and completes the task.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Expert action
SEEK proportion
Move forward
30%
Turn left/right ( 15∘ )
30%
Turn left/right ( 30∘ )
50%
Turn left/right ( 45∘ )
75%
Turn left/right (other angles)
100%
Stop
0%
Appendix
Table 2: Target SEEK proportions by expert action. For turns, the angle denotes the cumulative rotation of consecutive turns in the same direction.
Component
Setting
Base model
Pretrained Aux-Think ( Wang et al., 2025a )
Trainable modules
Language model and multimodal projector; visual encoder frozen
Maximum sequence length
512 tokens
Learning rate
2×10−5
Special tokens
<nav> , <seek> , </seek> , <think> , </think>
Observations input
Current observation and up to eight sampled historical frames; left, front, and right views are added in SEEK mode
Appendix
Table 3: FRG SFT training configuration.
Component
Setting
Model and data
Initialization
SeekVLN-FRG-SFT
Trainable modules
Language model and multimodal projector; visual encoder frozen
Value function
Separate model initialized from the SFT checkpoint
Visual context
Nine frames per decision
Training data
640 episodes sampled from the R2R-CE train split
Appendix
Table 4: C2PO RFT training configuration.
R2R-CE Val-Unseen-613
NE ↓
SR ↑
SPL ↑
Aux-Think (base model)
5.80
52.0
45.0
w/o Evidence Seeking
4.86
58.6
53.3
w/o Progress Reasoning
4.99
58.7
53.0
Full FRG
4.84
60.7
55.8
Appendix
Table 5: FRG ablations on the R2R-CE Val-Unseen-613 subset.
R2R-CE Val-Unseen-613
NE ↓
SR ↑
SPL ↑
Aux-Think (base model)
5.80
52.0
45.0
w/o Evidence Seeking
4.86
58.6
53.3
w/o Progress Reasoning
4.99
58.7
53.0
Full FRG
4.84
60.7
55.8
Appendix
Table 5: FRG ablations on the R2R-CE Val-Unseen-613 subset.
R2R-CE Val-Unseen-613
NE ↓
SR ↑
SPL ↑
Seek Rate (%)
SeekVLN-FRG-SFT
4.84
60.7
55.8
15.9
w/o CF Reward
3.94
65.1
60.4
34.3
Full C2PO
3.66
68.5
62.4
29.3
Appendix
Table 6: C2PO ablations on the R2R-CE Val-Unseen-613 subset.
Figure 6: Real-world SeekVLN deployment. The Go2 robot handles sensing and motion, while a remote GPU server performs policy inference. The hardware setup is shown on the right.
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to visual evidence and high-level decisions lack precise spatial goals to guide low-level motion. Cognitive science suggests that human navigation bridges these levels hierarchically by anchoring cognition to relevant landmarks and guiding locomotion toward spatial goals. Motivated by this principle, we propose GroundingVLN, which uses visual grounding as a shared interface between reasoning and action. GroundingVLN first reasons with grounding by anchoring task-relevant visual evidence to precise image locations throughout structured reasoning. It then acts through grounding by predicting a progress-aligned pixel goal that a geometric planner translates into primitive actions. To learn these capabilities, we construct GroundingCOTVLN-188K, a dataset of temporally aligned grounded reasoning traces, and introduce Grounded and Execution-Aware Reinforcement Learning (GEAR), which aligns grounded reasoning and spatial decisions with downstream execution. Experiments demonstrate that GroundingVLN achieves state-of-the-art performance (69.9% SR on R2R-CE and 75.1% SR on RxR-CE) with high sample efficiency, using just 0.9% as much training data as the strongest baseline. It also generalizes strongly across datasets, attaining 59.9% SR on RxR-CE when trained solely on R2R, a gain of 20.1% over the strongest baseline. Code and models will be released after review.
Kailing Li, Yu Han, Tianwen Qian +4
School of Computer Science and Technology, East China Normal University · NeoteAI · King Abdullah University of Science and Technology
Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3% SR and 51.4% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3% SR without additional navigation training data or a geometry encoder at inference.
Anh Dao, Quan-Dung Pham, Le Danh Vinh +6
VinMotion, Inc., Vietnam · University of Southern California, USA