Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation
Authors: Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, +3 more
Organizations: Tsinghua University · Pengcheng Laboratory · South China University of Technology · International Digital Economy Academy · Ping An Technology (Shenzhen) Co., Ltd., Shenzhen, China · Harbin Institute of Technology (Shenzhen)
Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.
Figures & tables
Figure 1: Illustration of Progress Myopia . Left: out-of-view landmarks (a) and ambiguous instructions (b) make the current observation insufficient to determine the correct decision. Right: representative navigation agents show similar decision confidence and action entropy on both successful and failed segments, indicating that unreliable decisions can remain highly confident.
Figure 2: Overview of SeekVLN . (1) The policy decides whether to navigate or seek evidence. (2) FRG reasons backward from future expert actions to determine where to seek evidence and annotate key evidence, while using historical context for progress annotation. (3) C2PO compares evidence-seeking and direct-navigation branches and jointly optimizes evidence-seeking and navigation.
Method
Observation
R2R-CE Val-Unseen
RxR-CE Val-Unseen
S.RGB
Depth
NE ↓
OSR ↑
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
nDTW ↑
BEVBert † An et al. (2022)
✓
4.57
67.0
59.0
50.0
–
–
–
–
ETPNav † An et al. (2024)
✓
4.71
65.0
57.0
49.0
5.64
54.8
44.9
61.9
ENP-ETPNav † Liu et al. (2024)
✓
4.69
65.0
58.0
50.0
5.51
55.3
45.1
63.0
Seq2Seq † Krantz et al. (2020)
✓
✓
7.77
37.0
25.0
22.0
12.10
13.9
11.9
–
CMA † Krantz et al. (2020)
✓
✓
7.37
40.0
32.0
30.0
–
–
–
–
Table 1: Comparison of different methods on the R2R-CE and RxR-CE Val-Unseen splits. Observations include a single RGB camera (S.RGB) and depth sensor (Depth). † indicates methods without using LLMs.
Figure 3: Deep analysis of evidence-seeking behavior. Left: navigation success and path efficiency under three trigger strategies, highlighting the effectiveness of adaptive evidence seeking. Center: beneficial action change rate (BACR) during reinforcement fine-tuning, measuring the proportion of beneficial action changes. Right: mean short-horizon progress gain of the evidence-seeking branch over the direct-navigation branch, quantifying the navigation benefit of acquired evidence.
Figure 4: Evidence seeking in simulated navigation. Two SEEK decisions along a representative trajectory. SeekVLN first identifies the dining area on the left; after completing that subgoal, it locates the hallway beyond the kitchen island on the right. In both cases, evidence seeking resolves uncertainty about instruction progress and grounds the next navigation action.
Figure 5: Evidence seeking in real-world navigation. At the end of the hallway, the destination is not visible ahead and the instruction leaves the turn direction unspecified. SeekVLN seeks task-relevant evidence, identifies the chair to its left, and completes the task.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Expert action
SEEK proportion
Move forward
30%
Turn left/right ( 15∘ )
30%
Turn left/right ( 30∘ )
50%
Turn left/right ( 45∘ )
75%
Turn left/right (other angles)
100%
Stop
0%
Appendix
Table 2: Target SEEK proportions by expert action. For turns, the angle denotes the cumulative rotation of consecutive turns in the same direction.
Component
Setting
Base model
Pretrained Aux-Think ( Wang et al., 2025a )
Trainable modules
Language model and multimodal projector; visual encoder frozen
Maximum sequence length
512 tokens
Learning rate
2×10−5
Special tokens
<nav> , <seek> , </seek> , <think> , </think>
Observations input
Current observation and up to eight sampled historical frames; left, front, and right views are added in SEEK mode
Appendix
Table 3: FRG SFT training configuration.
Component
Setting
Model and data
Initialization
SeekVLN-FRG-SFT
Trainable modules
Language model and multimodal projector; visual encoder frozen
Value function
Separate model initialized from the SFT checkpoint
Visual context
Nine frames per decision
Training data
640 episodes sampled from the R2R-CE train split
Appendix
Table 4: C2PO RFT training configuration.
R2R-CE Val-Unseen-613
NE ↓
SR ↑
SPL ↑
Aux-Think (base model)
5.80
52.0
45.0
w/o Evidence Seeking
4.86
58.6
53.3
w/o Progress Reasoning
4.99
58.7
53.0
Full FRG
4.84
60.7
55.8
Appendix
Table 5: FRG ablations on the R2R-CE Val-Unseen-613 subset.
R2R-CE Val-Unseen-613
NE ↓
SR ↑
SPL ↑
Aux-Think (base model)
5.80
52.0
45.0
w/o Evidence Seeking
4.86
58.6
53.3
w/o Progress Reasoning
4.99
58.7
53.0
Full FRG
4.84
60.7
55.8
Appendix
Table 5: FRG ablations on the R2R-CE Val-Unseen-613 subset.
R2R-CE Val-Unseen-613
NE ↓
SR ↑
SPL ↑
Seek Rate (%)
SeekVLN-FRG-SFT
4.84
60.7
55.8
15.9
w/o CF Reward
3.94
65.1
60.4
34.3
Full C2PO
3.66
68.5
62.4
29.3
Appendix
Table 6: C2PO ablations on the R2R-CE Val-Unseen-613 subset.
Figure 6: Real-world SeekVLN deployment. The Go2 robot handles sensing and motion, while a remote GPU server performs policy inference. The hardware setup is shown on the right.