Deploying vision-and-language navigation (VLN) agents in unseen environments remains challenging because unfamiliar layouts and visual conditions can cause execution to go off track. Rather than relying on continuous human supervision, a practical strategy is to selectively request corrective guidance, recover the ongoing task, and reuse corrective interactions to improve subsequent navigation. We propose Abstention-aware Visual Error Recovery and Training for Vision-and-Language Navigation (AVERT-VLN), a closed-loop framework that uses a plug-in vision-language Monitor for online human-assisted recovery and offline preference learning. The Monitor operates separately from navigation decision generation and assesses instruction-execution consistency from the instruction, visual history, and current observation. To train the Monitor for deviation recognition, we construct LOSTNAV DATASET with 20K counterfactual risk trajectories and rule-based deviation labels. The Monitor is first fine-tuned on 40K normal trajectories to assess instruction progress and then jointly fine-tuned on normal and risk trajectories to recognize semantic deviations. At runtime, Asynchronous Sidecar Monitoring evaluates execution alongside the navigation model. When the controller accepts a LOST verdict, it suspends autonomous execution and requests human guidance for recovery. For offline policy improvement, Trajectory-Anchored Preference Learning converts deviation-associated failures into decision-level preference pairs under shared decision contexts, restricting supervision to the decisions targeted for correction. Under human-assisted evaluation, the full AVERT-VLN system achieves success rates of 76.2% and 66.3% on the val-unseen splits of R2R-CE and RxR-CE, respectively. The same monitoring and human-assisted recovery interface also improves success rates across the three evaluated navigation architectures.
Figures & tables
Fig. 1: Conceptual comparison of conventional VLN and AVERT-VLN. AVERT-VLN forms a closed deployment loop in which the Monitor detects off-track execution, human guidance supports recovery when needed, and the resulting corrections are reused for offline policy improvement.
Fig. 2: Online recovery and offline preference learning in AVERT-VLN. (1) Asynchronous Sidecar Monitoring assesses execution alongside the navigation model and triggers abstention and human-assisted recovery upon a Lost verdict. (2) Trajectory-Anchored Preference Learning traces deviations to the nearest preceding Pass boundary, pairs correction-derived preferred with the original output under a shared context, and updates the navigation model offline using decision-focused DPO.
Fig. 3: Construction of LostNav Dataset . Simulator rollouts of validated target-selection interventions yield counterfactual risk trajectories with rule-based deviation labels and diagnostic annotations.
22 25 23 26 24 27-28 Method
Observation
R2R-CE Val-Unseen
RxR-CE Val-Unseen
Pano.
Odo.
Depth
S.RGB
NE ↓
OS ↑
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
nDTW ↑
HPN+DN ∗ [ 26 ]
✓
✓
✓
6.31
40.0
36.0
34.0
–
–
–
–
CMA ∗ [ 27 ]
✓
✓
✓
6.20
52.0
41.0
36.0
8.76
26.5
22.1
47.0
GridMM ∗ [ 28 ]
✓
✓
✓
5.11
61.0
49.0
41.0
–
–
–
–
ETPNav ∗ [ 29 ]
✓
✓
✓
4.71
65.0
57.0
49.0
5.64
54.7
44.8
61.9
ScaleVLN ∗ [ 30 ]
✓
✓
✓
4.80
–
55.0
51.0
–
–
–
–
TABLE I: Navigation Performance on R2R-CE and RxR-CE Val-Unseen Splits.
Fig. 4: Diagnostic reasoning and lost-state recognition. (a) Six-dimensional diagnostic comparison of Monitor-4B and six selected baselines. Annotations report Monitor scores and gaps to the highest displayed baseline score in each dimension. (b) Lost-state detection recall and F1 (%) for all 18 models on the 2,190 states in LostAware Benchmark , with Lost as the positive class. Higher is better for both metrics.
Fig. 5: Qualitative comparison of autonomous navigation before and after DPO. Two selected RxR-CE val-unseen episodes are shown for the original System 2 policy (Original) and its DPO-trained counterpart (+DPO), with online monitoring and human assistance disabled. Each row presents six chronologically ordered observations and the final trajectory map; columns are not temporally aligned across rows. Gold borders mark shared observations; arrows indicate camera motions or navigation actions.
6 Configuration
NE ↓
SR ↑
OS ↑
nDTW ↑
System 2
5.77
52.88
65.30
64.16
System 2 + SFT
5.69
53.41
65.30
64.20
System 2 + DPO
5.76
55.03
65.56
63.21
System 2 + Monitor
4.53
57.78
68.71
57.25
System 2 + DPO + Monitor
4.67
60.21
70.29
57.13
TABLE II: Component Ablation on RxR-CE Val-Unseen.
Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN
Vision-Language Navigation (VLN) requires agents to continuously ground task progress from long-horizon instructions and partial egocentric observations. Existing VLM-based navigation agents typically reason only over available observations and may remain confident even when task-relevant evidence is missing. For example, an agent may confidently proceed forward and get lost even though the landmark indicating the next turn lies outside its current field of view. We term this failure mode Progress Myopia: the agent fails to recognize unreliable progress grounding and continues acting on insufficient evidence. To address it, we propose SeekVLN, an evidence-seeking framework that couples semantic progress reasoning with active acquisition of task-relevant observations. SeekVLN is trained in two stages: First, Future-guided Reverse Generation (FRG) uses future expert actions to augment offline expert trajectories with supplementary views and evidence annotations. Supervised fine-tuning on these trajectories establishes a prior for evidence seeking and progress reasoning without additional expert interaction. However, imitation alone does not reveal whether seeking improves subsequent navigation. We therefore introduce Counterfactual Contrastive Policy Optimization (C2PO) for reinforcement fine-tuning. By comparing each evidence-seeking branch with a counterfactual direct-navigation branch from the same state, C2PO uses a contrastive reward to assign credit to seeking decisions based on subsequent navigation benefit. Experiments on simulated benchmarks show that SeekVLN achieves state-of-the-art performance, improving success rate by 12.7% and 7.5% over the base model on R2R-CE and RxR-CE, respectively. Both simulated and real-world evaluations exhibit human-like evidence-seeking behaviors for more reliable progress grounding.
Zhimin Wang, Meiyuan Zhu, Duo Wu +8
Tsinghua University · Pengcheng Laboratory · South China University of Technology +3
Vision-and-Language Navigation (VLN) necessitates an embodied agent to navigate in the physical world by adhering to natural language instructions. Recent advancements in Vision-Language Models (VLM) have propelled the development of VLM-based VLN methods with two predominant paradigms: (1) imitation learning (IL) on expert demonstrations, followed by the Dataset Aggregation (DAgger) algorithm to bolster error recovery capabilities; (2) reinforcement learning (RL) driven by verifiable rewards to enhance reasoning and exploration. A notable gap is the absence of integration between these two distinct paradigms. This paper introduces JOP-VLN, a novel VLN framework that synergistically combines off-policy imitation learning and on-policy exploration within a three-stage training pipeline. Initially, IL is employed on expert demonstrations to acquire basic navigation skills. Subsequently, the DAgger algorithm is utilized to generate heuristic exploration trajectories, which are then used for imitation learning to improve error recovery capabilities. Finally, a joint on-and-off policy learning framework is implemented, featuring high-entropy trajectory sampling to enhance RL training efficiency and an error-correction-prioritized trajectory sorting strategy for effective error correction. Extensive experiments demonstrate the efficacy of JOP-VLN, achieving success rates of 69.9% and 68.0% on the VLN-CE R2R and RxR benchmarks, respectively, setting a new state-of-the-art on R2R. Project page: https://qingrongh.github.io/JOP-VLN.