Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3% success and 1.51,m navigation error across eight challenging routes evaluated three times each.
Figures & tables
Figure 1: NavHarness connects adaptive goal execution, verified progress, and multimodal memory. Adaptive local goals guide navigation, while goal-specific questions verify whether observed outcomes meet completion conditions. Verified completion marks boundaries for summarizing completed interactions while retaining information for ongoing navigation and recovery. NavHarness significantly outperforms trained policies and other agent frameworks in VLN-CE.
Figure 2: NavHarness framework. Adaptive goals guide execution; observation-based verification provides feedback for the next step and goal revision or advancement. Verified completion triggers multimodal memory compression. The two situations illustrate route following and recovery.
R2R-CE
RxR-CE
Method / backbone
Views
Training
NE ↓
OSR ↑
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
nDTW ↑
Trained navigation models
NaVid ( Zhang et al., 2024 )
M
Trained
5.47
49.1
37.4
35.9
8.41
23.8
21.2
—
NaVILA ( Cheng et al., 2025 )
M
Trained
5.22
62.5
54.0
49.0
6.77
49.3
44.0
58.8
StreamVLN ( Wei et al., 2026 )
M
Trained
4.90
63.6
56.4
50.2
5.65
54.4
45.4
63.7
JanusVLN ( Zeng et al., 2026 )
M
Trained
4.78
65.2
60.5
56.8
6.06
56.2
47.5
62.1
Table 1: Navigation performance on R2R-CE and RxR-CE. Bold : best reported value per metric; underline : best baseline within each group unless bold. Shaded rows denote NavHarness. M = monocular, P = panorama, D = depth; 3-cam = three fixed RGB cameras. NE is in meters; other metrics are percentages. Results retain their original protocols: SmartWay and AgenticNav use Open-Nav's 100 episodes; Dashes denote unreported or pending results.
Navigation Quality
Token Consumption
Decision Context
Compression
Model
Harness
Compress.
SR ↑
SPL ↑
Cache
Input ↓
Total ↓
Image/Dec. ↓
Text/Dec. (k) ↓
Total ratio (%) ↓
R2R-CE
GPT-5.6-Sol
Codex CLI
Native
62.6
44.5
1351.40
1398.72
1410.31
12.43
68.99
—
NavHarness
Off
67.0
47.9
1234.50
1334.64
1345.23
13.10
80.57
—
NavHarness
Goal
66.0
50.3
770.45
851.85
861.82
11.80
32.31
64.06%
GPT-6-Astra
Codex CLI
Native
74.0
62.5
656.30
687.77
689.81
9.42
73.72
—
Table 2: Navigation quality and context consumption. Cache/Input/Total: 103 tokens per episode. Cache is included in Input. Image/Dec.: mean input images per external decision. Text/Dec.: mean input text length ( 103 characters) per external decision. Total ratio: Goal/Off (%).
Figure 3: Input context during navigation. Vertical markers indicate compression events; percentages report selected local input reductions, not episode-level savings.
Figure 4: Framework ablation. R2R-CE SR (%), rounded. Plain NavHarness is Off.
Method
SR (%) ↑
NE (m) ↓
NaVILA
20.83
3.91
AwareVLN
20.83
3.89
StreamVLN
25.00
3.78
JanusVLN
50.00
2.34
NavHarness
83.30
1.51
Table 3: Real-world navigation. Baselines use the supplied aggregates; NavHarness uses eight routes with three trials each.
Figure 5: Indoor scene transitions. Following ordered landmarks, clearing the doorway before turning, and approaching the yellow chair with a white ball.
Figure 6: Adaptive outdoor navigation. Obstacle-responsive goal adaptation toward the sculpture (top) and landmark-conditioned progress toward the trees (bottom).
Appendix figures & tables1 asset
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Successive landmark turns. The route links a doorway, a yellow chair, a white table, and a black table through successive landmark-conditioned turns. Panel (h) shows a clear view during the final approach; the run subsequently calls STOP.
Existing vision-language navigation methods often couple a VLM with waypoint decoders to produce multi-step action plans, but they typically lack an explicit closed-loop mechanism for tracking semantic progress, diagnosing execution failures, and recovering from error accumulation in long-horizon navigation. To address this gap, we propose ReflectVLN, an agentic VLN framework that organizes decision-making through bidirectionally interactive intention and execution agents. The intention agent performs subtask decomposition and reflection, generating executable subtask descriptions as corrective plans. Conditioned on these descriptions, the execution agent grounds them into short-horizon actions under current observations while monitoring sub-goal progress and detecting off-track behavior. Crucially, ReflectVLN enables closed-loop bidirectional communication: the execution agent emits progress and deviation signals to trigger reflection and subtask updates on demand, and the intention agent returns structured guidance that reconditions subsequent actions for recovery. To encourage temporally coherent decisions with interpretable intermediate rationales, we introduce Action Chain-of-Thought (Action-CoT), a path-conditioned dual-query training scheme for action generation. Experiments on standard VLN benchmarks show that ReflectVLN improves success rates and path efficiency under a constrained data budget, with favorable training cost and fewer high-level intention calls at inference time, while providing interpretable intermediate decisions for analysis and collaboration. Code is available at: https://github.com/AIprogrammer/ReflectVLN
Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.
Kai Sheng, Liuyi Wang, Jinlong Li +3
College of Electronic and Information Engineering, Tongji University, Shanghai, China
Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories to directly predict actions, and the zero-shot modular pipeline integrating pre-trained Multimodal Large Language Model (MLLM) for training-free generalization to unseen environments. However, end-to-end methods struggle with long-horizon navigation and lack dynamic reasoning, whereas zero-shot methods are constrained by limited spatial grounding for reliable planning and also require substantial reasoning time. To bridge this gap, we introduce SEDualVLN, a spatially-enhanced dual-system VLN framework. System 1 is a VLM model enhanced with both global and local spatial awareness, used for action generation. System 2 integrates a general MLLM with a mapping module, wherein the MLLM plans waypoints by leveraging top-down views of the real-time 3D map alongside streams of rendered path images. Both systems leverage different forms of spatial enhancement to cultivate the agent's sense of direction in VLN tasks. Ultimately, they cooperate to complete the navigation task through a fast-slow coordinated approach. SEDualVLN achieves state-of-the-art performance on VLN-CE benchmarks, and further ablation studies demonstrate the effectiveness of each system and module.
Jingzhi Huang, Junkai Huang, Wenxuan Song +4
Hong Kong Polytechnic University · Institute of Automation, Chinese Academy of Sciences · Hong Kong University of Science and Technology (Guangzhou)