Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.
Figures & tables
Figure 1 : Performance and efficiency advantages of NavJev on R2R-CE. NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step and substantially lower inference cost than representative MLLM-based VLN methods.
Figure 2 : Overview of NavJev. ACVC compresses candidate observations into compact action-centric representations, DASM filters shared semantics and maintains discriminative action-specific evidence across navigation steps, and Jev performs efficient typed action selection.
Figure 3 : Inference efficiency analysis of NavJev. (a) Average step time compared with representative zero-shot VLN methods. NavJev requires only 0.65 s per step, achieving approximately 41.5 × , 8.2 × , and 7.6 × speedups over Open-Nav, SmartWay, and P2DNav, respectively. (b) Wall-clock latency breakdown of NavJev. Jev API calls account for most of the inference time, while waypoint prediction, ACVC perception, and other modules introduce only limited additional overhead.
Method
Decision Model
NE ↓
OSR ↑
SR ↑
SPL ↑
Supervised Learning
Seq2Seq [ 17 ]
–
7.77
37.0
25.0
22.0
MEE [ 14 ]
–
6.82
44.6
35.9
32.3
NaVid [ 34 ]
Vicuna-7B
5.47
49.1
37.4
35.9
MLANet [ 15 ]
–
6.30
42.0
38.0
35.0
Uni-NaVid [ 33 ]
Vicuna-7B
5.58
53.3
47.0
42.7
Table 1 : Comparison with representative supervised and zero-shot methods on the R2R-CE val-unseen split.
Method
Average Step Time (s) ↓
Average Episode Time (s) ↓
Average Steps per Episode
Peak GPU Memory (GiB) ↓
NaVid
0.28
24.60
89.2
17.95
NaVILA
0.21
24.30
118.7
17.46
StreamVLN
0.23
16.69
73.6
23.63
JanusVLN
0.64
44.22
68.9
35.80
NavJev
0.65
6.03
9.3
4.46
Table 2 : Inference efficiency comparison with representative task-fine-tuned MLLM-based VLN methods. These methods typically predict low-level discrete actions at each step, whereas NavJev performs waypoint-level action selection.
Decision Model
SR ↑
SPL ↑
OSR ↑
nDTW ↑
Decision Latency (s) ↓
Average Step Time (s) ↓
Total Cost ↓
GPT-4o
14.0
8.5
42.0
23.8
1.54 ± 0.42
1.66 ± 0.42
$3.67
Qwen3.8-Max
25.0
16.0
46.0
34.2
1.03 ± 0.25
1.15 ± 0.25
$2.19
NavJev
27.0
22.4
35.0
43.5
0.53 ± 0.11
0.65 ± 0.11
$0.06
Table 3 : Comparison of decision models under the same ACVC–DASM representation on the R2R-CE 100-episode split. Qwen3.8-Max and GPT-4o are evaluated with reasoning disabled.
Figure 4 : Distribution of per-episode average step time (AST) for different decision models under the same ACVC–DASM representation on the R2R-CE 100-episode split.
Figure 5 : Failure analysis of NavJev on the R2R-CE 100-episode split. Left: distribution of navigation outcomes and failure modes, where incorrect waypoint selection accounts for most failures. Middle and right: representative failure cases comparing the waypoint selected by NavJev (red) with the oracle waypoint (green). When instructions require precise spatial-relation understanding or different viewpoints exhibit similar visual semantics, NavJev may select a plausible but incorrect direction.
Figure 6 : Real-world navigation visualization of NavJev. The agent follows natural-language instructions through sequential waypoint decisions and successfully stops at the referred vending machine and elevator. Red annotations indicate key selected actions and their corresponding action-semantic cues.
BLIP
RAM
DASM
TL ↓
NE ↓
OSR ↑
SR ↑
SPL ↑
nDTW ↑
✓
✗
✗
12.10
8.80
33.0
19.0
14.4
36.2
✓
✓
✗
15.83
8.13
34.0
21.0
16.1
35.0
✓
✓
✓
11.58
7.48
35.0
27.0
22.4
43.5
Table 4 : Component ablation of ACVC and DASM on the R2R-CE 100-episode split, where BLIP and RAM are components of ACVC.
Decision Model
Scene 1 (Office)
Scene 2 (Café)
Average
Decision Latency (s) ↓
Average Step Time (s) ↓
OSR ↑
SR ↑
OSR ↑
SR ↑
OSR ↑
SR ↑
Qwen3.8-Max
30.0
30.0
40.0
40.0
35.0
35.0
1.10 ± 0.30
1.23 ± 0.31
NavJev
50.0
40.0
70.0
60.0
60.0
50.0
0.65 ± 0.51
0.79 ± 0.51
Table 5 : Real-world evaluation of NavJev in office and café environments. We compare NavJev with Qwen3.8-Max under the same waypoint candidates, ACVC perception, and DASM representation. OSR and SR are computed over 10 tasks per scene and averaged over all 20 tasks.
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to maintain long-horizon visual history for trajectory consistency while executing actions with low latency. Existing video-based VLN approaches typically struggle to satisfy both demands simultaneously. To address these challenges, we propose MemVLN, a novel VLN framework that achieves state-of-the-art performance with real-time inference efficiency (14 FPS). MemVLN utilizes a visual encoder to process continuous observations and a Large Language Model (LLM) to interpret instructions and generate actions. Central to our approach is an Episodic Memory management that applies pyramidal resolutions. This mechanism concentrates computation on immediate percepts while retaining compressed long-term history. Complementing to this design, we introduce Procedural Memory for fast action with a compact vocabulary of atomic mid-level actions to bypass auto-regressive decoding latency. Experiments on VLN-CE show that MemVLN-4B surpasses the baseline Qwen3-VL-4B architecture by 5.8% SR in R2R and 9.7% SR in RxR, while achieving a 7× speedup in inference latency.
Yuqi Liu, Shengju Qian, Tianyuan Qu +5
1The Chinese University of Hong Kong · 2LIGHTSPEED · 3The University of Hong Kong +1
Vision-Language Navigation (VLN) requires embodied agents to generate actions based on instructions and observations. General-purpose multimodal agents offer a promising basis for this task, but selecting plausible local actions does not ensure that execution remains consistent with the intended route, particularly in long-horizon tasks. Moreover, the accumulated interaction history increases the input required for subsequent decisions, resulting in a significant inference overhead. To this end, we introduce \method, an Agentic VLN framework that includes a Goal Agent that sets adaptive goals for local actions, a Verify Agent that dynamically verifies whether a goal has been completed, a Memory Agent for multimodal context compression, and a Visuomotor Agent to execute adaptive goals. Specifically, the Goal Agent formulates adaptive goals based on the instruction, current observation, and execution history. Then the Visuomotor Agent executes navigation actions to achieve each goal, while the Verify Agent uses a goal-specific verification question to dynamically assess whether the observed outcomes satisfy the intended completion condition. Verified goal completion then marks a boundary for the Memory Agent to compress the corresponding multimodal interaction history while preserving information needed for subsequent navigation. We evaluate navigation on R2R-CE and RxR-CE, examine framework variants across three model backbones, and study context evolution during execution. For Real-World evaluation, \method achieves 83.3% success and 1.51,m navigation error across eight challenging routes evaluated three times each.
Haoxiang Shi, Zaijing Li, Muhe Ding +3
Harbin Institute of Technology (Shenzhen) · Pengcheng Laboratory
While Multimodal Large Language Models (MLLMs) have demonstrated significant promise in Vision-Language Navigation (VLN), existing agents remain heavily constrained by systemic bottlenecks across inference, training, and data collection. Specifically, they suffer from prohibitive latency due to visual history reprocessing, action leakage during sequence-packed training, and suboptimal exploration in self-correction data collection. To overcome these intertwined challenges, we present Efficient-VLN, a highly efficient and robust baseline that systematically resolves these issues through three simple-yet-effective mechanisms. (1) Inference: We introduce KV-cache reuse with contiguous RoPE, enabling the model to process only the newly observed frame at each step for real-time inference. (2) Training: We propose packed training with an action-isolating mask to accelerate throughput while effectively bridging the training-inference gap by preventing action leakage. (3) Data Collection: We employ an Adaptive DAgger to dynamically balance autonomous exploration and oracle guidance, enhancing error-recovery capability without escalating computational costs. Extensive evaluations show that Efficient-VLN significantly advances the state-of-the-art across the R2R-CE (73.2% SR) and RxR-CE (75.6% SR) benchmarks. Meanwhile, it yields a 28% latency reduction compared to the previous state-of-the-art StreamVLN, establishing a new paradigm for streaming MLLM-based navigation.