Recent zero-shot Vision-and-Language Navigation (VLN) methods increasingly rely on multimodal large language models (MLLMs) to reason over visual observations, navigation instructions, and candidate actions. Although effective, repeatedly invoking autoregressive multimodal reasoning at every navigation step introduces substantial inference latency, limiting the responsiveness of embodied agents. We propose NavJev, an efficient VLN framework that reformulates online navigation from repeated multimodal generation into compact visual compression followed by lightweight typed action selection. Specifically, Action-Centric Visual Compression (ACVC) integrates waypoint geometry, BLIP captions, and RAM semantic tags into compact representations of candidate actions, while Discriminative Action-Semantic Memory (DASM) filters shared semantics and maintains discriminative action-specific evidence across navigation steps. Based on these representations, Jev directly performs structured probabilistic decisions over the available action set. Experiments on R2R-CE show that NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step, while substantially reducing inference latency and cost compared with MLLM-based VLN methods. The project page is available at https://kai-sheng-caesar.github.io/NavJev/.
Figures & tables
Figure 1 : Performance and efficiency advantages of NavJev on R2R-CE. NavJev achieves 27.0% SR and 22.4% SPL with only 0.65 s per navigation step and substantially lower inference cost than representative MLLM-based VLN methods.
Figure 2 : Overview of NavJev. ACVC compresses candidate observations into compact action-centric representations, DASM filters shared semantics and maintains discriminative action-specific evidence across navigation steps, and Jev performs efficient typed action selection.
Figure 3 : Inference efficiency analysis of NavJev. (a) Average step time compared with representative zero-shot VLN methods. NavJev requires only 0.65 s per step, achieving approximately 41.5 × , 8.2 × , and 7.6 × speedups over Open-Nav, SmartWay, and P2DNav, respectively. (b) Wall-clock latency breakdown of NavJev. Jev API calls account for most of the inference time, while waypoint prediction, ACVC perception, and other modules introduce only limited additional overhead.
Method
Decision Model
NE ↓
OSR ↑
SR ↑
SPL ↑
Supervised Learning
Seq2Seq [ 17 ]
–
7.77
37.0
25.0
22.0
MEE [ 14 ]
–
6.82
44.6
35.9
32.3
NaVid [ 34 ]
Vicuna-7B
5.47
49.1
37.4
35.9
MLANet [ 15 ]
–
6.30
42.0
38.0
35.0
Uni-NaVid [ 33 ]
Vicuna-7B
5.58
53.3
47.0
42.7
Table 1 : Comparison with representative supervised and zero-shot methods on the R2R-CE val-unseen split.
Method
Average Step Time (s) ↓
Average Episode Time (s) ↓
Average Steps per Episode
Peak GPU Memory (GiB) ↓
NaVid
0.28
24.60
89.2
17.95
NaVILA
0.21
24.30
118.7
17.46
StreamVLN
0.23
16.69
73.6
23.63
JanusVLN
0.64
44.22
68.9
35.80
NavJev
0.65
6.03
9.3
4.46
Table 2 : Inference efficiency comparison with representative task-fine-tuned MLLM-based VLN methods. These methods typically predict low-level discrete actions at each step, whereas NavJev performs waypoint-level action selection.
Decision Model
SR ↑
SPL ↑
OSR ↑
nDTW ↑
Decision Latency (s) ↓
Average Step Time (s) ↓
Total Cost ↓
GPT-4o
14.0
8.5
42.0
23.8
1.54 ± 0.42
1.66 ± 0.42
$3.67
Qwen3.8-Max
25.0
16.0
46.0
34.2
1.03 ± 0.25
1.15 ± 0.25
$2.19
NavJev
27.0
22.4
35.0
43.5
0.53 ± 0.11
0.65 ± 0.11
$0.06
Table 3 : Comparison of decision models under the same ACVC–DASM representation on the R2R-CE 100-episode split. Qwen3.8-Max and GPT-4o are evaluated with reasoning disabled.
Figure 4 : Distribution of per-episode average step time (AST) for different decision models under the same ACVC–DASM representation on the R2R-CE 100-episode split.
Figure 5 : Failure analysis of NavJev on the R2R-CE 100-episode split. Left: distribution of navigation outcomes and failure modes, where incorrect waypoint selection accounts for most failures. Middle and right: representative failure cases comparing the waypoint selected by NavJev (red) with the oracle waypoint (green). When instructions require precise spatial-relation understanding or different viewpoints exhibit similar visual semantics, NavJev may select a plausible but incorrect direction.
Figure 6 : Real-world navigation visualization of NavJev. The agent follows natural-language instructions through sequential waypoint decisions and successfully stops at the referred vending machine and elevator. Red annotations indicate key selected actions and their corresponding action-semantic cues.
BLIP
RAM
DASM
TL ↓
NE ↓
OSR ↑
SR ↑
SPL ↑
nDTW ↑
✓
✗
✗
12.10
8.80
33.0
19.0
14.4
36.2
✓
✓
✗
15.83
8.13
34.0
21.0
16.1
35.0
✓
✓
✓
11.58
7.48
35.0
27.0
22.4
43.5
Table 4 : Component ablation of ACVC and DASM on the R2R-CE 100-episode split, where BLIP and RAM are components of ACVC.
Decision Model
Scene 1 (Office)
Scene 2 (Café)
Average
Decision Latency (s) ↓
Average Step Time (s) ↓
OSR ↑
SR ↑
OSR ↑
SR ↑
OSR ↑
SR ↑
Qwen3.8-Max
30.0
30.0
40.0
40.0
35.0
35.0
1.10 ± 0.30
1.23 ± 0.31
NavJev
50.0
40.0
70.0
60.0
60.0
50.0
0.65 ± 0.51
0.79 ± 0.51
Table 5 : Real-world evaluation of NavJev in office and café environments. We compare NavJev with Qwen3.8-Max under the same waypoint candidates, ACVC perception, and DASM representation. OSR and SR are computed over 10 tasks per scene and averaged over all 20 tasks.