Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3% SR and 51.4% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3% SR without additional navigation training data or a geometry encoder at inference.
Figures & tables
Fig. 1 : StageVLN overview. Left: comparison with explicit 3D inputs and runtime geometry fusion; StageVLN uses spatial and trajectory guidance only during training and retains the navigator alone at inference. Right: navigation success and inference efficiency. VRAM and latency are measured on a single NVIDIA RTX 6000 Pro.
Fig. 2 : StageVLN training and inference pipeline. During training, a frozen geometry teacher provides multi-level spatial guidance, while relative-heading and expert-route progress objectives supervise trajectory state. At inference, the teacher, alignment projectors, and auxiliary heads are excluded, leaving only the base RGB–language navigation policy.
Method
Model Params
Observation
R2R Val-Unseen
RxR Val-Unseen
External Training Data
Pano.
Odo.
Depth
S.RGB
NE ↓
OS ↑
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
nDTW ↑
HPN+DN [ 21 ]
–
✓
✓
✓
6.31
40.0
36.0
34.0
–
–
–
–
–
CMA [ 22 ]
–
✓
✓
✓
6.20
52.0
41.0
36.0
8.76
26.5
22.1
47.0
–
Sim2Sim [ 23 ]
–
✓
✓
✓
6.07
52.0
43.0
36.0
–
–
–
–
–
VLN-BERT [ 24 ]
–
✓
✓
✓
5.74
53.0
44.0
39.0
8.98
27.0
22.6
46.7
–
Ego2-Map [ 25 ]
–
✓
✓
✓
5.54
56.0
47.0
41.0
–
–
–
–
–
TABLE I : Comparison on the R2R-CE and RxR-CE validation-unseen splits. ↓ / ↑ indicate lower/higher is better. Best and second-best results are shown in bold and underline , respectively. StreamVLN ∗ uses EnvDrop augmentation, and NaVILA ∗ excludes human-following data. † denotes methods with an additional geometry encoder at inference; parameter counts are reported as VLM + geometry encoder. StageVLN uses no additional navigation training data beyond R2R-CE and RxR-CE; pretrained model data are excluded from the external-data counts.
Method
Peak allocated VRAM (GiB) ↓
Mean step time (ms) ↓
RTX 5090
JanusVLN
OOM
–
NaVILA
17.47
365.2
StreamVLN
17.48
233.4
StageVLN
10.24
304.5
RTX 6000 Pro
TABLE II : Inference memory and mean step time, grouped by GPU. StageVLN is shaded; bold values indicate the best result within each group. OOM denotes an out-of-memory failure under the tested configuration.
Fig. 3 : Simulation rollouts and real-world deployment. (a) Unitree G1 platform with a ZED X Mini camera and Jetson AGX Orin. (b) Simulated episode comparing JanusVLN and StageVLN under the same instruction and initial condition. (c) Two real-world deployment sequences, shown in chronological order.
Nav.
Geo.
Head.
Prog.
NE ↓
OS ↑
SR ↑
SPL ↑
✓
–
5.19
56.01
46.49
42.32
✓
1L
6.33
61.83
48.50
41.86
✓
ML
6.12
61.01
52.26
46.72
✓
ML
✓
6.17
63.30
49.97
43.80
✓
ML
✓
5.66
63.89
53.18
48.67
✓
–
✓
✓
5.86
64.06
54.43
48.58
TABLE III : Component ablation on R2R-CE validation-unseen. VGGT is used as the geometry teacher for all variants. Nav. : navigation supervision; Geo. : geometric representation distillation; Head. : relative heading; Prog. : expert-route progress. 1L uses only the final alignment pair (navigator layer 24, teacher index 23), while ML uses all three alignment pairs. The first row is supervised fine-tuning (SFT); the shaded row is full StageVLN. Best results are shown in bold .
Geometry teacher
NE ↓
OS ↑
SR ↑
SPL ↑
VGGT
5.58
67.37
55.36
48.44
VGGT- Ω
5.17
64.71
56.28
51.37
TABLE IV : Geometry-teacher ablation on R2R-CE validation-unseen. Each variant uses the full StageVLN objective with a frozen teacher. Best results are shown in bold
Vision-Language Navigation (VLN) approaches have currently followed two primary paradigms: the end-to-end Vision-Language Model (VLM) policy fine-tuned on navigation trajectories to directly predict actions, and the zero-shot modular pipeline integrating pre-trained Multimodal Large Language Model (MLLM) for training-free generalization to unseen environments. However, end-to-end methods struggle with long-horizon navigation and lack dynamic reasoning, whereas zero-shot methods are constrained by limited spatial grounding for reliable planning and also require substantial reasoning time. To bridge this gap, we introduce SEDualVLN, a spatially-enhanced dual-system VLN framework. System 1 is a VLM model enhanced with both global and local spatial awareness, used for action generation. System 2 integrates a general MLLM with a mapping module, wherein the MLLM plans waypoints by leveraging top-down views of the real-time 3D map alongside streams of rendered path images. Both systems leverage different forms of spatial enhancement to cultivate the agent's sense of direction in VLN tasks. Ultimately, they cooperate to complete the navigation task through a fast-slow coordinated approach. SEDualVLN achieves state-of-the-art performance on VLN-CE benchmarks, and further ablation studies demonstrate the effectiveness of each system and module.
Jingzhi Huang, Junkai Huang, Wenxuan Song +4
Hong Kong Polytechnic University · Institute of Automation, Chinese Academy of Sciences · Hong Kong University of Science and Technology (Guangzhou)
Vision-language navigation requires aligning language with visual observations while maintaining spatial understanding over time. Geometry foundation models (GFMs) expose intermediate representations throughout their hierarchy, but how navigation policies should use these features and retain historical geometric evidence remains unresolved. We introduce \method{}, a streaming VLN framework that addresses these questions across \textbf{representation depth} and \textbf{navigation time}. Hierarchical GFM--VLM fusion couples earlier, intermediate, and later GFM representations to successive policy stages instead of repeatedly injecting a terminal feature. Navigation-aware GFM memory retains historical VGGT global-attention KV states according to instruction relevance, geometric confidence, and transition novelty under a bounded per-layer budget. Retained states provide geometric context for subsequent observations before fusion with the policy. Across R2R-CE and RxR-CE, \method{} achieves strong performance using a single RGB stream without additional navigation-specific external data. Controlled ablations show that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations. Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention. These findings support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference. Code will be released upon acceptance at https://humanoid-research.github.io/adageovln/.
Quan-Dung Pham, Anh Dao, Danh Vinh Le +7
VinMotion, Inc., Vietnam · Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA · University of Southern California, USA
Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to visual evidence and high-level decisions lack precise spatial goals to guide low-level motion. Cognitive science suggests that human navigation bridges these levels hierarchically by anchoring cognition to relevant landmarks and guiding locomotion toward spatial goals. Motivated by this principle, we propose GroundingVLN, which uses visual grounding as a shared interface between reasoning and action. GroundingVLN first reasons with grounding by anchoring task-relevant visual evidence to precise image locations throughout structured reasoning. It then acts through grounding by predicting a progress-aligned pixel goal that a geometric planner translates into primitive actions. To learn these capabilities, we construct GroundingCOTVLN-188K, a dataset of temporally aligned grounded reasoning traces, and introduce Grounded and Execution-Aware Reinforcement Learning (GEAR), which aligns grounded reasoning and spatial decisions with downstream execution. Experiments demonstrate that GroundingVLN achieves state-of-the-art performance (69.9% SR on R2R-CE and 75.1% SR on RxR-CE) with high sample efficiency, using just 0.9% as much training data as the strongest baseline. It also generalizes strongly across datasets, attaining 59.9% SR on RxR-CE when trained solely on R2R, a gain of 20.1% over the strongest baseline. Code and models will be released after review.
Kailing Li, Yu Han, Tianwen Qian +4
School of Computer Science and Technology, East China Normal University · NeoteAI · King Abdullah University of Science and Technology