Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3% SR and 51.4% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3% SR without additional navigation training data or a geometry encoder at inference.
Figures & tables
Fig. 1 : StageVLN overview. Left: comparison with explicit 3D inputs and runtime geometry fusion; StageVLN uses spatial and trajectory guidance only during training and retains the navigator alone at inference. Right: navigation success and inference efficiency. VRAM and latency are measured on a single NVIDIA RTX 6000 Pro.
Fig. 2 : StageVLN training and inference pipeline. During training, a frozen geometry teacher provides multi-level spatial guidance, while relative-heading and expert-route progress objectives supervise trajectory state. At inference, the teacher, alignment projectors, and auxiliary heads are excluded, leaving only the base RGB–language navigation policy.
Method
Model Params
Observation
R2R Val-Unseen
RxR Val-Unseen
External Training Data
Pano.
Odo.
Depth
S.RGB
NE ↓
OS ↑
SR ↑
SPL ↑
NE ↓
SR ↑
SPL ↑
nDTW ↑
HPN+DN [ 21 ]
–
✓
✓
✓
6.31
40.0
36.0
34.0
–
–
–
–
–
CMA [ 22 ]
–
✓
✓
✓
6.20
52.0
41.0
36.0
8.76
26.5
22.1
47.0
–
Sim2Sim [ 23 ]
–
✓
✓
✓
6.07
52.0
43.0
36.0
–
–
–
–
–
VLN-BERT [ 24 ]
–
✓
✓
✓
5.74
53.0
44.0
39.0
8.98
27.0
22.6
46.7
–
Ego2-Map [ 25 ]
–
✓
✓
✓
5.54
56.0
47.0
41.0
–
–
–
–
–
TABLE I : Comparison on the R2R-CE and RxR-CE validation-unseen splits. ↓ / ↑ indicate lower/higher is better. Best and second-best results are shown in bold and underline , respectively. StreamVLN ∗ uses EnvDrop augmentation, and NaVILA ∗ excludes human-following data. † denotes methods with an additional geometry encoder at inference; parameter counts are reported as VLM + geometry encoder. StageVLN uses no additional navigation training data beyond R2R-CE and RxR-CE; pretrained model data are excluded from the external-data counts.
Method
Peak allocated VRAM (GiB) ↓
Mean step time (ms) ↓
RTX 5090
JanusVLN
OOM
–
NaVILA
17.47
365.2
StreamVLN
17.48
233.4
StageVLN
10.24
304.5
RTX 6000 Pro
TABLE II : Inference memory and mean step time, grouped by GPU. StageVLN is shaded; bold values indicate the best result within each group. OOM denotes an out-of-memory failure under the tested configuration.
Fig. 3 : Simulation rollouts and real-world deployment. (a) Unitree G1 platform with a ZED X Mini camera and Jetson AGX Orin. (b) Simulated episode comparing JanusVLN and StageVLN under the same instruction and initial condition. (c) Two real-world deployment sequences, shown in chronological order.
Nav.
Geo.
Head.
Prog.
NE ↓
OS ↑
SR ↑
SPL ↑
✓
–
5.19
56.01
46.49
42.32
✓
1L
6.33
61.83
48.50
41.86
✓
ML
6.12
61.01
52.26
46.72
✓
ML
✓
6.17
63.30
49.97
43.80
✓
ML
✓
5.66
63.89
53.18
48.67
✓
–
✓
✓
5.86
64.06
54.43
48.58
TABLE III : Component ablation on R2R-CE validation-unseen. VGGT is used as the geometry teacher for all variants. Nav. : navigation supervision; Geo. : geometric representation distillation; Head. : relative heading; Prog. : expert-route progress. 1L uses only the final alignment pair (navigator layer 24, teacher index 23), while ML uses all three alignment pairs. The first row is supervised fine-tuning (SFT); the shaded row is full StageVLN. Best results are shown in bold .
Geometry teacher
NE ↓
OS ↑
SR ↑
SPL ↑
VGGT
5.58
67.37
55.36
48.44
VGGT- Ω
5.17
64.71
56.28
51.37
TABLE IV : Geometry-teacher ablation on R2R-CE validation-unseen. Each variant uses the full StageVLN objective with a frozen teacher. Best results are shown in bold
VinMotion, Inc., Vietnam · Department of Electrical and Computer Engineering, University of California San Diego, La Jolla, CA 92093, USA · University of Southern California, USA