cs.CVOct 5, 2026

StageVLN: Spatial and Trajectory Auxiliary Guidance for Efficient Vision-Language Navigation

Authors: Anh Dao, Quan-Dung Pham, Le Danh Vinh, The Anh Nguyen, Nguyen Viet Tri Pham, Yiyu Chen, Pham Tuyen Le, Van-Truong Nguyen, +1 more

Organizations: VinMotion, Inc., Vietnam · University of Southern California, USA

Abstract

Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3% SR and 51.4% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3% SR without additional navigation training data or a geometry encoder at inference.

Figures & tables

Explore similar work

CardsList
  1. SEDualVLN: A Spatially-Enhanced Dual-System for Vision-Language Navigation

    May 17, 2026Jingzhi Huang, Junkai Huang, Wenxuan Song +4Vision-Language NavigationMultimodal Large Language Models

  2. AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation

    Sep 16, 2026Quan-Dung Pham, Anh Dao, Danh Vinh Le +7Vision-Language NavigationRepresentation Quality

  3. GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

    Sep 16, 2026Kailing Li, Yu Han, Tianwen Qian +4Vision-Language NavigationGrounding