Recent vision-language models (VLMs) have advanced vision-and-language navigation (VLN), enabling models to predict navigation actions from visual observations and language instructions. In this work, we explore VLN with panoramic observations and introduce PanoVLN. The motivation is straightforward: more complete visual context should enable better-informed navigation decisions. For example, a panorama can reveal a passage outside a perspective camera's field of view, allowing the model to identify the intended route without additional exploration. However, we find that simply replacing perspective images with panoramas yields only limited gains. Our diagnosis suggests that fully exploiting wider visibility requires modifications to action prediction, training supervision, and visual representation. First, wider visibility supports longer-horizon action planning. We make the model predict longer action sequences, enabling larger turns and subsequent movement from a single panorama. Specifically, we introduce a confidence-guided execution (CGE) strategy that dynamically determines how many predicted actions to execute before replanning. Second, wider visibility also brings more complex route choices. We therefore construct training routes with frequent branching points and clear instructions to provide targeted supervision for route selection. Third, panoramic navigation requires understanding spatial relationships across viewing directions, beyond recognizing individual landmarks. We combine semantic and geometric features from RGB panoramas to capture both scene content and spatial layout without adding visual tokens. With a 4B backbone and RGB-only input, PanoVLN surpasses the previous SOTA by 11.9% and 8.7% in success rate on R2R-CE and RxR-CE Val-Unseen. Real-world experiments on a quadruped further demonstrate faster navigation with fewer pauses than prior VLN methods.
Figures & tables
Figure 1: PanoVLN fully exploits the wider visual context of panoramas to make vision-and-language navigation more accurate and efficient.
Figure 2: PanoVLN pipeline. Current and historical panoramas share a fixed visual-token budget. Aligned geometric features enrich the current visual tokens, and uncertainty determines how much of the predicted action sequence to execute.
Figure 3: Action uncertainty. Uncertainty rises after an initial dip and varies across policy calls, motivating adaptive execution.
Figure 4: Data construction pipeline . We construct routes with frequent branching points, generate and verify instructions, and sample training states.
Figure 5: Real-world navigation. PanoVLN deployed on a quadruped robot and transfers to real-world instruction following without scene-specific fine-tuning.
Figure 6: Real-world navigation performance. We compare SR and NE on 20 shared instruction–route pairs each in Hallway, Office, and Campus and achieves the best navigation performance.
Method
Navigation
Continuity
Planning
Time (s) ↓
Speed (cm/s) ↑
Wait (%) ↓
Pauses ↓
Calls ↓
Latency (s) ↓
NaVid
135.2
12.9
23.2
15.7
29.4
0.90
NaVILA
194.0
8.2
24.7
29.6
41.3
1.05
StreamVLN
117.8
13.3
27.1
8.3
34.3
0.58
JanusVLN
412.9
5.3
39.2
95.3
101.7
1.32
PanoVLN (Ours)
86.4
25.7
13.6
5.4
7.4
1.08
Table 2: Real-world execution efficiency. PanoVLN achieves the best overall navigation efficiency among the compared methods.
Figure 7: Prediction horizon ablation. Panoramic policies benefit from longer supervision and outperform perspective policies at longer horizons.
Sampling
NE ↓
OS ↑
SR ↑
SPL ↑
Random
5.37
70.7
56.4
49.5
Ours
4.40
69.9
62.8
57.2
Table 3: Training data comparison.
Figure 8: Training data comparison. We compare data sources across training scales on R2R-CE Val-Unseen. Our data consistently delivers the best navigation performance.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Value
VLM / frozen geometry encoder
Qwen3.5-4B / PanoVGGT
Geometry features / fusion
Last aggregator layer; 2×2 grouping; fusion after the visual merger
Figure 10: Reference-path turn distribution in R2R-CE and RxR-CE, normalized over 54,457 turns of at least 30∘ . The 45∘ – 180∘ categories account for 53.3% of turns.
Figure 11: Additional paired attention cases, each using the same ERP and instruction. Red bold text marks the relevant route clause; dashed lines bound the forward 90∘ sector. Each case note identifies the visible hotspot locations.
Figure 12: Real-world navigation cases in Hallway, Office, and Campus. Each row shows the instruction and successive observations from one trial, with route annotations.
Figure 13: Navigation cases on R2R-CE Val-Unseen. Each case shows four panoramic observations in reading order and the full instruction. Numbered markers link each observation to its location on the trajectory; blue and green denote the executed and reference routes.
Figure 14: Navigation cases on RxR-CE Val-Unseen. Four panoramic observations per case are linked to the trajectory by numbered markers. Blue and green denote the executed and reference routes. Instruction excerpts retain the original wording; ellipses mark omissions.
Vision-and-language navigation (VLN) requires an embodied agent to ground natural-language instructions into executable navigation actions in unseen environments. Existing zero-shot methods typically rely on additional waypoint prediction modules, which often entangle high-level directional reasoning with fine-grained local grounding, leading to error-prone and unstable decisions. In this paper, we propose P2DNav, a hierarchical framework for zero-shot vision-and-language navigation. P2DNav consists of three core components: Panorama-to-Downview (P2D), Sliding-Window Dialogue Memory (SDM), and Reflective Reorientation Mechanism (RRM). P2D explicitly decomposes navigation decision-making into two stages: panoramic direction selection and downview local grounding. It first selects the instruction-relevant direction from a 360° panorama, and then predicts a pixel-level target point from the downview RGB observation in that direction. In addition, SDM organizes navigation history as a multi-turn dialogue context and maintains recent visual observations within a sliding window to support long-horizon navigation. RRM further enables reflective reorientation by assessing the reliability of local grounding based on the downview observation and returning to panoramic direction selection when necessary. Experiments on the R2R-CE benchmark show that P2DNav achieves strong performance among zero-shot methods. In particular, compared with the state-of-the-art (SOTA) zero-shot waypoint-based and waypoint-free methods, P2DNav achieves SR gains of 146.6% and 58.9%, respectively, demonstrating the effectiveness of P2D, SDM, and RRM for zero-shot VLN. Code will be released for public use.
Kai Sheng, Liuyi Wang, Haojie Dai +5
Department of Control Science and Engineering, Tongji University, Shanghai 210804, China
Vision-and-Language Navigation (VLN) policies increasingly benefit from strong semantic priors provided by large vision-language models (VLMs). However, standard action supervision does not explicitly encourage intermediate representations to preserve scene geometry, relative orientation, or global episode progress. Incorporating depth estimators, explicit maps, point clouds, or geometry foundation models at inference can provide such structure but introduces additional computation, memory overhead, and architectural dependence during deployment. We introduce StageVLN, a training framework that shapes navigation representations through privileged spatial and trajectory guidance while preserving the original inference pathway. A frozen geometry foundation model provides multi-level spatial guidance to hierarchical navigator states, while relative-heading and expert-route progress objectives provide complementary trajectory-state supervision. All auxiliary components are used only during training and removed at deployment. On R2R-CE validation-unseen, StageVLN achieves 56.3% SR and 51.4% SPL with a 4B-parameter backbone, without an additional geometry encoder at inference. On RxR-CE, it achieves 54.3% SR without additional navigation training data or a geometry encoder at inference.
Anh Dao, Quan-Dung Pham, Le Danh Vinh +6
VinMotion, Inc., Vietnam · University of Southern California, USA
Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While this fusion of language instructions and visual observations allows multimodal reasoning, it obscures how information is routed across modalities or what mechanisms drive navigation decisions. Thus, it remains unclear whether VLN models ground their predictions in relevant semantic cues or can track task progress. In this work, we study the interpretability and steerability of VLN models. We use intervention-based metrics that measure how visual observations, instructions, and visual memory causally influence navigation decisions. Our results show that these navigation policies are sensitive to all input modalities and do not depend on a single one. We further show that these agents encode navigation progress and retain semantic structure from their VLM backbones, enabling concept-level steering through internal activations. Finally, we extract activation vectors for abstract behaviors to transfer them zero-shot to out-of-distribution real-world scenarios, improving performance without additional fine-tuning.