Organizations: The Graduate School of Comprehensive Human Sciences, University of Tsukuba, Ibaraki, Japan. Work done during an internship at CyberAgent AI Lab. · CyberAgent AI Lab, Tokyo, Japan · The Department of Mechanical Systems Engineering, Nagoya University, Aichi, Japan
Visual Navigation Models (VNMs) enable robots to navigate from egocentric visual observations without geometric localization and planning, but long-range navigation still requires pre-built maps. This paper presents the Remote Visual Navigation Model (ReVNM), which uses a single remote surveillance camera to serve as both an observation source and an implicit environmental map for visual navigation. While the use of remote cameras could eliminate the need for pre-built maps as well as onboard vision processing, their limited field of view instead of egocentric observations makes it hard to achieve collision-free navigation. The lack of existing data with diverse remote viewpoints, which are crucial for training robust VNMs, further complicates the challenge. In this work, we propose a learning-by-synthesis approach to address this two-fold challenge. Our ReVNM extends a state-of-the-art VNM architecture with an exocentric-to-egocentric (exo2ego) module that predicts an egocentric depth observation from remote-camera observations. This helps the VNM to plan a path while considering obstacles in front of the robot. Trained only on randomly generated worlds with diverse obstacle layouts and camera viewpoints, ReVNM can generalize well to real robot navigation without additional fine-tuning. Experiments in both simulation and real-world environments confirmed the effectiveness of the proposed approach.
Figures & tables
Fig. 2: Model architecture of ReVNM. (a) The navigation policy is an ACT-style CVAE. It encodes the ego and exo depth with ResNet-34 and predicts a chunk of waypoints from these, the robot pixel, orientation, and goal. (b) exo2ego is a Diffusion Transformer that synthesizes the ego depth from the cropped exo depth, conditioned on the robot pixel and orientation.
Method
Random Pillar
Book Store
Warehouse
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
NoMaD w/o FT
5 ± 3%
0.02 ± 0.01
18 ± 5%
0.08 ± 0.02
20 ± 5%
0.08 ± 0.02
NoMaD w/ FT
40 ± 3%
0.33 ± 0.03
22 ± 3%
0.09 ± 0.02
22 ± 4%
0.09 ± 0.02
IBVS
78 ± 2%
0.61 ± 0.05
75 ± 5%
0.53 ± 0.04
52 ± 7%
0.36 ± 0.06
Ours
88 ± 3%
0.78 ± 0.03
88 ± 5%
0.65 ± 0.05
70 ± 6%
0.46 ± 0.05
TABLE I: Navigation performance in simulation.
Fig. 3: Results in the simulation experiments. Each column is one environment, with the remote camera view on top and a top view of the same trials below. S and G mark the start and the goal, and a cross marks where a run failed.
Fig. 4: Egocentric depth synthesized by the exo2ego module in Book Store (top) and Warehouse (bottom).
Fig. 5: Synthesized ego depth with and without the DAgger-based recovery data. Only w/ DAgger reproduces the obstacle in front of the robot (top), while in open space the two are comparable (bottom).
Method
DAgger
Random Pillar
Book Store
Warehouse
SR ↑
SPL ↑
SR ↑
SPL ↑
SR ↑
SPL ↑
Ours
✓
88 ± 3%
0.78 ± 0.03
88 ± 5%
0.65 ± 0.05
70 ± 6%
0.46 ± 0.05
87 ± 4%
0.69 ± 0.04
74 ± 6%
0.59 ± 0.06
43 ± 6%
0.29 ± 0.05
Ours w/o exo2ego
✓
91 ± 3%
0.79 ± 0.03
84 ± 5%
0.61 ± 0.05
26 ± 8%
0.17 ± 0.06
78 ± 6%
0.64 ± 0.06
44 ± 7%
0.33 ± 0.05
12 ± 4%
0.07 ± 0.03
Ours w/ oracle-ego
✓
99 ± 1%
0.89 ± 0.01
90 ± 4%
0.78 ± 0.04
83 ± 5%
0.58 ± 0.05
TABLE II: Ablation study.
exo2ego
Random Pillar
Book Store
Warehouse
PSNR ↑
SSIM ↑
FID ↓
PSNR ↑
SSIM ↑
FID ↓
PSNR ↑
SSIM ↑
FID ↓
w/ DAgger
15.2
0.611
14.5
13.1
0.454
24.6
13.8
0.480
23.7
w/o DAgger
14.4
0.601
17.3
12.7
0.494
25.9
12.5
0.488
26.0
TABLE III: Quality of the synthesized ego depth.
Method
Forest
Wall
SR ↑
SR ↑
IBVS
12% (3/25)
0% (0/5)
Ours w/o exo2ego
92% (23/25)
40% (2/5)
Ours
76% (19/25)
100% (5/5)
TABLE IV: Navigation performance in the real-world.
Fig. 6: Results in the real-world experiments, overlaid on the remote camera view.
While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.
Visual navigation policy is widely regarded as a promising direction, as it mimics humans by using egocentric visual observations for navigation. However, optical information of visual observations is difficult to be explicitly modeled like LiDAR point clouds or depth maps, which subsequently requires intelligent models and large-scale data. To this end, we propose to leverage the intelligence of the Vision-Language-Action (VLA) model to learn diverse navigation capabilities from synthetic expert data in a teacher-student manner. Specifically, we implement the VLA model, MM-Nav, as a multi-view VLA (with 360 observations) based on pretrained large language models and visual foundation models. For large-scale navigation data, we collect expert data from three reinforcement learning (RL) experts trained with privileged depth information in three challenging tailor-made environments for different navigation capabilities: reaching, squeezing, and avoiding. We iteratively train our VLA model using data collected online from RL experts, where the training ratio is dynamically balanced based on performance on individual capabilities. Through extensive experiments in synthetic environments, we demonstrate that our model achieves strong generalization capability. Moreover, we find that our student VLA model outperforms the RL teachers, demonstrating the synergistic effect of integrating multiple capabilities. Extensive real-world experiments further confirm the effectiveness of our method.
Tianyu Xu, Jiawei Chen, Jiazhao Zhang +5
Peking University, Beijing, China · Galbot, Beijing, China · Shanghai Jiao Tong University, Shanghai, China +1
We introduce VEGA, an approach for training navigation VisionLanguage-Action (VLA) models from unlabeled egocentric navigation videos. Internet-scale egocentric videos provide a scalable source of navigation-relevant visual observations, capturing cluttered scenes, close-range obstacles, and natural human motion through real-world spaces. However, these videos are not directly usable for policy learning because they do not provide obstacle-aware trajectories conditioned on explicit navigation goals in the robot's coordinate frame. VEGA addresses this gap by reconstructing local scene geometry from monocular video, sampling navigation goals (represented as text, image, or spatial waypoints) and generating obstacle-aware trajectories using the constructed geometry. The resulting trajectory distribution is then used to train a flow-matching VLA navigation policy. By using geometry exclusively during training, VEGA distills obstacle-aware planning directly into a vision-based policy. Furthermore, we introduce VEGA-Bench, a benchmark containing 250k scenes and approximately 5 million navigation goals paired with scene geometry, designed to evaluate goal progress, collision avoidance, and obstacle clearance of VLAs. Our evaluation shows that VEGA achieves competitive goal progress while reducing collisions by 33.0% and improving obstacle clearance by 17.9% over the strongest baseline on VEGABench, while improving success by at least 150.0%, reducing collisions by at least 66.7%, and improving obstacle clearance by at least 60.0% in real-world trials. Ultimately, we demonstrate that video-derived geometric supervision provides a scalable and effective signal for training obstacle-aware navigation VLAs. The code and benchmark will be released at the time of publication.
Gershom Seneviratne, Yohan Abeysinghe, Jianyu An +2