ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics
Authors: Jie Chen, Yuxin Cai, Yizhuo Wang, Ruofei Bai, Yuhong Cao, Jun Li, Wei-Yun Yau, Guillaume Sartoretti
Organizations: Department of Mechanical Engineering, National University of Singapore, Singapore · Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR), Singapore · Nanyang Technological University (NTU), Singapore
Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation. To address this gap, we propose ImagiNav, a novel hierarchical paradigm that formulates navigation in visual space. Instead of predicting waypoints, ImagiNav synthesizes a future egocentric video conditioned on language instructions, serving as a high-level plan, interpreted by an inverse dynamics model to extract metric trajectories for execution. By decoupling planning from robot actuation, the paradigm enables direct utilization of diverse in-the-wild navigation videos. To support this, we develop an auto-labeling data pipeline that enhances motion annotation accuracy. ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn navigation directly from unlabeled, open-world data. The project page is available at: https://j1dan.github.io/ImagiNav
Figures & tables
Fig. 1 : Overview of the ImagiNav Framework. Top: Navigation is formulated in visual space. A generative visual planner synthesizes future egocentric observations, which are geometrically grounded into pose trajectories via inverse dynamics and executed by a tracking controller. Bottom: Geometry-first scalable training pipeline from in-the-wild egocentric videos.
Fig. 2 : Dataset Statistics. (a) Distribution of raw videos across different real-world environments. (b) Distribution of motion primitives in the final augmented dataset, showing balanced coverage.
Method
Depth
Sim → Sim
Data Size
TL
NE ↓
OS ↑
SR ↑
SPL ↑
In-Domain Supervised
InternVLA-N1 [ 2 ]
✓
✓
∼ 1000h
6.42
3.48
0.63
0.53
0.44
Zero-Shot Transfer
NavDP [ 20 ]
✓
✓
∼ 101h
3.48
4.19
0.37
0.32
0.31
ImagiNav-Sim
×
✓
∼ 0.5h
3.92
3.94
0.45
0.39
0.37
ImagiNav-Real
×
×
∼ 0.5h
4.03
4.13
0.41
0.36
0.35
TABLE I : Quantitative Results on InteriorNav Dataset. We report performance relative to input modality (Depth), domain alignment (Sim → Sim), and data scale.
Fig. 3 : Qualitative Visualization of the Imagination-Action Cycle. (Left) The imagined future frame generated by the model. (Middle) The actual robot observation after executing the extracted plan. (Right) Top-down trajectory map, where the red line indicates the executed path and the green line represents the reference ground truth. The close alignment confirms the physical consistency of the generated plans.
TABLE II : Quantitative Evaluation of LTX-2B Video Model with Sim/Real Data
Fig. 4 : Qualitative Comparison of Simulation vs. Real-World Training. We visualize the generated future trajectories across Ground Truth (Top Row), ImagiNav-Real (Middle Row), and ImagiNav-Sim (Bottom Row). Real-world data demonstrates superior understanding of static geometry (Scene A, the navigable walkway is circled in green ), and dynamic behavior (Scene B).
Fig. 5 : Controllability and Geometric Grounding. Given a single start frame, the model generates diverse trajectories conditioned on different language instructions (from left to right: “Move forward and turn left” , “Move forward” , and “Move forward and turn right” ).
Open-world navigation requires robots to make decisions in complex everyday environments while adapting to flexible task requirements. Conventional navigation approaches often rely on dense 3D reconstruction and hand-crafted goal metrics, which limits their generalization across tasks and environments. Recent advances in vision-language navigation (VLN) and vision-language-action (VLA) models enable end-to-end policies conditioned on natural language, but typically require interactive training, large-scale data collection, or task-specific fine-tuning with a mobile agent. We formulate navigation as a sparse subgoal identification and reaching problem and observe that providing visual anchoring targets for high-level semantic priors enables highly efficient goal-conditioned navigation. Based on this insight, we select visual frontiers as semantic anchors and propose OpenFrontier, a navigation framework that requires no task-specific training or fine-tuning and seamlessly integrates diverse vision-language prior models. OpenFrontier enables efficient navigation with a lightweight system design, without dense 3D semantic mapping, task-specific policy training, or model fine-tuning. We evaluate OpenFrontier across multiple navigation benchmarks and demonstrate strong zero-shot performance, as well as effective real-world deployment on a mobile robot.
Esteban Padilla-Cerdio, Boyang Sun, Marc Pollefeys +1
1ETH Zurich · 2Microsoft Spatial AI Lab · University of Bonn
Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control. Existing navigation systems instead rely on task- or embodiment-specific components, fragmenting perception, reasoning, and action while offering limited generalization. Here we present LightNav-0, a compact generalist embodied navigation model that elicits the spatial intelligence of a pretrained VLM and aligns it with navigation, without task-specific prediction heads. LightNav-0 represents diverse navigation tasks through a unified token interface: dual-channel pointing expresses task-, scene-, and embodiment-agnostic spatial intent, while a residual vector-quantized action tokenizer maps this intent to precise, embodiment-specific trajectories. Together with temporally aware visual history compression, ER mid-training, supervised fine-tuning, and reinforcement learning, this formulation supports instruction following, open-vocabulary object navigation, and visual tracking within a single model. The navigation training corpus spans 2K+ scenes and 4K+ hours of embodied navigation data. LightNav-ER, the embodied-reasoning checkpoint used to initialize LightNav-0, attains the highest complete-set average across 8 embodied-reasoning benchmarks, while LightNav-0 achieves state-of-the-art monocular success rates across all 10 public navigation simulation settings. Real-world evaluations further demonstrate zero-shot generalization across robot embodiments, diverse scenes, and static and dynamic targets. These results establish compact VLMs as a unified and transferable backbone for generalist embodied navigation.
Vision-and-Language Navigation (VLN) is a cornerstone of embodied intelligence. However, current agents often suffer from significant performance degradation when transitioning from simulation to real-world deployment, primarily due to perceptual instability (e.g., lighting variations and motion blur) and under-specified instructions. While existing methods attempt to bridge this gap by scaling up model size and training data, we argue that the bottleneck lies in the lack of robust spatial grounding and cross-domain priors. In this paper, we propose StereoNav, a robust Vision-Language-Action framework designed to enhance real-world navigation consistency. To address the inherent gap between synthetic training and physical execution, we introduce Target-Location Priors as a persistent bridge. These priors provide stable visual guidance that remains invariant across domains, effectively grounding the agent even when instructions are vague. Furthermore, to mitigate visual disturbances like motion blur and illumination shifts, StereoNav leverages stereo vision to construct a unified representation of semantics and geometry, enabling precise action prediction through enhanced depth awareness. Extensive experiments on R2R-CE and RxR-CE demonstrate that StereoNav achieves state-of-the-art egocentric RGB performance, with SR and SPL scores of 81.1% and 68.3%, and 67.5% and 52.0%, respectively, while using significantly fewer parameters and less training data than prior scaling-based approaches. More importantly, real-world robotic deployments confirm that StereoNav substantially improves navigation reliability in complex, unstructured environments. Project page: https://yunheng-wang.github.io/stereonav-public.github.io.