ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics
Organizations: Department of Mechanical Engineering, National University of Singapore, Singapore · Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR), Singapore · Nanyang Technological University (NTU), Singapore
Abstract
Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation. To address this gap, we propose ImagiNav, a novel hierarchical paradigm that formulates navigation in visual space. Instead of predicting waypoints, ImagiNav synthesizes a future egocentric video conditioned on language instructions, serving as a high-level plan, interpreted by an inverse dynamics model to extract metric trajectories for execution. By decoupling planning from robot actuation, the paradigm enables direct utilization of diverse in-the-wild navigation videos. To support this, we develop an auto-labeling data pipeline that enhances motion annotation accuracy. ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn navigation directly from unlabeled, open-world data. The project page is available at: https://j1dan.github.io/ImagiNav
Figures & tables
| Method | Depth | Sim Sim | Data Size | TL | NE | OS | SR | SPL |
| In-Domain Supervised | ||||||||
| InternVLA-N1 [ 2 ] | ✓ | ✓ | 1000h | 6.42 | 3.48 | 0.63 | 0.53 | 0.44 |
| Zero-Shot Transfer | ||||||||
| NavDP [ 20 ] | ✓ | ✓ | 101h | 3.48 | 4.19 | 0.37 | 0.32 | 0.31 |
| ImagiNav-Sim | ✓ | 0.5h | 3.92 | 3.94 | 0.45 | 0.39 | 0.37 | |
| ImagiNav-Real | 0.5h | 4.03 | 4.13 | 0.41 | 0.36 | 0.35 | ||
| Model Variant | FVD | LPIPS | PSNR | SSIM | Mot. Fid. | RPE-T (norm.) | RPE-R ( ∘ ) |
|---|---|---|---|---|---|---|---|
| Base LTX-2B (Zero-shot) | 391.14 | 0.520 | 12.61 | 0.37 | 0.40 | 0.036 | 1.39 |
| Real-Finetuned w/o AC-MoE | 73.08 | 0.486 | 13.36 | 0.40 | 0.60 | 0.028 | 1.31 |
| Sim-Finetuned | 72.72 | 0.514 | 13.03 | 0.39 | 0.67 | 0.051 | 1.73 |
| Real-Finetuned | 65.39 | 0.481 | 13.27 | 0.40 | 0.72 | 0.024 | 1.18 |
| Note: Mot. Fid.: Motion Fidelity, RPE-T: Translational Relative Pose Error, RPE-R: Rotational Relative Pose Error. | |||||||
| Best results are bolded . | |||||||