cs.ROMar 14, 2026

ImagiNav: Scalable Embodied Navigation via Generative Visual Prediction and Inverse Dynamics

Authors: Jie Chen, Yuxin Cai, Yizhuo Wang, Ruofei Bai, Yuhong Cao, Jun Li, Wei-Yun Yau, Guillaume Sartoretti

Organizations: Department of Mechanical Engineering, National University of Singapore, Singapore · Institute for Infocomm Research (I2R), Agency for Science, Technology and Research (A*STAR), Singapore · Nanyang Technological University (NTU), Singapore

Abstract

Enabling robots to navigate open-world environments via natural language is critical for general-purpose autonomy. Yet, Vision-Language Navigation has relied on end-to-end policies trained on expensive, embodiment-specific robot data. While recent foundation models trained on vast simulation data show promise, the challenge of scaling and generalizing persists due to the limited scene diversity and visual fidelity in simulation. To address this gap, we propose ImagiNav, a novel hierarchical paradigm that formulates navigation in visual space. Instead of predicting waypoints, ImagiNav synthesizes a future egocentric video conditioned on language instructions, serving as a high-level plan, interpreted by an inverse dynamics model to extract metric trajectories for execution. By decoupling planning from robot actuation, the paradigm enables direct utilization of diverse in-the-wild navigation videos. To support this, we develop an auto-labeling data pipeline that enhances motion annotation accuracy. ImagiNav demonstrates strong zero-shot transfer to robot navigation without requiring robot demonstrations, paving the way for generalist robots that learn navigation directly from unlabeled, open-world data. The project page is available at: https://j1dan.github.io/ImagiNav

Figures & tables

Explore similar work

CardsList
  1. OpenFrontier: General Navigation with Visual-Language Grounded Frontiers

    Mar 5, 2026Esteban Padilla-Cerdio, Boyang Sun, Marc Pollefeys +1Robot NavigationVision-Language Navigation

  2. What Limits Vision-and-Language Navigation ?

    May 13, 2026Yunheng Wang, Yuetong Fang, Taowen Wang +9Sim-to-Real TransferVision-Language Navigation