Direct visual navigation policies generate trajectories efficiently but do not explicitly evaluate their future consequences. Generative navigation world models provide this foresight through visual rollouts, which are costly when evaluating multiple candidates. We present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories. In offline evaluations on RECON, SCAND, and SACSoN, LiteNWM reduces macro-averaged trajectory error by 17.56% relative to NoMaD+NWM-XL and achieves a 128.00-fold end-to-end speedup on an RTX 5090. The same evaluator transfers from NoMaD to MBRA without proposer-specific retraining, reducing MBRA's macro-averaged trajectory error by 16.2%. In real-robot experiments in unseen indoor and outdoor environments, LiteNWM improves navigation success from 43.3% to 83.3% relative to NoMaD. These results demonstrate that LiteNWM can be deployed for future-aware planning and closed-loop navigation on a physical robot.
Figures & tables
Fig. 2: Overall workflow of LiteNWM. (A) Given a four-frame observation history and a visual goal, a frozen goal-conditioned policy produces K=16 four-step trajectory challengers together with a retained incumbent τ0 . (B) A shared frozen DINOv3 encoder and a lightweight adapter encode the history and goal once per state. (C) The action-conditioned world model jointly predicts H1–H4 future latents for the incumbent and all challengers. A goal-conditioned, proposal-aware scorer compares these predicted evolutions, ranks eligible challengers, and either selects the best one or retains the incumbent. Further details are provided in Sec. III-C .
Fig. 3: Captured component-latency profile under matched RTX 5090, RECON-100, H=4 , and K=16 conditions. (a) Candidate-wise CDiT rollout accounts for 84.8% and 97.6% of the captured NoMaD+NWM-S and NoMaD+NWM-XL latency, respectively. (b) End-to-end outer-process wall time per state for the methods reported in Table III. LiteNWM amortizes candidate-wise RGB rollout through shared visual encoding, batched latent prediction, and trajectory scoring.
Fig. 4: Backbone of LiteNWM for joint H1–H4 latent prediction. Pose-conditioned warping provides W for query construction and R for residual prediction. Six factorized blocks integrate action and history conditions. Dashed arrows denote conditioning.
Fig. 5: Step error over the prediction horizon (H). The curves report the macro error and the errors on RECON, SCAND, and SACSoN for NoMaD, NoMaD+NWM-XL, and LiteNWM. LiteNWM shows a consistently slower error increase at later timesteps across all three domains.
Fig. 6: Representative offline trajectory selections on three datasets. For each dataset, the left panels show the start and goal observations for NoMaD+NWM-S, NoMaD+NWM-XL, and LiteNWM, while the right panel compares the selected trajectories with ground truth over the first 3 m of longitudinal progress.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 8: Predictor-action interventions in the fixed-pool diagnostic. Points show intervention minus matched-action macro ATE; bars show 95% paired cluster-bootstrap CIs using 100/7/100 groups over 300 queries and four proposal seeds. Model weights, visual inputs, candidate trajectories, and scorer-side action descriptors remain fixed.
Fig. 9: Robustness across endpoint tolerances and evaluation states. (a) Equal-domain macro endpoint agreement, P(FDE≤τ) ; the vertical line marks SR@1. (b) Paired endpoint-agreement advantage over NoMaD+NWM-XL, with macro and domain-level estimates. (c) Per-state ATE reduction relative to NoMaD+NWM-XL. Grey points, violins, and diamonds show observations, distributions, and means, respectively. Positive values favor LiteNWM; percentages report per-state win rates. Shading and horizontal bars indicate 95% trajectory-cluster bootstrap CIs from 20,000 resamples; bands in (a,b) are pointwise. Each primary domain contains 100 states. SCAND-bag contains 48 states from six clusters and is a descriptive sensitivity cohort excluded from the macro average. DWM denotes the historical LiteNWM configuration; success measures offline endpoint agreement.
Fig. 10: Indoor and outdoor route progression. (a) Doorway traversal, corridor navigation, turning, and approach to a chair goal. (b) Outdoor navigation across paved areas and beside vegetation toward the visual target. Four observations in each row follow video order from left to right, followed by a cropped visual-goal reference. Trajectory overlays reproduce the source-video visualization. Frame spacing does not represent uniform elapsed time.
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and evaluate it on image-goal navigation against planning-based world models and a representative direct navigation policy. Across offline benchmarks and closed-loop real-robot deployment, NavWAM improves over planning-based world-model baselines in our evaluations while using the default policy mode without CEM-style action search. Project page: https://dachii-azm.github.io/navwam/
Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6
The University of Tokyo · National Institute of Informatics · AIRoA +1
Conventional visual navigation policies often struggle with myopic decision-making and mode collapse in complex environments. While world models offer a promising alternative, existing paradigms typically isolate perception, generation, and control, failing to capture their shared spatio-temporal dynamics. In this paper, we propose NavWM, a unified navigation world model that seamlessly integrates latent world reasoning, multimodal action prediction, and controllable visual generation. At its core, NavWM leverages latent world tokens to distill geometric and semantic priors, endowing the agent with robust structural understanding. To overcome the limitations of deterministic policies, we introduce an anchor-based multimodal trajectory forecasting framework that generates a diverse action space. This inherent diversity explicitly empowers the generative world model to act as a robust closed-loop planner, utilizing visual foresight to evaluate and select the optimal path. Extensive experiments across diverse robotics datasets demonstrate that NavWM significantly advances the state-of-the-art, delivering remarkable improvements in both high-fidelity future state generation and zero-shot navigation success.
Yanghong Mei, Longteng Guo, Ming-Ming Yu +3
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Beihang University
Image-goal navigation with latent world models requires not only accurate future prediction, but also a planning cost that reliably ranks candidate action sequences. We define the cost as the cosine distance between the predicted future embedding and the goal embedding, and show that poor cost ordering can mislead sampling-based planners such as Cross-Entropy Method (CEM). To address this, we propose a latent world model built on a frozen DINO-family encoder and train it with two complementary objectives. An autoregressive rollout loss reduces the gap between training and multi-step planning rollouts, while a Monotone Cost Ranking (MCR) loss directly encourages increasingly perturbed action sequences to receive higher planning costs. We also study InfoNCE-based action-contrastive training and find that temporal permutation negatives distort the latent geometry and degrade planning performance. On the GNM navigation dataset, our method outperforms Navigation World Models (NWM), DINO-WM, OmniVLA, and NoMaD, achieving state-of-the-art image-goal navigation performance while reducing orientation error by 2.7× over the same-encoder DINO WM baseline. We also deploy the model zero-shot on a physical robot, where it follows goal-directed paths in unseen indoor and outdoor environments.