Direct visual navigation policies generate trajectories efficiently but do not explicitly evaluate their future consequences. Generative navigation world models provide this foresight through visual rollouts, which are costly when evaluating multiple candidates. We present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories. In offline evaluations on RECON, SCAND, and SACSoN, LiteNWM reduces macro-averaged trajectory error by 17.56% relative to NoMaD+NWM-XL and achieves a 128.00-fold end-to-end speedup on an RTX 5090. The same evaluator transfers from NoMaD to MBRA without proposer-specific retraining, reducing MBRA's macro-averaged trajectory error by 16.2%. In real-robot experiments in unseen indoor and outdoor environments, LiteNWM improves navigation success from 43.3% to 83.3% relative to NoMaD. These results demonstrate that LiteNWM can be deployed for future-aware planning and closed-loop navigation on a physical robot.
Figures & tables
Fig. 2: Overall workflow of LiteNWM. (A) Given a four-frame observation history and a visual goal, a frozen goal-conditioned policy produces K=16 four-step trajectory challengers together with a retained incumbent τ0 . (B) A shared frozen DINOv3 encoder and a lightweight adapter encode the history and goal once per state. (C) The action-conditioned world model jointly predicts H1–H4 future latents for the incumbent and all challengers. A goal-conditioned, proposal-aware scorer compares these predicted evolutions, ranks eligible challengers, and either selects the best one or retains the incumbent. Further details are provided in Sec. III-C .
Fig. 3: Captured component-latency profile under matched RTX 5090, RECON-100, H=4 , and K=16 conditions. (a) Candidate-wise CDiT rollout accounts for 84.8% and 97.6% of the captured NoMaD+NWM-S and NoMaD+NWM-XL latency, respectively. (b) End-to-end outer-process wall time per state for the methods reported in Table III. LiteNWM amortizes candidate-wise RGB rollout through shared visual encoding, batched latent prediction, and trajectory scoring.
Fig. 4: Backbone of LiteNWM for joint H1–H4 latent prediction. Pose-conditioned warping provides W for query construction and R for residual prediction. Six factorized blocks integrate action and history conditions. Dashed arrows denote conditioning.
Fig. 5: Step error over the prediction horizon (H). The curves report the macro error and the errors on RECON, SCAND, and SACSoN for NoMaD, NoMaD+NWM-XL, and LiteNWM. LiteNWM shows a consistently slower error increase at later timesteps across all three domains.
Fig. 6: Representative offline trajectory selections on three datasets. For each dataset, the left panels show the start and goal observations for NoMaD+NWM-S, NoMaD+NWM-XL, and LiteNWM, while the right panel compares the selected trajectories with ground truth over the first 3 m of longitudinal progress.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Fig. 8: Predictor-action interventions in the fixed-pool diagnostic. Points show intervention minus matched-action macro ATE; bars show 95% paired cluster-bootstrap CIs using 100/7/100 groups over 300 queries and four proposal seeds. Model weights, visual inputs, candidate trajectories, and scorer-side action descriptors remain fixed.
Fig. 9: Robustness across endpoint tolerances and evaluation states. (a) Equal-domain macro endpoint agreement, P(FDE≤τ) ; the vertical line marks SR@1. (b) Paired endpoint-agreement advantage over NoMaD+NWM-XL, with macro and domain-level estimates. (c) Per-state ATE reduction relative to NoMaD+NWM-XL. Grey points, violins, and diamonds show observations, distributions, and means, respectively. Positive values favor LiteNWM; percentages report per-state win rates. Shading and horizontal bars indicate 95% trajectory-cluster bootstrap CIs from 20,000 resamples; bands in (a,b) are pointwise. Each primary domain contains 100 states. SCAND-bag contains 48 states from six clusters and is a descriptive sensitivity cohort excluded from the macro average. DWM denotes the historical LiteNWM configuration; success measures offline endpoint agreement.
Fig. 10: Indoor and outdoor route progression. (a) Doorway traversal, corridor navigation, turning, and approach to a chair goal. (b) Outdoor navigation across paved areas and beside vegetation toward the visual target. Four observations in each row follow video order from left to right, followed by a cropped visual-goal reference. Trajectory overlays reproduce the source-video visualization. Frame spacing does not represent uniform elapsed time.
Institute of Automation, Chinese Academy of Sciences · School of Artificial Intelligence, University of Chinese Academy of Sciences · Beihang University