cs.CVMar 23, 2026

Language-Conditioned World Modeling for Visual Navigation

Authors: Yifei Dong, Fengyi Wu, Yilong Dai, Lingdong Kong, Guangyu Chen, Yetong Sha, Qiyu Hu, Feng Liu, +3 more

Organizations: University of Washington · National University of Singapore · Clemson University · Drexel University · Microsoft Research

Abstract

Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal images, the agent must rely on language to shape its perception and continuous control. We introduce the LCVN Dataset, a benchmark of 39,016 trajectories and 117,048 human-verified instructions spanning diverse environments and instruction styles. Building on this benchmark, we study two complementary paradigms: (i) latent-imagination policy learning, in which a diffusion-based world model (LCVN-WM) imagines future observations and an actor-critic agent (LCVN-AC) learns its policy entirely within the imagined latent space; and (ii) unified autoregressive prediction, in which a single multimodal backbone (LCVN-Uni) jointly predicts actions and observations in one forward pass over a shared token sequence. Experiments show that two paradigms offer complementary strengths: latent imagination produces more temporally coherent rollouts, whereas unified prediction generalizes better to unseen environments. Targeted ablations further isolate the contributions of language guidance, conditioning signals, and instruction style, clarifying when language grounding versus dynamics modeling is the performance bottleneck. Together, these findings position LCVN as a testbed for studying how language, imagination, and decision-making interact in embodied agents.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FutureNav: Unified World-Action Modeling for Vision-and-Language Navigation

    Jun 29, 2026Lingfeng Zhang, Zeying Gong, Xiaoshuai Hao +7Vision-Language NavigationNavigation

  2. NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

    Sep 30, 2026Haoxiang Shi, Zaijing Li, Muhe Ding +3Vision-Language NavigationNavigation

  3. Uni-LaViRA: Language-Vision-Robot Actions Translation for Unified Embodied Navigation

    May 26, 2026Hongyu Ding, Sizhuo Zhang, Ziming Xu +13Object Goal NavigationHeterogeneous Robot Teams