Goal-conditioned visual navigation has been a long-standing testbed for embodied AI. We study a natural language-conditioned variant, language-conditioned visual navigation (LCVN), in which an embodied agent must follow a natural language instruction given only an initial egocentric observation. Without access to goal images, the agent must rely on language to shape its perception and continuous control. We introduce the LCVN Dataset, a benchmark of 39,016 trajectories and 117,048 human-verified instructions spanning diverse environments and instruction styles. Building on this benchmark, we study two complementary paradigms: (i) latent-imagination policy learning, in which a diffusion-based world model (LCVN-WM) imagines future observations and an actor-critic agent (LCVN-AC) learns its policy entirely within the imagined latent space; and (ii) unified autoregressive prediction, in which a single multimodal backbone (LCVN-Uni) jointly predicts actions and observations in one forward pass over a shared token sequence. Experiments show that two paradigms offer complementary strengths: latent imagination produces more temporally coherent rollouts, whereas unified prediction generalizes better to unseen environments. Targeted ablations further isolate the contributions of language guidance, conditioning signals, and instruction style, clarifying when language grounding versus dynamics modeling is the performance bottleneck. Together, these findings position LCVN as a testbed for studying how language, imagination, and decision-making interact in embodied agents.
Figures & tables
Figure 1 : Language-Conditioned Visual Navigation (LCVN). Given only an initial egocentric observation and a language instruction, the agent generates the entire future trajectory without environmental feedback, imagining intermediate states ( ➀ – ➃ ) along the described route. LCVN-WM + LCVN-AC imagines future latent observations and selects actions within the imagined latent space, while LCVN-Uni jointly predicts the next action and observation in a single autoregressive pass, conditioned on both visual and language inputs.
Figure 2 : Language-conditioned world model (LCVN-WM) and associated LCVN-AC. Phase 1: LCVN-WM is trained with Diffusion Forcing ( Chen et al., 2024 ; Song et al., 2025 ) to predict future latent observations from noisy context latents at independent noise levels, conditioned on actions, instruction, and time shift. Phase 2: LCVN-AC is trained in LCVN-WM’s latent space, aligning expert and learner plans via KL divergence and optimizing an actor–critic objective with intrinsic rewards that measure agreement between predicted and expert latent rollouts. Inference: LCVN-WM predicts the next latent observation, on which LCVN-AC conditions to generate the next action.
Figure 3 : LCVN-Uni architecture. LCVN-Uni unifies navigation planning and world modeling within an autoregressive MLLM backbone. Actions a^t , instructions I , and observations os,o^t are tokenized by bin, BPE, and VQ tokenizers, respectively, then fused into a unified sequence.
Figure 4 : Qualitative Comparisons on LCVN val seen split across NWM and LCVN agents. LCVN agents exhibit stronger sensitivity to directional changes and are less prone to losing landmarks during world modeling.
Methods
Context Size
Validation Seen
Validation Unseen
Test
ATE ↓
RPE ↓
SR ↑
ATE ↓
RPE ↓
SR ↑
ATE ↓
RPE ↓
SR ↑
GNM ( Shah et al., 2022 ) (lang)
4
1.18
0.39
0.21
2.72
0.91
0.10
1.89
0.55
0.16
NoMaD ( Sridhar et al., 2024 ) (lang)
4
1.08
0.36
0.24
2.61
0.88
0.11
1.78
0.53
0.18
Diamond ( Alonso et al., 2024 ) + LCVN-AC
4
1.35
0.42
0.18
2.84
0.95
0.09
2.05
0.58
0.14
NWM ( Bar et al., 2025 ) + LCVN-AC
4
0.75
0.24
0.28
2.17
0.73
0.12
1.33
0.42
0.21
NWM ( Bar et al., 2025 ) (lang) + LCVN-AC
4
0.72
0.26
0.29
2.21
0.70
0.13
1.31
0.44
0.22
Table 1 : Comparison with SOTA Methods upon Language-Conditioned Visual Navigation on LCVN val seen, val unseen and test splits with SR, ATE, and RPE. We equip two LCVN agents with context sizes k set to 1, 2, 4.
Methods
Context Size
Single-frame Generation
Long-horizon Generation @8
SSIM ↑
PSNR ↑
LPIPS ↓
DreamSIM ↓
SSIM @8↑
PSNR @8↑
LPIPS @8↓
DreamSIM @8↓
Diamond ( Alonso et al., 2024 )
4
0.315
9.850
0.427
0.135
0.114
4.526
0.679
0.283
NWM ( Bar et al., 2025 )
4
0.370
11.425
0.314
0.096
0.181
6.057
0.629
0.196
NWM ( Bar et al., 2025 ) (lang)
4
0.382
11.612
0.309
0.094
0.176
6.041
0.605
0.191
LCVN-Uni
1
0.398
12.881
0.306
0.076
0.201
7.057
0.466
0.128
LCVN-Uni (w/o ins)
2
0.387
12.642
0.319
0.082
0.192
6.874
0.508
0.135
Table 2 : Comparison with SOTA methods on imagination performance (single-frame / long-horizon) on LCVN val seen split. We equip two LCVN agents with context sizes k set to 1, 2, 4. Best results are in bold.
Navigation
Imagination
Language
Action
Time
ATE ↓
RPE ↓
SR ↑
SSIM ↑
DreamSIM ↓
SSIM @8↑
DreamSIM @8↓
×
×
✓
1.82
0.63
0.12
0.251
0.627
0.095
0.798
✓
×
×
1.12
0.35
0.19
0.341
0.137
0.187
0.226
×
✓
×
0.54
0.22
0.31
0.388
0.112
0.229
0.169
✓
×
✓
0.89
0.31
0.22
0.352
0.142
0.201
0.214
✓
✓
×
0.37
0.14
0.41
0.422
0.081
0.275
0.138
Table 3 : Ablations of language, action and time conditioning on navigation and imagination performance of LCVN-WM ( k =4) on LCVN val seen split. ✓ denotes with and × denotes without. Best results are in bold.
Figure 5 : Qualitative comparisons of language guidance showcase both baselines better preserve semantic consistency across states with language guidance.
Navigation
Imagination
Method
Ins. Style
ATE ↓
RPE ↓
SR ↑
SSIM ↑
DreamSIM ↓
SSIM @8↑
DreamSIM @8↓
LCVN-Uni
Concise
0.37
0.13
0.41
0.421
0.076
0.218
0.121
Intricate
0.39
0.11
0.38
0.415
0.074
0.211
0.123
Landmark.
0.32
0.10
0.47
0.433
0.069
0.225
0.116
Concise
0.35
0.12
0.42
0.435
0.078
0.293
0.127
Intricate
0.36
0.14
0.40
0.429
0.081
0.285
0.128
Table 4 : Ablations of instruction style for LCVN-Uni ( k =2) and LCVN-WM ( k =4) on LCVN val seen split.
Navigation
Imagination
Encoding Space
ATE ↓
RPE ↓
SR ↑
SSIM ↑
DreamSIM ↓
SSIM @8↑
DreamSIM @8↓
Pixel
0.42
0.16
0.36
0.381
0.095
0.257
0.150
Latent
0.34
0.12
0.43
0.435
0.078
0.293
0.127
Table 5 : Pixel vs. Latent space on the performance of LCVN-WM ( k =4) on LCVN val seen split.
Method
External Data
SSIM ↑
DreamSIM ↓
SSIM @8↑
DreamSIM @8↓
NWM
×
0.180
0.325
0.102
0.543
NWM
✓ (Ego4D)
0.197
0.318
0.115
0.528
LCVN-WM
×
0.302
0.239
0.167
0.376
LCVN-WM
✓ (Ego4D)
0.313
0.226
0.184
0.357
Table 6 : External data scaling on imagination performance of NWM and LCVN-WM ( k =4) on LCVN val unseen. ✓ denotes training with additional Ego4D data. Best results are in bold.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : Overview of annotation pipeline of the LCVN dataset.
Trajectory example 1 (concise style). Walk straight ahead and stop at the intersection.
Trajectory example 1 (intricate style). You are on a wide sidewalk lined with trees on both sides. A pedestrian is approaching from the front. Continue walking until a red stall appears on your right, with three people gathered around it, and stop in front of the stall .
Trajectory example 1 (landmark-grounded style). Walk along the sidewalk until you reach a red stall ahead, and stop in front of it.
Trajectory example 2 (concise style). Walk straight ahead, turn left at the first corner, and stop along the wall.
Trajectory example 2 (intricate style). Proceed straight along the gray corridor, passing a row of white cabinets and a glass-walled office on your left where people appear to be working. Continue forward until you reach a gray door , with a small white trash bin positioned to its left. Then turn left into another gray corridor and stop along the wall on your left.
Trajectory example 2 (landmark-grounded style). Walk straight down the corridor, passing a glass-walled office on your left. Continue until you reach a gray door , then turn left and stop along the wall on your left.
Appendix
Table A1 : Multi-style Instruction Samples from the LCVN Dataset. Text in purple denotes directional guidance, while text in blue marks landmark references. Each trajectory in LCVN is paired with three distinct instruction styles (concise, intricate and landmark-grounded), enabling a comprehensive evaluation of model generalization across varying levels of specificity and emphasis ( Kolagar and Zarcone, 2024 ) .
Figure A2 : Prompt design example of LCVN-Uni (context size = 1).
NWM
LCVN-Uni
LCVN-WM
LCVN-WM (+Distillation.)
LCVN-WM (+Quant. 4-bit)
11.2
20.5
6.4
0.6
0.1 (est. ( Frantar et al., 2022 ) )
Appendix
Table A2: Inference time (seconds) per step, averaged on the LCVN dataset val seen split, including both action prediction and imagination. For NWM and LCVN-WM, the times include the latency of LCVN-AC’s action prediction.
Navigation
Imagination
Method
Speed
ATE ↓
RPE ↓
SR ↑
SSIM ↑
DreamSIM ↓
SSIM @8↑
DreamSIM @8↓
LCVN-Uni (Inter.)
1 ×
0.34
0.12
0.44
0.438
0.074
0.201
0.115
LCVN-Uni
1.3 ×
0.36
0.11
0.42
0.423
0.072
0.218
0.119
Appendix
Table A3 : Interleave vs. Predict Both on navigation and imagination performance of LCVN-Uni ( k =2) on LCVN val seen.