Organizations: MoE Key Lab of Artificial Intelligence, Institute of AI, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Beihang University, Beijing, China · The Hong Kong University of Science and Technology, Hong Kong, China · Tsinghua University, Beijing, China · The University of Texas at Austin, Austin, USA · Intelligent Game and Decision Lab (IGDL), Beijing, China
Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.
Figures & tables
Fig. 1: From imitation learning to recursive world–action models self-improvement. (a) Imitation learning trains the policy from expert trajectories but provides limited supervision for policy-induced failures. (b) A frozen world model enables learning from imagined outcomes, but its predictions become misaligned with the evolving policy. (c) RIWANav closes the training loop: imagined outcomes improve the policy, while grounded policy-induced samples refine the world model for subsequent policy learning.
Fig. 2: Recursive self-improvement loop in RIWANav . (1) The current world model imagines action consequences that improve the policy. (2) The improved policy constructs a grounded self-curriculum that improves the next world model. Repeating these updates recursively changes both components and the feedback exchanged between them. The flame marks the component optimized in each phase.
Test Seen
Training framework
SR (%) ↑
SPL (%) ↑
NE ↓
nDTW (%) ↑
SDTW (%) ↑
OSR (%) ↑
Imitation learning
70.91
70.55
3.8427
83.81
69.16
71.18
PPO [ 35 ]
72.04±0.06
71.40±0.09
3.8634±0.0136
84.02±0.08
70.29±0.06
72.36±0.07
GRPO [ 8 ]
75.11±0.75
74.45±0.75
3.7012±0.1803
86.40±0.47
73.37±0.76
75.32±0.81
PPO-WM
78.90±0.18
78.25±0.19
2.9530±0.0647
88.06±0.13
77.36±0.18
79.39±0.20
GRPO-WM
83.35±1.25
81.27±1.38
2.7059±0.1858
90.55±0.73
81.66±1.29
84.31±1.17
TABLE I: Navigation performance on the UrbanNav test-seen and test-unseen splits. Values are reported as mean ± standard deviation.
Test Seen
Category
Method
SR (%) ↑
SPL (%) ↑
NE ↓
nDTW (%) ↑
SDTW (%) ↑
OSR (%) ↑
Imitation Learning
UrbanNav [ 2 ]
60.36±5.39
59.75±5.34
5.3958±0.6689
76.86±3.50
58.22±5.34
60.48±5.36
NoMaD [ 11 ]
82.35±0.78
77.59±0.55
2.5532±0.0608
90.15±0.19
77.97±0.79
82.47±0.76
OmniVLA [ 12 ]
77.49±1.66
77.25±1.66
3.1856±0.1415
90.25±0.47
73.63±1.72
77.66±1.62
WM-based
Fast-WAM [ 6 ]
58.48±0.75
39.57±0.18
8.4899±0.0723
64.77±0.20
50.62±0.56
62.07±0.31
NWM [ 4 ]
71.26±0.21
70.67±0.21
3.7117±0.0200
83.88±0.05
69.45±0.20
71.63±0.20
TABLE II: Comparison with representative prior approaches on the UrbanNav test-seen and test-unseen splits. Values are reported as mean ± standard deviation.
WM
Policy-induced
Grounded
SR (%) ↑
SPL (%) ↑
update
replay
self-curriculum
×
×
×
80.31±0.94
76.09±1.05
✓
×
×
78.23±1.91
74.72±1.92
✓
✓
×
84.83±0.33
80.11±0.32
✓
✓
✓
88.25±0.43
82.59±0.25
TABLE III: Ablation study on the UrbanNav test-unseen split. Values are reported as mean ± standard deviation.
World model
Evaluation
L1↓
LPIPS ↓
actions
Initial ϕ0
Expert
0.1396±0.0032
0.4532±0.0065
Policy-induced
0.1521±0.0042
0.4539±0.0059
Static expert
Expert
0.1352±0.0034
0.4265±0.0079
Policy-induced
0.1505±0.0044
0.4314±0.0067
RIWANav ϕK
Expert
0.1329±0.0034
0.4203±0.0079
TABLE IV: Endpoint world-model prediction quality on expert actions and actions generated by the improved policy.
Fig. 3: World-model prediction quality across recursive rounds. The left and right panels show L1 and LPIPS on two fixed evaluation sets: expert actions and actions generated by the improved policy. Points are means over 128 samples and error bars denote standard error. Round 0 uses the initial world model, whereas the final round K=19 uses the final RIWANav world model. Lower is better.
Fig. 4: Qualitative evidence for recursive world-model improvement on actions generated by the improved policy. All three models receive the same observation, language goal, and action; the aligned recorded future is shown as reference. The rows depict night traffic, a bicycle-lined sidewalk, and a rainy sidewalk. In these cases, the RIWANav world model after recursive updates better preserves traversable corridors and scene layout, with fewer long-horizon artifacts.
Fig. 5: Real-world urban navigation with RIWANav . Orange dashed arrows indicate schematic paths, and green boxes highlight the targets.
Goal-conditioned visual navigation requires a robot to act under partial observability by anticipating how its motion will change the future egocentric view and whether that change brings it closer to the goal. Navigation world models provide such visual foresight, but they remain prediction modules that require an external planner to convert predicted futures into closed-loop control. We propose Navigation World Action Model (NavWAM), a diffusion-transformer policy that turns navigation world-model prediction into executable action by representing future observations, goal-progress values, and action chunks in a shared latent sequence. By learning future prediction jointly with the action and value targets that determine closed-loop behavior, NavWAM makes visual foresight directly usable for robot control. We build NavWAM through simulation pretraining and real-robot adaptation, and evaluate it on image-goal navigation against planning-based world models and a representative direct navigation policy. Across offline benchmarks and closed-loop real-robot deployment, NavWAM improves over planning-based world-model baselines in our evaluations while using the default policy mode without CEM-style action search. Project page: https://dachii-azm.github.io/navwam/
Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6
The University of Tokyo · National Institute of Informatics · AIRoA +1
Learning robust navigation policies remains a core challenge in robotics. Offline imitation learning suffers from distribution shift and compounding errors at rollout, while reinforcement learning requires reward engineering and learns inefficiently. In this paper, we propose NavOL, an online imitation learning paradigm that interacts with a simulator and updates itself using expert demonstrations gathered online. Built upon a pretrained navigation diffusion policy that maps local observations to future waypoints, NavOL trains in a rollout update loop: during rollout, the policy acts in the simulator and queries a global planner which has privileged access to the global environment for the optimal path segment as ground truth trajectory labels; during update, the policy is trained on the online collected observation trajectory pairs. This online imitation loop removes the need for reward design, improves learning efficiency, and mitigates distribution shift by training on the policy own explored rollouts. Built on IsaacLab with fast, high-fidelity parallel rendering and domain randomization of camera pose and start-goal pairs, our system scales across 50 scenes on 8 RTX 4090 GPUs, collecting over 2,000 new trajectories per hour, each averaging more than 400 steps. We also introduce an indoor visual navigation benchmark with predefined start and goal positions for zero-shot generalization. Extensive evaluations on simulation benchmarks, including the NavDP benchmark and our proposed benchmark, as well as carefully designed real-world experiments, demonstrate the effectiveness of NavOL, showing consistent performance gains in online imitation learning.
Xiaofei Wei, Chun Gu, Li Zhang
School of Data Science, Fudan University · 2Shanghai Innovation Institute.
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC2-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. Our method derives feedback from world-model foresight to perform state-level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world-aware adaptation, which enables model-level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN-CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at https://github.com/sunrise-ikun/SC2_WM.
Xuan Yao, Yuze Zhu, Junyu Gao +2
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA) · School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS) · HiThink Research, China +1