Organizations: MoE Key Lab of Artificial Intelligence, Institute of AI, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Beihang University, Beijing, China · The Hong Kong University of Science and Technology, Hong Kong, China · Tsinghua University, Beijing, China · The University of Texas at Austin, Austin, USA · Intelligent Game and Decision Lab (IGDL), Beijing, China
Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.
Figures & tables
Fig. 1: From imitation learning to recursive world–action models self-improvement. (a) Imitation learning trains the policy from expert trajectories but provides limited supervision for policy-induced failures. (b) A frozen world model enables learning from imagined outcomes, but its predictions become misaligned with the evolving policy. (c) RIWANav closes the training loop: imagined outcomes improve the policy, while grounded policy-induced samples refine the world model for subsequent policy learning.
Fig. 2: Recursive self-improvement loop in RIWANav . (1) The current world model imagines action consequences that improve the policy. (2) The improved policy constructs a grounded self-curriculum that improves the next world model. Repeating these updates recursively changes both components and the feedback exchanged between them. The flame marks the component optimized in each phase.
Test Seen
Training framework
SR (%) ↑
SPL (%) ↑
NE ↓
nDTW (%) ↑
SDTW (%) ↑
OSR (%) ↑
Imitation learning
70.91
70.55
3.8427
83.81
69.16
71.18
PPO [ 35 ]
72.04±0.06
71.40±0.09
3.8634±0.0136
84.02±0.08
70.29±0.06
72.36±0.07
GRPO [ 8 ]
75.11±0.75
74.45±0.75
3.7012±0.1803
86.40±0.47
73.37±0.76
75.32±0.81
PPO-WM
78.90±0.18
78.25±0.19
2.9530±0.0647
88.06±0.13
77.36±0.18
79.39±0.20
GRPO-WM
83.35±1.25
81.27±1.38
2.7059±0.1858
90.55±0.73
81.66±1.29
84.31±1.17
TABLE I: Navigation performance on the UrbanNav test-seen and test-unseen splits. Values are reported as mean ± standard deviation.
Test Seen
Category
Method
SR (%) ↑
SPL (%) ↑
NE ↓
nDTW (%) ↑
SDTW (%) ↑
OSR (%) ↑
Imitation Learning
UrbanNav [ 2 ]
60.36±5.39
59.75±5.34
5.3958±0.6689
76.86±3.50
58.22±5.34
60.48±5.36
NoMaD [ 11 ]
82.35±0.78
77.59±0.55
2.5532±0.0608
90.15±0.19
77.97±0.79
82.47±0.76
OmniVLA [ 12 ]
77.49±1.66
77.25±1.66
3.1856±0.1415
90.25±0.47
73.63±1.72
77.66±1.62
WM-based
Fast-WAM [ 6 ]
58.48±0.75
39.57±0.18
8.4899±0.0723
64.77±0.20
50.62±0.56
62.07±0.31
NWM [ 4 ]
71.26±0.21
70.67±0.21
3.7117±0.0200
83.88±0.05
69.45±0.20
71.63±0.20
TABLE II: Comparison with representative prior approaches on the UrbanNav test-seen and test-unseen splits. Values are reported as mean ± standard deviation.
WM
Policy-induced
Grounded
SR (%) ↑
SPL (%) ↑
update
replay
self-curriculum
×
×
×
80.31±0.94
76.09±1.05
✓
×
×
78.23±1.91
74.72±1.92
✓
✓
×
84.83±0.33
80.11±0.32
✓
✓
✓
88.25±0.43
82.59±0.25
TABLE III: Ablation study on the UrbanNav test-unseen split. Values are reported as mean ± standard deviation.
World model
Evaluation
L1↓
LPIPS ↓
actions
Initial ϕ0
Expert
0.1396±0.0032
0.4532±0.0065
Policy-induced
0.1521±0.0042
0.4539±0.0059
Static expert
Expert
0.1352±0.0034
0.4265±0.0079
Policy-induced
0.1505±0.0044
0.4314±0.0067
RIWANav ϕK
Expert
0.1329±0.0034
0.4203±0.0079
TABLE IV: Endpoint world-model prediction quality on expert actions and actions generated by the improved policy.
Fig. 3: World-model prediction quality across recursive rounds. The left and right panels show L1 and LPIPS on two fixed evaluation sets: expert actions and actions generated by the improved policy. Points are means over 128 samples and error bars denote standard error. Round 0 uses the initial world model, whereas the final round K=19 uses the final RIWANav world model. Lower is better.
Fig. 4: Qualitative evidence for recursive world-model improvement on actions generated by the improved policy. All three models receive the same observation, language goal, and action; the aligned recorded future is shown as reference. The rows depict night traffic, a bicycle-lined sidewalk, and a rainy sidewalk. In these cases, the RIWANav world model after recursive updates better preserves traversable corridors and scene layout, with fewer long-horizon artifacts.
Fig. 5: Real-world urban navigation with RIWANav . Orange dashed arrows indicate schematic paths, and green boxes highlight the targets.
State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences (CASIA) · School of Artificial Intelligence, University of Chinese Academy of Sciences (UCAS) · HiThink Research, China +1