cs.ROOct 6, 2026

RIWANav: Recursive World-Action Models with Self-Improvement for Urban Navigation

Authors: Jing Xie, Shouwei Ruan, Yubin Wang, Yuxiang Zhang, Haitao Yang, Songchang Jin, Dianxi Shi

Organizations: MoE Key Lab of Artificial Intelligence, Institute of AI, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China · Beihang University, Beijing, China · The Hong Kong University of Science and Technology, Hong Kong, China · Tsinghua University, Beijing, China · The University of Texas at Austin, Austin, USA · Intelligent Game and Decision Lab (IGDL), Beijing, China

Abstract

Long-horizon urban navigation requires sequential local decisions whose errors can compound over time. Imitation learning (IL) rarely learns from failures, while physical trial-and-error reinforcement learning (RL) is costly. Action-conditioned world models can provide imagined feedback by predicting visual consequences for candidate actions. However, a frozen world model may become less reliable as the policy evolves. In this paper, we introduce RIWANAV, a post-training framework that casts the coupled adaptation of a world model and an action model (policy) as task-specific recursive self-improvement (RSI). Each cycle alternates two updates. The world model evaluates policy actions through imagined outcomes, providing comparative feedback for group-relative policy optimization (GRPO). The improved policy then constructs a grounded self-curriculum, selecting expert-consistent action-video pairs by behavioral novelty and prediction error. The refined world model supplies feedback for the next policy update, closing the recursive self-improvement loop. Experiments show that RIWANAV outperforms training baselines and prior methods, validating the proposed recursive self-improvement loop between the policy and world model. Real-world trials further demonstrate its practical applicability.

Figures & tables

Explore similar work

CardsList
  1. NavWAM: A Navigation World Action Model for Goal-Conditioned Visual Navigation

    Jun 11, 2026Daichi Azuma, Taiki Miyanishi, Koya Sakamoto +6Efficient World-Action ModelObject Goal Navigation

  2. NavOL: Navigation Policy with Online Imitation Learning

    May 12, 2026Xiaofei Wei, Chun Gu, Li ZhangNavigation

  3. SC2^{2}-WM: A Self-Correcting World Model with Closed-Loop Feedback for Vision-and-Language Navigation in Continuous Environments

    Aug 1, 2026Xuan Yao, Yuze Zhu, Junyu Gao +2Vision-Language NavigationObject Goal Navigation