cs.ROSep 30, 2026

Social-WM: Safety-Aware Latent World Models for Robot Social Navigation

Authors: Zhihao Zheng, Mooi Choo Chuah

Organizations: Computer Science and Engineering department, P.C. Rossin College of Engineering and Applied Science, Lehigh University, Bethlehem, PA 18015, USA.

Abstract

Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.

Figures & tables

Explore similar work

CardsList
  1. G2-Nav: Grounded and Guarded Vision-Language Costmaps for Robot Social Navigation

    Jul 18, 2026Yuwen Liao, Yihang Lan, Yizhuo Yang +4Robot NavigationAutonomous Navigation

  2. Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

    Jun 9, 2026Qingzi Wang, Xiyang Wu, Guangyao Shi +3Safe NavigationCollision Avoidance

  3. Grounded World Model: Latent Planning with Language Goals

    Apr 13, 2026Quanyi Li, Lan Feng, Haonan Zhang +4World ModelsWorld