Social-WM: Safety-Aware Latent World Models for Robot Social Navigation
Authors: Zhihao Zheng, Mooi Choo Chuah
Organizations: Computer Science and Engineering department, P.C. Rossin College of Engineering and Applied Science, Lehigh University, Bethlehem, PA 18015, USA.
Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.
Figures & tables
Fig. 1: Motivation of Social-WM. A nominal action that appears feasible from the current observation may become unsafe when executed. Social-WM predicts its future consequence and the corresponding realizable action before execution, allowing the planner to reject unsafe candidates and select a safer alternative.
Fig. 2: Overview of Social-WM. Candidate action chunks are proposed and imagined through the latent world model. Realizable-action inverse dynamics estimates the motion associated with each imagined transition, enabling candidate selection based on both goal progress and action realizability.
Fig. 3: Qualitative visualization of nominal and realizable futures. For each example, the top row shows the ground-truth observations obtained from the action chunk proposed by the CVAE, while the bottom row visualizes the future predicted by Social-WM. In the top-down view, the solid purple trajectory shows the proposed nominal motion and the dashed orange trajectory shows the ID-predicted realizable motion. The small yellow dot marks the pedestrian position, and the surrounding large yellow circle denotes its safety region; the black circle marks the robot’s starting position. When a nominal action conflicts with a pedestrian, Social-WM predicts a constrained future and the inferred realizable motion deviates from the nominal motion, exposing the unsafe candidate before execution.
Methods
Social-HM3D
Social-MP3D
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Rule-based
A* [ 19 ]
44.81
43.99
90.38
54.80
45.67
44.69
91.97
54.00
ORCA [ 6 ]
37.44
32.91
92.23
39.77
38.81
34.65
94.03
39.86
Reinforcement Learning-based
Habitat-official [ 20 ]
38.99
33.53
90.37
55.48
37.00
31.76
92.03
52.33
TABLE I: Social Navigation Results on Social-HM3D and Social-MP3D (zero-shot transfer from Social-HM3D). We report SR/SPL (success/efficiency), PSC (social compliance), and H-Coll (human collisions). Best and 2nd-best are in bold and underline .
Fig. 4: Qualitative comparison on Social-HM3D. Red boxes highlight unsafe interactions with insufficient pedestrian clearance. Social-WM instead slows or waits to maintain enough separation from pedestrians.
Goal Guidance
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Final Image
42.91
27.07
92.11
26.26
Image Subgoal
59.52
35.23
93.16
27.04
Final Position
63.77
49.62
94.37
21.67
Position Subgoal
68.65
58.72
93.28
22.09
TABLE II: Ablation of different goal settings on Social-HM3D. We compare final-goal and intermediate-subgoal settings under position and image goals.
Planner
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Latency(ms) ↓
MPC
56.08
32.72
91.34
31.33
139
CVAE
59.52
35.23
93.16
27.04
51
Diffusion Policy
58.84
40.15
94.08
31.49
126
TABLE III: Ablation of different planners on Social-HM3D.
Variant
Lid
Safety Eval.
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Latent WM
✗
✗
51.33
26.49
90.26
44.97
+ Realizable ID
✓
✗
53.08
30.16
90.34
42.12
Social-WM
✓
✓
56.08
32.72
91.34
31.33
TABLE IV: Effect of realizable-action learning on Social-HM3D.
Subset
Blocked Resid.
Free Resid.
AUROC ↑
All
0.97
0.26
0.961
Easy
0.99
0.09
0.988
Hard
0.99
0.10
0.995
TABLE V: nominal–realizable residual discrepancy. Mean residuals are reported for blocked and free transitions.
Representation
Rx2↑
Ry2↑
Ryaw2↑
Err@Min ↓
Regret ↓
LeWM-CLS
-0.36
-0.00
0.37
0.659
0.422
LeWM-Patch
-0.41
0.15
0.58
0.740
0.503
DINOv2-CLS
-0.10
0.16
0.31
0.756
0.520
DINOv2-Patch
0.12
0.29
0.42
0.320
0.083
TABLE VI: Latent representation ablation. Pose decodability and latent goal-ranking quality on frame pairs.
Fig. 5: Timeout failure example under persistent human blockage. The gray region denotes navigable floor and white regions denote non-navigable space. The green curve is the geodesic path to the red goal, while the dark blue trajectory shows the robot trajectory. Colored trajectories indicate the full-episode motion of pedestrians 2–4. A pedestrian repeatedly occupies the narrow passage along the nominal route, causing Social-WM to remain in the initial room. The model avoids unsafe motion near the pedestrian, but lacks long-horizon reasoning to determine whether to search for an alternative route, eventually resulting in a timeout.
Social navigation requires the robot to reason and respond in complex real-world environments. While recent works attempt to incorporate human-level intelligence into robot planning using large Vision-Language Models (VLMs), end-to-end frameworks often create an unpredictable black-box, and existing instruction-following methods are not designed for full autonomy. To bridge this gap, we present G2-Nav, a novel framework that grounds abstract social reasoning and guards safe real-world deployment. Instead of asking the VLM for direct planning decisions, G2-Nav translates its semantic reasoning into a vision-language costmap with reliability and interpretability. The VLM evaluates traversable regions and social agents from open-set perception, mapping social context into the costmap. To improve real-world robustness, the VLM performs semantic verification on upstream tracking, and we introduce a high-frequency safety check to guard against system latency prior to trajectory generation. We demonstrate through real-world experiments that G2-Nav delivers safe, efficient, and socially compliant autonomous navigation in unstructured environments. Code is available at https://github.com/centiLinda/G2-Nav.
Yuwen Liao, Yihang Lan, Yizhuo Yang +4
School of Electrical and Electronic Engineering Nanyang Technological University Singapore
Safe social navigation requires robots to distinguish people from ordinary obstacles and to react before danger becomes imminent. We show that pretrained Vision-Language-Action (VLA) models already encode pedestrian-object distinctions and future collision signals in their internal representations, but behavior cloning fails to translate these signals into socially appropriate actions. To address this mismatch, we propose SALSA, a two-stage annotation-free post-training framework: (1) social behavioral alignment bridges intermediate-layer social features to the action head and trains on counterfactual human-object scene pairs to break visual saliency shortcuts; (2) temporal safety alignment provides automatically generated future-risk supervision to enable anticipatory collision avoidance. On SCAND and real-world deployment, SALSA reduces near-collisions by 86.4% and improves social counterfactual accuracy from 53% to 93%, demonstrating that safer social navigation can be achieved by teaching VLA policies to act on representations they already possess. These results show that pretrained VLA policies can be adapted for safer social navigation by better aligning their latent representations with action generation.
Qingzi Wang, Xiyang Wu, Guangyao Shi +3
University of Maryland · University of Southern California
World models such as DINO-WM and LeWM specify the goal with an image, which is difficult to obtain in advance for novel tasks. We present the Grounded World Model (GWM), a latent world model that enables zero-shot planning in the real world from language goals alone. Given a candidate action sequence and the current observation, GWM predicts the future in the visual space of a pretrained video-language embedding model. The frozen readout of this embedding model maps this imagined future and the task description into the same embedding space, where their negative cosine similarity serves as the planning cost. Training GWM requires only offline and task-agnostic video-action pairs and no language labels. In simulated experiments on WISER, planning with GWM, which executes the candidate action of lowest cost, solves 87% of 288 tasks with unseen instructions and visual signals, while ten fine-tuned VLAs average 22%. We then scale GWM up with real robot data, and use it for zero-shot planning in realistic simulation and real scenes. In the IsaacSim evaluation, planning with GWM completes all 70 trials across 14 tasks that require reasoning over referring expressions, matching a modular planner grounded by a frontier VLM, while pi0.5 reaches 37/70. Deployed on a real Franka, the same stack completes 55/60 separately evaluated pick-and-place sub-tasks, comparable to the modular planner's 52/60, with the full system running locally on a single consumer GPU. Project website: https://quanyili.github.io/gwm-wiser/.