Social-WM: Safety-Aware Latent World Models for Robot Social Navigation
Authors: Zhihao Zheng, Mooi Choo Chuah
Organizations: Computer Science and Engineering department, P.C. Rossin College of Engineering and Applied Science, Lehigh University, Bethlehem, PA 18015, USA.
Safe social navigation requires a robot to anticipate not only the future consequences of its actions, but also whether a nominal action can actually be executed under surrounding physical and social constraints. We present Social-WM, an efficient latent world-model planning framework trained from egocentric RGB video sequences. Our key observation is that social-navigation experience contains a systematic discrepancy between the nominal action and the realizable action: a nominal forward action may be fully executed in free space, but needs to be constrained when heading towards a pedestrian or obstacle. Social-WM learns these safety-relevant consequences directly through action-conditioned future prediction, where the target is the actual observed future following each command. We further introduce a realizable inverse-dynamics objective that associates observed latent transitions with the action actually realized rather than the nominal one. At deployment, candidate actions are imagined through the latent world model, and the inverse dynamics model estimates their realizability; nominal--realizable discrepancy then provides a safety signal before execution. The learned dynamics and realizability model remain goal-independent and support both position- and image-goal navigation. On Social-HM3D, Social-WM achieves 63.77% success while reducing human collisions to 21.67%, and maintains strong performance under zero-shot transfer to Social-MP3D, without explicit pedestrian tracking, privileged human state, or online reinforcement learning.
Figures & tables
Fig. 1: Motivation of Social-WM. A nominal action that appears feasible from the current observation may become unsafe when executed. Social-WM predicts its future consequence and the corresponding realizable action before execution, allowing the planner to reject unsafe candidates and select a safer alternative.
Fig. 2: Overview of Social-WM. Candidate action chunks are proposed and imagined through the latent world model. Realizable-action inverse dynamics estimates the motion associated with each imagined transition, enabling candidate selection based on both goal progress and action realizability.
Fig. 3: Qualitative visualization of nominal and realizable futures. For each example, the top row shows the ground-truth observations obtained from the action chunk proposed by the CVAE, while the bottom row visualizes the future predicted by Social-WM. In the top-down view, the solid purple trajectory shows the proposed nominal motion and the dashed orange trajectory shows the ID-predicted realizable motion. The small yellow dot marks the pedestrian position, and the surrounding large yellow circle denotes its safety region; the black circle marks the robot’s starting position. When a nominal action conflicts with a pedestrian, Social-WM predicts a constrained future and the inferred realizable motion deviates from the nominal motion, exposing the unsafe candidate before execution.
Methods
Social-HM3D
Social-MP3D
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Rule-based
A* [ 19 ]
44.81
43.99
90.38
54.80
45.67
44.69
91.97
54.00
ORCA [ 6 ]
37.44
32.91
92.23
39.77
38.81
34.65
94.03
39.86
Reinforcement Learning-based
Habitat-official [ 20 ]
38.99
33.53
90.37
55.48
37.00
31.76
92.03
52.33
TABLE I: Social Navigation Results on Social-HM3D and Social-MP3D (zero-shot transfer from Social-HM3D). We report SR/SPL (success/efficiency), PSC (social compliance), and H-Coll (human collisions). Best and 2nd-best are in bold and underline .
Fig. 4: Qualitative comparison on Social-HM3D. Red boxes highlight unsafe interactions with insufficient pedestrian clearance. Social-WM instead slows or waits to maintain enough separation from pedestrians.
Goal Guidance
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Final Image
42.91
27.07
92.11
26.26
Image Subgoal
59.52
35.23
93.16
27.04
Final Position
63.77
49.62
94.37
21.67
Position Subgoal
68.65
58.72
93.28
22.09
TABLE II: Ablation of different goal settings on Social-HM3D. We compare final-goal and intermediate-subgoal settings under position and image goals.
Planner
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Latency(ms) ↓
MPC
56.08
32.72
91.34
31.33
139
CVAE
59.52
35.23
93.16
27.04
51
Diffusion Policy
58.84
40.15
94.08
31.49
126
TABLE III: Ablation of different planners on Social-HM3D.
Variant
Lid
Safety Eval.
SR ↑
SPL ↑
PSC ↑
H-Coll ↓
Latent WM
✗
✗
51.33
26.49
90.26
44.97
+ Realizable ID
✓
✗
53.08
30.16
90.34
42.12
Social-WM
✓
✓
56.08
32.72
91.34
31.33
TABLE IV: Effect of realizable-action learning on Social-HM3D.
Subset
Blocked Resid.
Free Resid.
AUROC ↑
All
0.97
0.26
0.961
Easy
0.99
0.09
0.988
Hard
0.99
0.10
0.995
TABLE V: nominal–realizable residual discrepancy. Mean residuals are reported for blocked and free transitions.
Representation
Rx2↑
Ry2↑
Ryaw2↑
Err@Min ↓
Regret ↓
LeWM-CLS
-0.36
-0.00
0.37
0.659
0.422
LeWM-Patch
-0.41
0.15
0.58
0.740
0.503
DINOv2-CLS
-0.10
0.16
0.31
0.756
0.520
DINOv2-Patch
0.12
0.29
0.42
0.320
0.083
TABLE VI: Latent representation ablation. Pose decodability and latent goal-ranking quality on frame pairs.
Fig. 5: Timeout failure example under persistent human blockage. The gray region denotes navigable floor and white regions denote non-navigable space. The green curve is the geodesic path to the red goal, while the dark blue trajectory shows the robot trajectory. Colored trajectories indicate the full-episode motion of pedestrians 2–4. A pedestrian repeatedly occupies the narrow passage along the nominal route, causing Social-WM to remain in the initial room. The model avoids unsafe motion near the pedestrian, but lacks long-horizon reasoning to determine whether to search for an alternative route, eventually resulting in a timeout.