Organizations: National University of Singapore, Singapore · Singapore Management University, Singapore · The Chinese University of Hong Kong, Hong Kong SAR, China · Max Planck Institute for Plasma Physics, Germany
Training robust social-navigation policies requires simulators with diverse scene layouts, terrain, and human motion, but constructing such environments and specifying pedestrian behavior is costly. We propose an efficient pipeline that converts ordinary monocular walking videos directly into closed-loop social-navigation training environments in the policy's state space. Our key observation is that local social navigation primarily depends on two types of information: where the robot can traverse and how nearby pedestrians move. We therefore represent the static scene as a metric traversability map, which can be rigidly transformed under counterfactual robot motion, while directly replaying the pedestrian trajectories recovered from the video over time. This abstraction allows us to define the forward dynamics directly in the policy's state space and efficiently simulate counterfactual robot states without reconstructing or rendering photorealistic observations. The resulting policy achieves 81.2% success in the independent Arena benchmark, compared with 75.0% for the strongest baseline, and succeeds in 19/20 real-robot trials without policy fine-tuning. Project page: https://jiaming.im/VideoSocNav
Figures & tables
Fig. 1: From video to social-navigation simulation. Ordinary monocular walking videos become closed-loop training environments in the policy’s state space. The recorded world is decomposed into a static traversability map, which is rigidly transformed under counterfactual robot motion, and dynamic pedestrian trajectories, which are replayed over time. This lightweight abstraction enables efficient simulation and reinforcement learning without reconstructing or rendering photorealistic scenes.
Fig. 2: Overview. (a) First-person walking videos are processed by (b) pretrained perception models to recover a gravity-aligned metric traversability map and time-indexed pedestrian states. (c) These quantities define the state-space replay simulator: world time follows the recording, while the robot pose evolves independently from the policy commands. The recorded map and pedestrians are transformed into the robot frame at each step, enabling counterfactual motion away from the camera path, collision checking, and closed-loop training without rendering. (d) The policy encodes the egocentric traversability map, nearby pedestrian states, goal, and previous action with a lightweight transformer and outputs velocity commands. (e) At deployment, the same policy state is reconstructed online from onboard perception, allowing the video-trained policy to transfer without modification.
Fig. 3: Frames from three held-out source videos: an evening crowd on Istiklal Street (Istanbul), a rainy shopping street (London), and a snow-covered street at dusk (Tromsø). Ordinary first-person walking tours like these are the only input of the pipeline.
Train
Test
Source videos (distinct cities)
25 (12)
19 (17)
20 s episodes extracted
8,147
3,126
Episodes passing gates
6,075
2,155
Hours of usable video
33.8
12.0
Pedestrian tracks
91,261
29,846
Tracks per episode (mean)
15.0
13.8
TABLE I: Corpus statistics.
Policy
Observed
Unsupported path
Success
cell (%)
(m/ep.)
(%)
(%) ↑
Recorded walker
23
1.7
11.2
–
Ours (released)
21
2.3
15.2
92.6
Control (30 M)
21
2.2
14.4
91.4
Unknown-aware (30 M)
21
1.8
12.7
87.4
TABLE II: Observation coverage on held-out episodes: fraction of local-map cells supported by recorded measurements, distance traveled on unobserved space (m/episode and % of path length), and success rate.
Success rate (95% Wilson interval)
92.6% [91.4, 93.7]
Pedestrian collision
3.7%
Static-obstacle collision
1.9%
Timeout (18 s)
1.8%
SPL
0.860
Path efficiency (shortest / traveled)
0.928
Time to goal (s)
12.2
TABLE III: Closed-loop evaluation on the held-out episodes in the state-space simulator.
Variant
Succ. ↑ [95% CI]
SPL ↑
Ped. coll. ↓
Stat. coll. ↓
Pers. (s) ↓
Control
91.4 [90.1, 92.5]
0.858
4.4
1.5
0.19
IL only
55.3 [53.2, 57.4]
0.548
32.1
11.2
0.46
PPO from IL initialization
90.0 [88.6, 91.2]
0.859
5.4
1.2
0.45
No spawn jitter
90.4 [89.1, 91.6]
0.860
5.4
1.6
0.58
90 ∘ FOV map
88.9 [87.5, 90.1]
0.846
4.5
1.3
0.09
Proxemics reward
85.0 [83.4, 86.5]
0.823
11.5
2.1
0.42
TABLE IV: Ablations on held-out episodes. Succ. is success rate, Ped. coll. and Stat. coll. are pedestrian and static-obstacle collision rates, and Pers. is time per episode within 1.2 m of a pedestrian. Brackets give 95% Wilson confidence intervals. All variants are trained for 30 M steps.
Inputs
Outcome (%)
Efficiency
Social
Method
LiDAR
Plan
Success [95% CI] ↑
Coll. ↓ (ped/wall)
Timeout ↓
Time (s) ↓
SPL ↑
vˉ (m/s)
Personal (s) ↓
Intimate (s) ↓
CrowdSurfer [ 13 ]
✓
✓
42.5 [35.1, 50.2]
50.0 (47.5/2.5)
7.5
46.2
0.418
0.49
19.9
1.9
AttnGraph [ 11 ]
✓
35.0 [28.0, 42.7]
45.6 (11.2/34.4)
19.4
52.0
0.299
0.47
10.3
1.2
HEIGHT [ 12 ]
✓
✓
75.0 [67.8, 81.1]
25.0 (19.4/5.6)
0.0
46.4
0.714
0.47
11.8
3.5
Ours (SF peds)
58.1 [50.4, 65.5]
41.9 (39.4/2.5)
0.0
42.6
0.559
0.49
8.4
1.8
Ours (Trajectron++ peds)
70.0 [62.5, 76.6]
25.0 (20.0/5.0)
5.0
51.9
0.599
0.48
12.2
2.3
TABLE V: Arena/Isaac Sim results on 160 scenarios. Personal and intimate denote seconds per episode within 1.2 m and 0.45 m of a pedestrian, respectively.
Fig. 4: Arena results by crowd density: E, M and H are the easy (0–4 pedestrians), medium (5–9) and hard (10–14) scenarios. The advantage of our video-trained policy grows with density.
Fig. 5: (a) Unitree Go2 with RGB-D camera and LiDAR; (b, c) the indoor and outdoor test scenes.
Scene
Method
Success ↑
Mean time (s)
Mean path (m)
Indoor
Nav2-MPPI
5/10
28.6
9.64
Ours (Traj. peds)
7/10
32.1
10.96
Ours
9/10
31.1
10.14
Outdoor
Nav2-MPPI
3/10
18.9
7.74
Ours (Traj. peds)
8/10
16.1
8.29
Ours
10/10
17.9
8.59
TABLE VI: Real-robot results (10 trials per method and scene; means over successes).
School of Computer Science and Technology, Tongji University, Shanghai 201804, China · Department of Electronics and Information Engineering, Tongji University, Shanghai 201804, China · Shanghai Institute of Intelligent Science and Technology, Tongji University, Shanghai 201210, China