Autonomous social navigation requires balancing efficiency, physical safety, and social compliance. Reinforcement Learning (RL) methods provide a viable and effective solution but often rely on unrealistic assumptions, such as the knowledge of humans' position and velocity. In this paper, we introduce JESSI (JAX-based E2E Safe Social Interpretable navigation), a lightweight end-to-end RL framework that maps raw LiDAR scans directly to kinematically feasible control commands. JESSI enhances safety via Dirichlet-parameterized continuous action spaces and deterministic bounding, while an integrated attention-based perception module extracts probabilistic human states for interpretable, socially aware decision-making. Through extensive simulations and real-world deployment on a differential-drive robot, we demonstrate that jointly optimizing the RL policy with a supervised perception signal in a multi-task paradigm enhances social behavior. Ultimately, JESSI is able to balance high navigation success rates and superior social behaviors compared to state-of-the-art baselines.
Figures & tables
Fig. 1: Overview of the proposed method. JESSI operates directly from raw LiDAR scans and leverages a perception module to extract probabilistic human states (human-centered bivariate gaussians) and plan socially-aware and kinematically feasible control commands through a scene attention mechanism and Dirichlet distributions.
Fig. 2: Triangular action space of the differential drive with linear velocity v∈[0,vˉ] and angular velocity ω∈[−ωˉ,ωˉ] .
Fig. 3: Complete architecture of the JESSI policy. The top part in blue represents the perception module while the bottom part in red depicts the actor-critic module. The inputs of the overall network are depicted in green blocks. In order, the dimensionalities Nt , Nb , Ne , Ns , Nk represent, the LiDAR stack dimension, the number of LiDAR beams for each stack, the embedding size, the number of angular sectors of the cross-attention layer, and the number of objects predicted by the perception module.
Fig. 4: Example of one scenario in the first experimental setup (parallel traffic) with 5 humans (in blue) and 5 box-shaped obstacles (in black). The robot (circle) and its goal (star) are depicted in red. Humans continuously flow in the corridor in the opposite direction of the robot.
Fig. 5: Example of the second experimental setup with 5 humans (in blue) and 5 circular obstacles (in black). The robot (circle) and its goal (star) are depicted in red. Humans continuously flow in the center of the scene.
Policy
SR (%)
CR-H (%)
CR-O (%)
TR (%)
TtG (s)
LJ (m/s 3 )
AJ (rad/s 3 )
SC (%)
JESSI-MT
94.2
0.0
1.8
4.0
18.25
1.30
4.84
90.1
JESSI-MD
94.9
0.3
1.7
3.1
17.88
1.46
4.88
88.2
JESSI-PO
95.1
0.1
2.2
2.6
18.03
1.77
5.45
87.9
V-E2E
86.3
0.2
12.6
0.9
16.93
0.76
2.27
87.6
BV-E2E
92.9
0.2
1.5
5.4
17.01
1.04
1.91
85.0
DIR-SAFE
73.2
1.1
4.5
21.2
23.17
2.99
10.82
92.1
TABLE I: First experimental setup results
Policy
SR (%)
CR-H (%)
CR-O (%)
TR (%)
TtG (s)
LJ (m/s 3 )
AJ (rad/s 3 )
SC (%)
JESSI-MT
99.9
0.0
0.0
0.1
20.55
0.97
2.96
89.7
JESSI-MD
99.8
0.0
0.0
0.2
19.82
1.11
2.98
87.5
JESSI-PO
100.0
0.0
0.0
0.0
19.73
1.25
3.07
87.9
V-E2E
99.0
0.0
1.0
0.0
18.82
0.65
1.61
87.3
BV-E2E
99.8
0.0
0.2
0.0
17.58
0.83
1.31
80.7
DIR-SAFE
98.1
0.0
0.1
1.8
22.65
2.63
7.52
91.4
TABLE II: Second experimental setup results
Fig. 6: Experimental setup of the real world test with the Loomo Segway robotic platform.
Deep Reinforcement Learning (DRL) has shown promise for social navigation, yet its real-world deployment remains hindered by a persistent sim-to-real gap arising from simplified first-order dynamics and context-specific human state estimation pipelines. This work presents a unified framework that addresses these limitations to produce dynamically feasible navigation policies suitable for real-world deployment. First, theoretical analysis reveals that tracking error between simulated and actual robot position decays exponentially with increased control order, motivating the use of higher-order control inputs as DRL action space. A second-order control formulation tailored to differential drive robots is developed, complemented by a stochastic iterative Linear Quadratic Regulator (iLQR) that pretrains the policy via a divergence minimization objective. Second, to avoid the added system complexity of camera-LiDAR fusion, a cluster-based human tracking pipeline using only 2D LiDAR is introduced. Human detections are associated according to both spatial proximity and velocity similarity, enabling reliable differentiation of nearby pedestrians and yielding stable velocity estimates through temporal aggregation. Third, we introduce an unbiased residual gating block to balance reaction- and memory-based behaviors while handling time-varying crowd sizes, both critical for social navigation. The resulting policy, KinematicRL, consistently improves kinematic performance and adapts to varying number of detected humans. Experiments in real-world environments demonstrate that, when combined with the proposed tracking pipeline, KinematicRL can be deployed on a real differential drive robot with minimal modifications.
Zhiming Xu, Haodong Yang, Chengju Liu +2
School of Computer Science and Technology, Tongji University, Shanghai 201804, China · Department of Electronics and Information Engineering, Tongji University, Shanghai 201804, China · Shanghai Institute of Intelligent Science and Technology, Tongji University, Shanghai 201210, China
Robots navigating among pedestrians typically sense their surroundings with a 2D LiDAR mounted close to the ground. At that height, the sensor mostly sees moving legs rather than whole people, yet most learning-based navigation methods still treat pedestrians as simple shapes like circles. This paper addresses that gap with CALF (Convolutional Attention for Leg Features), an end-to-end neural architecture that combines convolutional layers, attention, and MLP to interpret leg motion directly from LiDAR scans and produce safe navigation commands. The CALF policy is trained using deep reinforcement learning algorithms within LegNav, a custom lightweight 2D simulator that combines 2D LiDAR ray tracing with a novel pedestrian gait model. The resulting policy is compared against classical and learning-based baselines in terms of navigation performance and social compliance. The approach is validated through real-world experiments via zero-shot deployment on a TurtleBot 4, yielding smooth and socially compliant trajectories. Written in JAX, the LegNav simulator enables the training of a deployment-ready CALF policy in under an hour on a single consumer GPU.
Alberto Vaglio, Andrea Garulli, Antonio Giannitrapani +2
Department of Information Engineering and Mathematics, University of Siena, Siena, Italy · Uninettuno University, Rome, Italy
As mobile robots become more integrated into everyday human environments, social robot navigation is becoming essential for ensuring human comfort, safety, and trust. While reinforcement learning (RL) navigation policies provide the fast inference and reactive behavior necessary for real-time deployment, they still lack flexible semantic reasoning capabilities and often fail to generalize to complex social scenarios. Recent approaches have increasingly turned to vision-language models (VLMs) in place of RL policies to improve semantic and social reasoning in robot navigation. Nevertheless, their high computational cost and slow inference remain major barriers to real-time deployment. To overcome these limitations, we introduce HUMA (Hybrid Understanding for Multi-modal social Navigation), a hybrid architecture that dynamically balances the computational efficiency of RL policies with the deep semantic understanding of VLMs. Our approach uses a reactive RL policy to handle low-density, routine navigation tasks, while conditioning it on a post-trained high-level VLM when a human enters sensitive situations, such as the robot's proximity zone. We evaluate HUMA on the Social-MP3D and Social-HM3D benchmarks, where it achieves task success improvements of 20% and 3%, respectively, while significantly reducing personal space violations and human collisions against state-of-the-art baselines. Extensive ablation studies validate each architectural component, and real-world deployment on the Mirokaï mobile robot further demonstrates the practical viability of our approach.
Ali Ahmadi, Hamed Rahimi, Adrien Jacquet Cretides +3
Institut des Syst`emes Intelligents et de Robotique (ISIR) Sorbonne Universit´e France