DODGER: Safety-Guided Reinforcement Learning for Robot Navigation Among Dynamic Obstacles
Authors: Sanghyuk Park, Kwanwoo Lee, Taekyung Kim, Seohyeon Lim, Yisoo Lee
Organizations: Department of Intelligence and Information, Seoul National University, Republic of Korea · Department of Robotics, University of Michigan, Ann Arbor, MI, USA · Department of Mechanical Engineering, Yonsei University, Republic of Korea · Center for Humanoid Research, Korea Institute of Science and Technology (KIST), Seoul, Republic of Korea
Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such safety information into learned policies. However, executing only safety-filtered actions during training can restrict policy exploration, a limitation that becomes particularly consequential in dynamic scenes where safety depends on relative robot-obstacle motion. We propose DODGER, a safety-guided RL framework that directly executes policy-generated actions to drive training rollouts while using CBF-filtered references and constraint violations to shape the policy toward collision-avoidance behavior. We evaluate DODGER through a Dubins-car safety analysis and demonstrate goal-directed navigation among multiple dynamic obstacles in full-order humanoid simulation and real-world humanoid experiments using LiDAR-based perception, without a runtime safety filter.
Figures & tables
Fig. 1: A humanoid avoids a moving pedestrian while navigating toward a goal using the DODGER policy. The policy is trained with CBF-based safety guidance and deployed without a runtime safety filter.
Fig. 2: Overview of DODGER. The high-level policy combines local robot-goal information with a graph-based representation of surrounding obstacles to generate navigation commands. During training, the policy action directly drives the environment rollout, while the CBF-filtered reference and constraint violations guide learning. The resulting navigation command is executed by a fixed pretrained locomotion controller.
Fig. 3: Comparison of DODGER and CBF-RL on the Dubins-car benchmark. The dashed black curve denotes the HJ safety boundary ∂VHJ , and the shaded region denotes U^ϵ .
Term
Definition
Weight
rCBF
exp(−∥Δu∥22/σ2)−1
100
rprog
Δρ
60
rhead
Δeˉϕexp[−(ρ/ℓh)2]
10
rgoal
\mathmybb1{goal reached}
20
robs
\mathmybb1{obstacle collision}
−55
routside
\mathmybb1{p∈/Ahard}
−100
TABLE I: Reward terms used for training DODGER.
Fig. 4: Evaluation success rates of DODGER, CBF-RL-Min, and CBF-RL during the first 300 million environment transitions for humanoid navigation with obstacle speeds up to 0.6m/s . The markers denote individual evaluations, and the solid lines denote the moving average over the three most recent evaluations.
CBF formulation
End (M)
Converged
Stage
DPCBF (DODGER)
113.97
✓
3
C3BF
600
✗
1
Dist.-ECBF
166.20
✓
3
Dist.-ECBF + yaw
154.01
✓
3
TABLE II: Training outcomes under different CBF formulations. “End” denotes the number of environment transitions (millions) at which training terminated; curriculum stages are indexed from 0 to 3.
Metric
DPCBF
Dist.
Dist.+Yaw
Cmd. RMS ( rad/s )
0.554
0.934
0.894
Cmd. near-sat. (%)
5.48
87.08
61.56
G1 RMS ( rad/s )
0.435
0.808
0.763
G1 near-sat. (%)
0.87
18.80
15.45
TABLE III: Yaw-rate statistics for different CBF formulations.
Fig. 5: Real-robot navigation with a Unitree G1 among multiple moving pedestrians. The upper panels show the estimated scene and camera views at six representative times, together with DPCBF boundaries and relative velocities; h<0 when a relative velocity enters the unsafe side of its corresponding parabolic boundary. The lower-left plot shows the DPCBF values h and margins d of the selected obstacles, with solid curves denoting their minima. Although the minimum h temporarily falls below zero, the minimum d remains positive throughout the experiment.
Component
Configuration
Graph Attention Encoder
Graph input
Robot, goal, and up to 10 obstacle nodes
Node representation
nj∈R9
Edge representation
erj∈R11
Message network ψ1
Linear (11→64) + ReLU, Linear (64→16)
Attention network ψ2
Linear (16→16) + ReLU, Linear (16→1) , masked softmax
TABLE IV: Network architectures used for humanoid navigation.
RGB-only end-to-end visual navigation policies remain vulnerable to collisions in real-world dynamic environments, motivating a dedicated safety layer. Existing visual Control Barrier Function (CBF) approaches seek to provide safety from RGB observations, but often rely on real-time rendering or explicit scene reconstruction and are primarily designed for static scenes, limiting their practicality for onboard deployment. We propose a teacher-student visual distillation framework that transfers the safety behavior of a privileged CBF teacher to an RGB-only student filter for dynamic environments. The student maps a short RGB history, robot velocity, and a nominal control action directly to a safe action, while the teacher uses ground-truth robot and obstacle states in a real-to-sim dynamic Gaussian Splatting environment. To reduce the teacher-student information gap, the teacher constructs safety constraints only from obstacles observable within the student's RGB history. It also accounts for obstacle-velocity uncertainty to improve robustness to motion variations, while action augmentation exposes the student to diverse safe and unsafe nominal actions to better capture the safety boundary. At deployment, the student requires only RGB observations and robot velocity, without explicit 3D reconstruction or online rendering. Experiments show that the proposed method outperforms visual CBF baselines and improves the safety of RGB-based navigation policies under dynamic obstacle motion. Project page: https://syeon-yoo.github.io/distill-cbf-site/.
Seungyeon Yoo, Gawon Lee, Seungwoo Jung +2
Department of Aerospace Engineering, Seoul National University, Seoul, Republic of Korea
Reinforcement learning (RL), while powerful and expressive, can often prioritize performance at the expense of safety. Yet safety violations can lead to catastrophic outcomes in real-world deployments. Control Barrier Functions (CBFs) offer a principled method to enforce dynamic safety -- traditionally deployed online via safety filters. While the result is safe behavior, the fact that the RL policy does not have knowledge of the CBF can lead to conservative behaviors. This paper proposes CBF-RL, a framework for generating safe behaviors with RL by enforcing CBFs in training. CBF-RL has two key attributes: (1) minimally modifying a nominal RL policy to encode safety constraints via a CBF term, (2) and safety filtering of the policy rollouts in training. Theoretically, we prove that continuous-time safety filters can be deployed via closed-form expressions on discrete-time roll-outs. Practically, we demonstrate that CBF-RL internalizes the safety constraints in the learned policy -- both enforcing safer actions and biasing towards safer rewards -- enabling safe deployment without the need for an online safety filter. We validate our framework through ablation studies on navigation tasks and on the Unitree G1 humanoid robot, where CBF-RL enables safer exploration, faster convergence, and robust performance under uncertainty, enabling the humanoid robot to avoid obstacles and climb stairs safely in real-world settings without a runtime safety filter.
Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent during training. Large-scale offline safe learning is impractical for covering all edge cases. Existing safety frameworks either rely on reduced-order models that cannot reason about whole-body behaviors or require conservative recovery controllers that degrade task performance. We propose a predictive safety filter that post-hoc filters the nominal contact locations fed to the RL policy. When a collision is predicted, a sampling-based optimizer asynchronously searches for safer contact sequences using a full-physics model, while a learned value function bootstraps long-horizon returns. Our three algorithmic components (geometric projection of sampled contacts, momentum-augmented updates, and replica-exchange) make the optimization tractable in a discontinuous contact landscape. We validate the filter on a quadruped robot in dense, cluttered environments, both in simulation and in the real world, showing substantial reductions in safety violations with minimal deviation from the nominal input.
Aditya Shirwatkar, Sebastian Sanokowski, Shishir Kolathaya +2
Robert Bosch Center for Cyber Physical Systems, Indian Institute of Science, Bangalore, India · Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich, Munich, Germany · Department of Computer Science & Automation, Indian Institute of Science, Bangalore, India +2