DODGER: Safety-Guided Reinforcement Learning for Robot Navigation Among Dynamic Obstacles
Organizations: Department of Intelligence and Information, Seoul National University, Republic of Korea · Department of Robotics, University of Michigan, Ann Arbor, MI, USA · Department of Mechanical Engineering, Yonsei University, Republic of Korea · Center for Humanoid Research, Korea Institute of Science and Technology (KIST), Seoul, Republic of Korea
Abstract
Robots operating in human-centered environments must safely navigate among multiple dynamic obstacles to avoid collisions with people and surrounding infrastructure. Control barrier functions (CBFs) provide an effective mechanism for safety filtering, and recent CBF-based reinforcement learning (RL) methods embed such safety information into learned policies. However, executing only safety-filtered actions during training can restrict policy exploration, a limitation that becomes particularly consequential in dynamic scenes where safety depends on relative robot-obstacle motion. We propose DODGER, a safety-guided RL framework that directly executes policy-generated actions to drive training rollouts while using CBF-filtered references and constraint violations to shape the policy toward collision-avoidance behavior. We evaluate DODGER through a Dubins-car safety analysis and demonstrate goal-directed navigation among multiple dynamic obstacles in full-order humanoid simulation and real-world humanoid experiments using LiDAR-based perception, without a runtime safety filter.
Figures & tables
| Term | Definition | Weight |
|---|---|---|
| CBF formulation | End (M) | Converged | Stage |
|---|---|---|---|
| DPCBF (DODGER) | 113.97 | ✓ | 3 |
| C3BF | 600 | ✗ | 1 |
| Dist.-ECBF | 166.20 | ✓ | 3 |
| Dist.-ECBF + yaw | 154.01 | ✓ | 3 |
| Metric | DPCBF | Dist. | Dist.+Yaw |
|---|---|---|---|
| Cmd. RMS ( ) | 0.554 | 0.934 | 0.894 |
| Cmd. near-sat. (%) | 5.48 | 87.08 | 61.56 |
| G1 RMS ( ) | 0.435 | 0.808 | 0.763 |
| G1 near-sat. (%) | 0.87 | 18.80 | 15.45 |
| Component | Configuration |
|---|---|
| Graph Attention Encoder | |
| Graph input | Robot, goal, and up to obstacle nodes |
| Node representation | |
| Edge representation | |
| Message network | Linear + ReLU, Linear |
| Attention network | Linear + ReLU, Linear , masked softmax |