Humanoid robots promise versatile mobility in cluttered, human-centric environments, but real deployment demands principled safety. Classical model-based gait generators yield interpretable motions but often lack the robustness and adaptability of modern reinforcement learning (RL) based approaches. We propose a model-informed reinforcement learning framework anchored to the analytical Angular Momentum Linear Inverted Pendulum (ALIP) template. We provide a step-to-step safety certificate for ALIP stepping via a discrete exponential control barrier function (DECBF) and use it as (i) a training-time shaping signal and (ii) a runtime action filter that minimally adjusts swing-foot placement to satisfy template-level constraints. Full-order safety is evaluated empirically on the Digit humanoid in MuJoCo with a whole-body controller stack. Compared to an unconstrained baseline, our approach reduces safety-violation events in the reported external-disturbance trial, while larger lateral-velocity transients reveal a safety-tracking tradeoff.
Figures & tables
Fig. 1: Overview of the ALIP-informed safe RL framework on the Digit humanoid. The policy commands swing-foot placement and torso pitch. DECBFs derived from the ALIP model constrain foot-placement actions during training (safety shaping) and execution (projection filter).
Fig. 2: Training vs. execution pipeline. During training, the policy outputs action ak=(uk,θk) ; the foot-placement component uk is evaluated through the ALIP step-to-step map to compute the safety certificate sk , yielding a shaping reward rsafe . During execution, the shaping reward is replaced by a projection that finds the closest uk satisfying sk≥0 before passing the command to the whole-body controller.
Fig. 3: Orbital energy and the orbital-energy safe regions used in this work. Left: orbital-energy phase portrait with iso-energetic lines. Middle: lateral-walking safe region ( py∈[−0.5,0.5] m and E∈[−0.464,−0.012] ). Right: sagittal safe region ( px∈[−0.7,0.7] m and E<1.125 ).
Quantity
Value
Rationale
px,min,px,max
(−0.70,0.70) m
stance reach
py,max
0.50 m
lateral reach
Wmin
0.08 m
no leg crossing
vx,max
1.5 m/s
energy surrogate
T
0.35 s
nominal step duration
TABLE I: Safety bounds for Digit
Fig. 4: Lateral-plane phase portrait showing the relationship between orbital energy and leg collision. During stable walking, the system remains in the safe region ( E<0 ). After a disturbance drives the orbital energy positive (red states), the CoM crosses the stance foot and a leg collision occurs.
Fig. 5: Velocity tracking (left) and lateral foot distance (right) under periodic external disturbances. Left: desired vs. average CoM velocity for (a) unguided, filter off; (b) unguided, filter on; (c) DECBF-guided, filter off; (d) DECBF-guided, filter on. Right: lateral distance from stance foot to swing foot at touchdown compared against the safety bound Wmin=0.08 m. Subplots (a)–(d) follow the same ordering.
Fig. 6: Sagittal phase portrait under periodic disturbances ( 25 ). (a) The unguided policy violates the orbital-energy and CoM-extension bounds. (b) The DECBF-guided policy remains within the safe region. The red star marks the initial state.
Perceptive bipedal locomotion over sparse terrain remains a difficult challenge: model-based methods are precise but brittle to uncertainty, while model-free methods are robust but struggle to discover the precise, constrained motions required for safety-critical locomotion where small errors can cause catastrophic failures. We propose a model-assisted reinforcement learning (RL) framework that combines both perspectives in three steps: (1) generate a safe reference trajectory using simplified models; (2) train a privileged teacher policy guided by a control Lyapunov function (CLF) reward built around the safe reference trajectory; and (3) distill the teacher into a vision-based student policy. We show that this model-assistance procedure produces physically grounded locomotion, improving sample efficiency, reducing the need for a complex learning curriculum, and achieving smoother locomotion behavior alongside stepping stone performance comparable to model-free baselines. We validate our approach in simulation and demonstrate successful deployment on a Unitree G1 humanoid robot navigating sparse footholds with lateral constraints.
Codrin Crismariu, Ryan K. Cosner
Department of Mechanical Engineering Tufts University, Medford, MA, United States
Robot learning has produced remarkably effective ``black-box'' controllers for complex tasks such as dynamic locomotion on humanoids. Yet ensuring dynamic safety, i.e., constraint satisfaction, remains challenging for such policies. Reinforcement learning (RL) embeds constraints heuristically through reward engineering, and adding or modifying constraints requires retraining. Model-based approaches, like control barrier functions (CBFs), enable runtime constraint specification with formal guarantees but require accurate dynamics models. This paper presents SHIELD, a layered safety framework that bridges this gap by: (1) training a generative, stochastic dynamics residual model using real-world data from hardware rollouts of the nominal controller, capturing system behavior and uncertainties; and (2) adding a safety layer on top of the nominal (learned locomotion) controller that leverages this model via a stochastic discrete-time CBF formulation enforcing safety constraints in probability. The result is a minimally-invasive safety layer that can be added to the existing autonomy stack to give probabilistic guarantees of safety that balance risk and performance. In hardware experiments on an Unitree G1 humanoid, SHIELD enables safe navigation (obstacle avoidance) through varied indoor and outdoor environments using a nominal (unknown) RL controller and onboard perception.
Lizhi Yang, Blake Werner, Ryan K. Cosner +3
Mechanical and Civil Engineering, California Institute of Technology · Aerospace Engineering and Engineering Mechanics, UT Austin · Computer Science, Cornell University
Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this work, we propose residual reinforcement learning as an implicit safety-filtering mechanism for safe humanoid control. Instead of relying on a single nominal policy to simultaneously balance performance, safety, and robustness, we decouple performance and safety. The nominal policy focuses solely on task performance, while a residual policy learns safety corrections. This decoupling leads to a better performance--safety Pareto trade-off and avoids the need for careful tuning of multiple competing reward terms within a single policy training. We show that the residual policy can act as an implicit safety filter.
Gechen Qu, Tong Zhang, Bike Zhang +4
University of California, Berkeley · University of California, Los Angeles