Safe reinforcement learning commonly places safety and task performance in the same policy objective, where they can introduce competing updates. Safety filters separate them at action execution, but classical designs require an analytic safety function and dynamics model, and standard minimal-intervention filters are myopic to long-horizon task return because they minimize only instantaneous action deviation. Hard projections are also undefined when no safe action exists. We present FAITH, a feasibility-aware, model-free framework that approximates the optimal state-action safety value and amortizes minimal-intervention filtering with a feedforward network. The task policy optimizes the task return through the filtered dynamics, which recovers the feasible constrained problem without a competing safety term in the task-policy update. When no action satisfies the learned safety condition, the same filter approaches the action with minimum predicted peak harm. On a double integrator example and a Safety Gym environment, FAITH achieves the highest return among methods with no feasible-start violations and matches the lowest harm from infeasible starts. On a 29-DoF humanoid, it reaches a 99.95% safety rate while retaining 97% of the unfiltered return in Walking-Avoid, and obtains the highest measured safety rate in Push-Avoid by learning to sacrifice balancing and fall away from the protected region. The same policies are also demonstrated on a real-world Unitree G1 humanoid.
Figures & tables
Fig. 1: FAITH separates safety learning from task-return optimization. (a) The task policy proposes a nominal action, and the safety filter produces the executed action. The dashed box denotes the filtered dynamics seen by the task policy . (b) The safety critic , backup actor , and safety filter learn from replay. The backup actor supplies actions for the safety critic’s Bellman target, while critic gradients guide the backup actor and safety filter . (c) PPO updates the task policy using nominal-action likelihood ratios and filtered-rollout returns. Only the task policy and safety filter are evaluated for deployment.
Fig. 2: Double integrator results: safety rate and task return of feasible starts, and mean peak harm of infeasible starts. Left: For feasible starts, FAITH achieves a 100% safety rate, and the highest return among all safe methods (error bars show min-max ranges). Right: For infeasible starts, FAITH achieves the lowest peak harm ( ± shows the standard deviation).
Fig. 3: Trajectories in the double integrator environment. FAITH matches the feasible set closely in visited states. From feasible starts ( ∘ ), FAITH rides the feasibility boundary to achieve the highest return and stays safe. When infeasible ( △ ), FAITH applies maximum braking to minimize peak harm.
Fig. 4: Car-Circle Results. With all initial states being feasible, FAITH has 100% safety and high task return (error bars show min-max ranges).
Fig. 5: Push-Avoid simulation. The red box marks the protected region and the arrow denotes the push. (a) The robot remains upright while staying outside the region. (b) When the push is too hard, the robot sacrifices balance and falls away while protecting its head and hands.
Method
Safety % ↑
Return ↑
No Filter
80.37
2428±288
ROCBF( 0 )
99.95
2094±342
ROCBF( 0.1 )
99.95
2182±330
ROCBF( 1 )
99.71
2362±307
ROCBF( 5 )
94.80
2373±303
FAITH
99.95
2365±345
TABLE I: Walking-Avoid. ± denotes episode-level standard deviation.
Method
Safety % ↑
Return ↑
Emaxkh (common failure) ↓
No Filter
21.53
123.2±96.0
0.253±0.149
ROCBF( 0 )
84.52
115.5±64.3
0.172±0.146
ROCBF( 0.1 )
83.35
116.3±41.3
0.175±0.140
ROCBF( 1 )
78.56
117.1±87.0
0.165±0.140
ROCBF( 5 )
60.11
115.5±96.6
0.166±0.133
FAITH
85.45
0080.0±104.9
0.153±0.184
TABLE II: Push-Avoid. Peak harm is evaluated on common failure episodes. ± denotes episode-level standard deviation.
Controller
Return J↑
Safety (%) ↑
Emaxkh↓
Oracle
−219.37
100.00
0.1499
FAITH
−220.64±0.23
100.00
0.1499±9e−6
PPO+GTCBF
−219.40±0.01
100.00
0.1500
PPO+ROCBF( 0.5 )
−241.39±0.01
99.95
0.1506
PPO+ROCBF( 3 )
−187.46±0.04
0.00
1.3501±1.7e−4
PPO + FAITH filter (fixed)
−221.09±0.22
99.98±0.04
0.1499±9e−6
TABLE III: Double Integrator ablations. ± denotes standard deviation.
Controller
Return J↑
Safety (%) ↑
FAITH
19.20±0.71
100.00
PPO + FAITH filter (fixed)
18.46±1.39
100.00
FAITH w/o filtered dyn.
8.96±5.21
98.44±2.07
FAITH w/o backup replay
18.17±1.29
97.66±3.41
FAITH w/ numerical search
18.84±0.37
96.88±2.07
FAITH ( 0.25λmax )
18.98±0.93
95.31±5.63
TABLE IV: Car-Circle ablations. ± denotes standard deviation.
Fig. 6: Hardware control pipeline. SLAM-based ground removal and angular binning produce a 64 -beam LiDAR observation. FAITH generates joint targets tracked by PD control. The dashed box marks simulated components, including G1 dynamics and analytically generated beam inputs.
Safe control of humanoid robots remains challenging due to their high-dimensional dynamics, contact-rich interactions, and sensitivity to disturbances. Although reinforcement learning has enabled effective locomotion and motion tracking, learned policies can still generate unsafe actions that lead to instability or falls. In this work, we propose residual reinforcement learning as an implicit safety-filtering mechanism for safe humanoid control. Instead of relying on a single nominal policy to simultaneously balance performance, safety, and robustness, we decouple performance and safety. The nominal policy focuses solely on task performance, while a residual policy learns safety corrections. This decoupling leads to a better performance--safety Pareto trade-off and avoids the need for careful tuning of multiple competing reward terms within a single policy training. We show that the residual policy can act as an implicit safety filter.
Gechen Qu, Tong Zhang, Bike Zhang +4
University of California, Berkeley · University of California, Los Angeles
Reinforcement learning (RL) policies enable dynamic legged locomotion but lack mechanisms to avoid violations of safety constraints that are absent during training. Large-scale offline safe learning is impractical for covering all edge cases. Existing safety frameworks either rely on reduced-order models that cannot reason about whole-body behaviors or require conservative recovery controllers that degrade task performance. We propose a predictive safety filter that post-hoc filters the nominal contact locations fed to the RL policy. When a collision is predicted, a sampling-based optimizer asynchronously searches for safer contact sequences using a full-physics model, while a learned value function bootstraps long-horizon returns. Our three algorithmic components (geometric projection of sampled contacts, momentum-augmented updates, and replica-exchange) make the optimization tractable in a discontinuous contact landscape. We validate the filter on a quadruped robot in dense, cluttered environments, both in simulation and in the real world, showing substantial reductions in safety violations with minimal deviation from the nominal input.
Aditya Shirwatkar, Sebastian Sanokowski, Shishir Kolathaya +2
Robert Bosch Center for Cyber Physical Systems, Indian Institute of Science, Bangalore, India · Munich Institute of Robotics and Machine Intelligence (MIRMI), Technical University of Munich, Munich, Germany · Department of Computer Science & Automation, Indian Institute of Science, Bangalore, India +2
Safety remains an open problem in reinforcement learning (RL), especially during training. While safety filters are promising to address safe exploration, they are generally poorly suited for high-dimensional systems with unknown dynamics. We propose Dyna-style Safety Augmented Reinforcement Learning (Dyna-SAuR), a novel algorithm that learns both a scalable safety filter and a control policy using a learned uncertainty-aware dynamics model, while requiring minimal domain knowledge. The filter avoids failures and high uncertainty regions. Thus, better models expand the set of safe and certain states, reducing filter conservatism. We present the effectiveness of Dyna-SAuR on goal-reaching CartPole as well as MuJoCo Walker, reducing failures compared to state-of-the-art methods by 2 orders of magnitude.
Artur Eisele, Bernd Frauenknecht, Friedrich Solowjow +1
Institute for Data Science in Mechanical Engineering, RWTH Aachen University, Aachen, Germany