Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms typically force a trade-off: they either require a priori knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a joint value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the safety value of the learning joint policy is estimated, while the joint value is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an occupation-averaged differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. Theoretically, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with neural approximations demonstrate stable convergence with near-zero safety violations at test time.
Figures & tables
Figure 1: (Left) The Joint Value function is conceptualized as a 2D figure, with the state-space indicated on the x -axis, and joint values on the y -axis. (Right) The corresponding regions within the safe set Ωπ are depicted. The locus of points where Vπϕ=Vπψ indicates the boundary between Ωϕ and Ωψ as defined in Definition 1 . Trajectories must stay within Ωψ to maximize episodic return.
Figure 2: Training performance across 5 seeds (mean ± standard deviation). The top row indicates episodic returns, while the bottom row indicates episodic costs as training progresses. JointSAC demonstrates stable convergence with near-zero episode costs.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Safe set expands initially in the direction of high rewards during training ( InvertedPendulum-v4 rewarded for positive cart velocity), before converging to the maximal safe set. Constraint boundary (red dotted line) and 0 -level set of safety critic (black line) are visualized.
Figure 4: Figure shows variation of episode reward and episode cost as K increases from 200 to 2000 . While episode reward generally tends to increase with an increase in K , scaling beyond a point can introduce training instabilities with neural approximations, as seen with Walker2d environment.
Figure 5: Training performance across 5 seeds (mean ± standard deviation) on three locomotion tasks from Gymnasium suite with adversarial disturbances. The top row indicates episodic returns, while the bottom row indicates episodic costs as training progresses. JointSAC demonstrates stable convergence with near-zero episode costs on two out of three tasks. A single cost violation incurs +1 penalty each timestep.
Figure 6: Training performance across 5 seeds (mean ± standard deviation) on eight single agent environments from SafeVelocity ( v0 ) suite from SafetyGymnasium with adversarial disturbances. The top two rows indicate episodic returns, while the bottom two rows indicate episodic costs as training progresses. JointSAC demonstrates stable convergence with near-zero episode costs.
In safety-critical domains, reinforcement learning (RL) systems must satisfy strict, zero-cost safety constraints while achieving meaningful task performance. Existing model-free methods can struggle to achieve high safety without substantially compromising task performance. We introduce \emph{Safety-Biased Trust Region Policy Optimisation (SB-TRPO)}, a principled approach to RL with zero-cost constraints, which requires only a fixed fraction of the maximal cost reduction achievable within the trust region, thus retaining flexibility for reward optimisation. We show that the idealised update nevertheless converges to zero cost and maximal reward amongst zero-cost policies in finite MDPs. A practical gradient-based approximation provides local improvements in both safety and reward under suitable gradient alignment. Experiments on \emph{Safety Gymnasium} demonstrate high safety alongside strong task performance.
Safe exploration remains a fundamental challenge in reinforcement learning (RL), limiting the deployment of RL agents in the real world. We propose Sampling-Based Safe Reinforcement Learning (SBSRL), a model-based RL algorithm that maintains safety throughout the learning process by enforcing constraints jointly across a finite set of dynamics samples. This formulation approximates an intractable worst-case optimization over uncertain dynamics and enables practical safety guarantees in continuous domains. We further introduce an exploration strategy based on constraining epistemic uncertainty, eliminating the need for explicit exploration bonuses. Under regularity conditions, we derive high-probability guarantees of safety throughout learning and a finite-time sample complexity bound for recovering a near-optimal policy. Empirically, SBSRL achieves safe and efficient exploration both in simulation and in real robotic hardware, and readily extends to practical deep-ensemble implementations that scale to high-dimensional continuous control problems.
Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in which an agent maximizes expected reward while keeping its expected cumulative cost below a specified safety bound. Existing safe RL benchmarks predominantly report whether an algorithm is safe on average, following this expectation-based guarantee. We argue that this convention is insufficient to reliably characterize an algorithm's true safety: it fails to capture how often and how severely the safety bound is violated, whether this holds consistently across tasks and safety bounds, and whether training-time behavior is representative of behavior of the final converged policy. Therefore, we introduce (i) evaluation metrics for safe RL that address each of these concerns and in addition allow for aggregation across tasks and safety bounds. We furthermore define (ii) a safety tier system to systematically categorize and compare algorithms in terms of safety and reliability at both training and for a final policy. Using this framework, we provide (iii) an empirical safety evaluation across multiple safety navigation tasks. Our results show that aggregate metrics, distributional reporting, and task- and safety bound-specific results each reveal information the other metrics cannot. We therefore recommend reporting all three jointly, rather than compressing this information into a single value, as is common practice. We provide SafeRLEval, an open-source evaluation suite to support the reliable characterization of safety in future safe RL research.
Lindsay Spoor, Aske Plaat, Thomas Moerland
Leiden Institute of Advanced Computer Science, Leiden University, Leiden, the Netherlands