Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using CVaR0.1, the average cost of the worst 10% of episodes. We classify a policy as tail-safe when CVaR0.1 is within the safety budget. This allows us first to identify policies that are safe on average but unsafe in the tail and then to study whether their tail violations can be controlled while preserving return. To identify tail-unsafe policies, we evaluate five standard algorithms on three Safety-Gymnasium navigation tasks. We then examine four constraint families on dense-hazard navigation and assess tail control across four navigation and four locomotion tasks.
Figures & tables
Figure 1: Mean compliance can hide episodic tail risk. Both panels describe the same illustrative policy. Its mean episodic cost is within budget, but the average cost of its worst 10% of episodes exceeds the budget. Our evaluation measures both quantities from complete post-training episodes.
Algo
Task
N
Return
Mean cost
CVaR0.1
×d
SAC-Lag
PointGoal1
3
25.4
48
116
4.6
TRPO-Lag
PointGoal1
3
25.7
47
110
4.4
FOCOPS
PointGoal1
3
25.4
52
134
5.4
PPO-Lag
PointGoal1
4
10.2
55
229
9.2
CPO
PointGoal1
3
2.9
19
139
5.6
TRPO-Lag
PointGoal2
3
7.9
87
350
14.0
Table 1: Episodic-cost tail under the shared post-training evaluation ( d=25 , 100 episodes per run). Statistics are computed for each run and then averaged. N is the number of available training runs. No reported policy is tail-safe.
Figure 2: The negative result on PointGoal1 ( d=25 ). The figure shows 17 operating points, including an unconstrained reference. Each point plots return against episodic-cost CVaR0.1 . Constraint tightening collapses return without bringing the tail within the shaded safe region.
OQ-SAC (ours)
SAC-Lag
PPO-Lag
Task
R↑
cost
CVaR↓
R↑
cost
CVaR↓
R↑
cost
CVaR↓
Locomotion
HalfCheetah
2868±46
0
0
3255±2757
308
597±425
3146±288
548
581±185
Hopper
877±556
0.1
1.0±1.3
1164±222
70
97±137
517±597
7.5
18±18
Ant
2656±329
0.9
0.9±0.6
2909±40
2.9
6.1±2.0
171±92
4.1
15±8
Walker2d
1617±949
2.4
10.7±10.1†
1273±1175
1.6
6.8±9.6
551±129
52
54±41
Table 2: Performance after 100 deterministic evaluation episodes with d=25 . Training uses 106 interactions except OQ-SAC on HalfCheetah , which uses 5×105 . Return and CVaR are means and standard deviations across seeds, while cost is the mean episodic cost. Results use three seeds except OQ-SAC on Walker2d , which uses five. Bold CVaR entries are safe on every seed. Underline marks the highest-return safe result. † OQ-SAC is safe on four of five Walker2d seeds, with the remaining seed at CVaR0.1=28.3 .
Figure 3: Episode return and cost over 200 evaluations per policy. Circles are safe episodes with C(τ)≤d , and crosses are unsafe. Safe, high-return episodes are frequent for the shown locomotion policies and absent or infrequent for the shown navigation policies.
Observed evidence
Locomotion
Navigation
Representative cost example
Separated clusters (HalfCheetah)
Diffuse costs (PointButton1)
Safe, high-return episodes in Figure 3
Frequent
Absent or infrequent
Safe tail at held return found
Yes
No
OQ-SAC safe on every seed
3 of 4 tasks
0 of 4 tasks
Table 3: Episode-level patterns associated with tail control in the evaluated tasks. These patterns summarize the experiments and are not necessary or sufficient conditions. The single-seed hazard-count study in Table 9 provides a within-navigation sensitivity analysis.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
OQ-SAC (ours)
SAC-Lagrangian
Task
mean
CVaR0.1
tail/mean
mean
CVaR0.1
tail/mean
Locomotion
HalfCheetah
0.0
0
–
308
597
1.9×
Hopper
0.1
1.0
10.0×
70
97
1.4×
Ant
0.9
0.9
1.0×
2.9
6.1
2.1×
Walker2d
2.4
10.7
4.5×
1.6
6.8
4.3×
Appendix
Table 4: Mean episode cost and CVaR0.1 under the shared 100 -episode deterministic evaluation with d=25 . The tail-to-mean ratio compares the average cost of the worst 10% of episodes with the mean across all episodes. Navigation entries are averaged across three training seeds. A dash indicates that both cost statistics are zero.
Figure 4: Cumulative cost over 40 deterministic episodes from one OQ-SAC policy per task. Orange curves identify the four episodes with the largest final costs. Blue curves show the remaining episodes, and the dashed line marks the episodic budget d=25 .
Mechanism family
PointGoal1
PointButton1
Total
Feasibility gating
13
0
13
Budget as state
4
0
4
Tail-driven Lagrangian
8
4
12
Per-state quantile-CVaR critic
4
5
9
Total
29
9
38
Appendix
Table 5: Archived training traces from the constraint-mechanism study. Each trace corresponds to one method configuration and one training seed.
Figure 5: Training trajectories for the 29 archived PointGoal1 runs. Panels A, B, and D show logged return and rolling empirical CVaR0.1 . Panel C shows deterministic test return and mean episode cost because the OmniSafe logs do not contain the episode-level costs required to reconstruct CVaR. Each line represents one configuration, and each dot marks its final logged value. The shaded green region marks costs at or below d=25 . These training statistics are separate from the shared post-training evaluation.
Figure 6: Training trajectories for the nine archived PointButton1 runs. Panel A shows deterministic test return and mean episode cost for four tail-driven Lagrangian configurations. Panel B shows logged return and rolling empirical CVaR0.1 for five per-state quantile-CVaR configurations. Each dot marks the final logged value. The shaded green region marks costs at or below d=25 .
Figure 7: Conceptual projection of the exploratory action shift. The blue arrow represents the change from the target-policy action aT to the exploratory action aE . This update follows the action-space gradient of the upper-confidence reward estimate minus the weighted lower-confidence cost estimate. The orange arrow illustrates a cost reduction accompanied by lower return. The arrows do not represent measured policy trajectories.
HalfCheetah
Hopper
Variant
R↑
CVaR↓
R↑
CVaR↓
Full OQ-SAC
2868±46
0±0
877±556
1.0±1.3
Mean-cost variant
2948±28
0±0
1246±398
0.5±0.4
No-shift variant
2908±32
0±0
1162±398
5.4±7.1
Appendix
Table 6: Component ablations on HalfCheetah and Hopper with three seeds and d=25 . Values are means and standard deviations across seeds. Every reported variant satisfies CVaR0.1≤d on every seed. Hopper uses 106 training interactions for every row. Full OQ-SAC on HalfCheetah uses 5×105 , while its two ablations use 106 .
OQ-SAC
SAC-Lagrangian
PPO-Lagrangian
Task
Return
CVaR
Return
CVaR
Return
CVaR
HalfCheetah
2868
0
1693(d′=10)
0
–
–
Hopper
877
1.0
1000(d′=5)
0
1319(d′=5)
0
Appendix
Table 7: Highest-return tail-safe policy found in the single-seed expected-cost target comparison. The internal training target d′ is shown in parentheses. OQ-SAC is included as a three-seed reference and is not part of this comparison. A dash indicates that no tested target produced a tail-safe policy.
OQ-SAC
WCSAC
Task
R↑
CVaR↓
R↑
CVaR↓
HalfCheetah
2868±46
0±0
2756±87
0±0
Ant
2656±329
0.9±0.6
−683±230
4.5±3.7
Walker2d
1617±949†
10.7±10.1
481±64
34.3±21.8
Hopper
877±556
1.0±1.3
1016±391
179.8±165.7
Appendix
Table 8: OQ-SAC and WCSAC on four locomotion tasks with d=25 . Values are means and standard deviations across training seeds. Bold CVaR entries satisfy the tail constraint on every seed. Underlined returns identify the higher-return method among those satisfying this criterion. † OQ-SAC is safe on four of five Walker2d seeds. The remaining seed has CVaR0.1=28.3 .
Figure 8: Per-seed return and episodic-cost CVaR0.1 on the four locomotion tasks. Each marker represents one trained policy evaluated for 100 deterministic episodes. The horizontal axis is linear within the budget and logarithmic above it. The green region marks CVaR0.1≤d=25 .
Figure 9: Episode-cost distributions from 200 deterministic evaluations of representative SAC-Lagrangian policies. On HalfCheetah , 85 episodes have cost above 10d . On PointButton1 , 6 episodes exceed this threshold. The shaded region marks episode costs at or below d=25 .
Figure 10: Observed return and episodic-cost CVaR0.1 on two representative tasks. The HalfCheetah panel shows seven OQ-SAC multiplier settings, with the connecting line ordered by multiplier value. The PointGoal1 panel shows 17 evaluated operating points across the tested constraint families. The green region marks CVaR0.1≤d=25 .
Hazards
Return
Safe%
Safe-episode return
CVaR0.1
CVaR/ d
1
27.9
90
28.0
40.9
1.6
2
−0.4
94
−0.4
116.2
4.6
4
27.3
62
27.4
82.3
3.3
8
26.8
34
27.7
176.9
7.1
16
6.8
63
4.9
397.2
15.9
Appendix
Table 9: Single-seed hazard-count comparison on PointGoal1 . Safe episodes satisfy C(τ)≤d . Safe-episode return is the mean return conditioned on this event. Every policy is evaluated for 100 deterministic episodes, and none satisfies CVaR0.1≤25 .
Figure 11: Mean return and CVaR0.1 as the number of hazards varies. The one-, four-, eight-, and sixteen-hazard results are connected in increasing order. The two-hazard result is shown separately because it does not follow this observed trend. The green region in the tail-cost panel marks CVaR0.1≤25 .
Safe reinforcement learning (RL) is commonly formalized as a Constrained Markov Decision Process (CMDP), in which an agent maximizes expected reward while keeping its expected cumulative cost below a specified safety bound. Existing safe RL benchmarks predominantly report whether an algorithm is safe on average, following this expectation-based guarantee. We argue that this convention is insufficient to reliably characterize an algorithm's true safety: it fails to capture how often and how severely the safety bound is violated, whether this holds consistently across tasks and safety bounds, and whether training-time behavior is representative of behavior of the final converged policy. Therefore, we introduce (i) evaluation metrics for safe RL that address each of these concerns and in addition allow for aggregation across tasks and safety bounds. We furthermore define (ii) a safety tier system to systematically categorize and compare algorithms in terms of safety and reliability at both training and for a final policy. Using this framework, we provide (iii) an empirical safety evaluation across multiple safety navigation tasks. Our results show that aggregate metrics, distributional reporting, and task- and safety bound-specific results each reveal information the other metrics cannot. We therefore recommend reporting all three jointly, rather than compressing this information into a single value, as is common practice. We provide SafeRLEval, an open-source evaluation suite to support the reliable characterization of safety in future safe RL research.
Lindsay Spoor, Aske Plaat, Thomas Moerland
Leiden Institute of Advanced Computer Science, Leiden University, Leiden, the Netherlands
Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet most existing safe reinforcement learning (RL) methods assess safety through average cumulative cost. Such metrics can mask dangerous tail-risk behaviors. To address this, we propose a framework that trains risk-sensitive policies through Conditional Value-at-Risk (CVaR) constrained optimization on an off-policy TD3 backbone and evaluates their safety margins post-training through neural network reachability verification. During training, the policy is optimized under CVaR constraints on cumulative costs, promoting sensitivity to high-cost tail outcomes rather than average behavior alone. After training, we compute action reachable sets under bounded observation uncertainty using Taylor Model analysis, yielding a safety rate metric that quantifies the proportion of evaluated states at which the policy's reachable action set remains within prescribed safety margins. A key finding is that policies trained with CVaR constraints maintain larger safety margins from obstacles across evaluated states. This makes them significantly more amenable to formal reachability verification. Experiments across ten navigation scenarios and six baselines show that our method achieves a 98.3% success rate, the highest safety verification rate among all compared methods, while revealing that average cost rankings and reachability-based safety rankings can diverge. This indicates that reachability verification captures risks which are missed by empirical cost metrics alone. We further validate our approach on a physical Clearpath Jackal robot, demonstrating successful sim-to-real transfer.
Qisong He, Xinmiao Huang, Jinwei Hu +4
University of Liverpool, Liverpool, UK. · Universit´e Grenoble Alpes, Grenoble, France.
Safe reinforcement learning typically enforces safety by bounding expected cumulative costs, a criterion that often fails to detect rare but catastrophic tail events. To overcome these limitations, this paper introduces SteinGate, a boundary-aware distributional safety certificate that replaces fragile tail fitting with a robust consistency check using Kernelized Stein Discrepancy while accounting for boundary atoms induced by clipped costs. SteinGate evaluates whether observed policy rollout costs remain consistent with a safe reference distribution, providing a non-parametric safety certificate. This certificate is used to dynamically adapt the learning regime: favoring reward-improving policy updates when rollouts remain consistent with the safe reference and switching to recovery behavior when the cost tail deviates. Experiments on continuous-control benchmarks demonstrate that SteinGate significantly reduces both the frequency and severity of constraint violations during training while maintaining competitive returns relative to state-of-the-art baselines.
Yassine Chemingui, Chenhua Fan, Honghao Wei +1
School of Electrical Engineering and Computer Science, Washington State University, Pullman, Washington, USA