cs.LGOct 7, 2026

Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?

Authors: Samuel Tetteh, Cody Fleming

Organizations: Iowa State University Ames, Iowa, USA

Abstract

Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using CVaR0.1\mathrm{CVaR}_{0.1}, the average cost of the worst 10%10\% of episodes. We classify a policy as tail-safe when CVaR0.1\mathrm{CVaR}_{0.1} is within the safety budget. This allows us first to identify policies that are safe on average but unsafe in the tail and then to study whether their tail violations can be controlled while preserving return. To identify tail-unsafe policies, we evaluate five standard algorithms on three Safety-Gymnasium navigation tasks. We then examine four constraint families on dense-hazard navigation and assess tail control across four navigation and four locomotion tasks.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Evaluation Metrics for Safe Reinforcement Learning

    Sep 14, 2026Lindsay Spoor, Aske Plaat, Thomas MoerlandRL BenchmarksConstrained RL

  2. Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

    May 13, 2026Qisong He, Xinmiao Huang, Jinwei Hu +4Constrained RLRisk-Sensitive RL

  3. SteinGate: Tail-Sensitive Safe Reinforcement Learning via Stein Discrepancy

    Jul 14, 2026Yassine Chemingui, Chenhua Fan, Honghao Wei +1Constrained RLRisk-Sensitive RL