cs.LGMay 11, 2026

Robust Probabilistic Shielding for Safe Offline Reinforcement Learning

Authors: Maris F. L. GaleslootThomas RhemrevNils Jansen

Organizations: Radboud University · Nijmegen, The Netherlands · Ruhr University & Radboud University · Bochum, Germany

Abstract

In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide guarantees on the (1) performance and (2) safety of the resulting policy. A technique called safe policy improvement (SPI) provides a performance guarantee: with high probability, the new policy outperforms a given baseline policy, which is assumed to be safe. Orthogonally, in the context of safe RL, a shield provides a safety guarantee by restricting the action space to those actions that are provably safe with respect to a given safety-relevant model. We integrate these paradigms by extending shielding to offline RL, relying solely on the available dataset and knowledge of safe and unsafe states. Then, we shield the policy improvement steps, guaranteeing, with high probability, a safe policy. Experimental results demonstrate that shielded SPI outperforms its unshielded counterpart, improving both average and worst-case performance, particularly in low-data regimes.

Explore similar work

CardsList
  1. Robust Shielding for Safe Reinforcement Learning

    May 29, 2026Edwin Hamel-De le Court, Thom Badings, Alessandro Abate +2ShieldingMarkov Decision Processes

  2. Easy-to-Use Shielding for Reinforcement Learning

    Jun 2, 2026Stefan Pranger, Bettina KönighoferShielding