cs.LGJun 30, 2026

Safe Online Learning via Smooth Safety-Structured Policy Composition

Authors: Hongpeng CaoLiqun ZhaoYuliang GuNaira HovakimyanLui ShaMarco Caccamo

Organizations: School of Engineering and Design Technical University of Munich Garching, Munich 85748, Germany · Department of Engineering Science University of Oxford Oxford OX1 3PJ, United Kingdom · Department of Mechanical Science and Engineering University of Illinois Urbana-Champaign Urbana, IL 61801, USA · Department of Computer Science University of Illinois Urbana-Champaign Urbana, IL 61801, USA

Abstract

Safe online reinforcement learning requires policies to respect safety constraints while maintaining smooth optimization dynamics. Existing approaches typically rely on either strict safety enforcement via action interventions, which introduce discontinuities in system interaction and learning, or soft safety constraint formulations, which preserve smooth learning but provide limited safety assurance. We propose AutoSafe, a safety-aware policy architecture that integrates structured safety monitoring and intervention directly into the action generation process. This design enables smooth, risk-dependent transitions between performance-driven and safety-preserving behaviors, resulting in continuous online interaction and learning dynamics. Empirical results across a suite of continuous-control benchmarks demonstrate strong safety enforcement without sacrificing learning smoothness. We further validate AutoSafe on a physical cart-pole system, highlighting its practical effectiveness for safe online learning in the real world.

Explore similar work

CardsList