cs.ROSep 5, 2026

SCoCaT: Success Conditioned Constrained Reinforcement Learning for Spacecraft Docking

Authors: Aman Arora, Ricard Marsal I Castan, Matteo El-Hariry, Miguel Olivares-Mendez

Organizations: SnT - Interdisciplinary Centre for Security, Reliability and Trust University of Luxembourg, Luxembourg

Abstract

Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. CSPO: Constraint-Sensitive Policy Optimization for Safe Reinforcement Learning

    Jun 12, 2026Ayoub Belouadah, Sylvain Kubler, Yves Le TraonSafety Constraints

  2. Safety-Constrained Reinforcement Learning with Post-Training Reachability Verification for Robot Navigation

    May 13, 2026Qisong He, Xinmiao Huang, Jinwei Hu +4Safety ConstraintsRobot Navigation

  3. TraCeS: Learning Per-Timestep Constraint-Violation Credit from Sparse Trajectory-Level Labels

    Apr 17, 2025Siow Meng Low, Ze Gong, Akshat KumarSafety ConstraintsTrajectory-Level Credit