Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.
Figures & tables
Figure 1: Left: Ablation Pareto scatter Each point is one reward × success indicator configuration; For floating platform, results show that SCoCaT occupies the ideal region (compliance ≥0.97 , GCO >0.5 ); CaT is compliant but collapses to near-zero task performance. Stars are the corresponding 6U CubeSat runs under the dense reward. Right: Constraint compliance vs. task success on two platforms. Each radar axis shows docking-corridor entry rate or per-constraint compliance ( 1−violation rate ); larger area is better. For the Floating Platform, PPO enters the corridor but violates safety constraints. CaT satisfies constraints but fails to enter the corridor (Section 3.1 ). SCoCaT achieves high corridor-entry and high constraint compliance. The same pattern holds on the 6U CubeSat: CaT enters the corridor in 3.7% of episodes and never completes the task, while SCoCaT reaches 79.4% task completion rate (TCR; Section 3.4 ) at 0.945 compliance against CaT’s 0.971 , a gap carried entirely by line-of-sight, the one constraint our tiering classifies as soft. PPO is unconstrained and buys its higher success at the lowest compliance of any method ( 0.839 ), with 12× the docking-velocity and 3× the angular-rate violation rate of SCoCaT.
Figure 2: Sim-to-real docking evaluation (Section 4.1 ). Left: Hardware setup and target interface. Right: SCoCaT Simulation result for Floating Platform showing average corridor occupancy.
Task
Constraint violation rate ( ↓ )
In-corridor velocity ( ↓ )
Method
df (cm) ↓
Hold (ts) ↑
SR
LoS
AngVel
Actions
vG (m/s)
ωG (rad/s)
PPO
2.4±2.1
41.8±12.3
4/4
0.855±0.174
0.592±0.045
0.012±0.019
0.022±0.002
0.111±0.006
CaT
4.0±2.2
11.5±3.7
0/4
0.121±0.136
0.628±0.160
0.000±0.000
0.012†
0.039†
SCoCaT
1.0±0.9
23.0±9.4
3/4
0.153±0.151
0.682±0.092
0.000±0.000
0.024±0.001
0.142±0.020
† n=1 run entered the corridor; single-trial value.
Table 1: Sim-to-real results on the floating platform testbed (4 runs per method, zero-shot). Constraint violation rates are the fraction of timesteps in violation ( ↓ ). vG and ωG : mean linear and angular body velocity measured while inside the 10 mm goal corridor. PPO uses the same terminal-proximity reward as CaT and SCoCaT (fair comparison). Bold : best per column.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Regime
Eff. pmax (dock / ang / fuel)
Dock.Vel. ↓
Fuel ↓
LoS ↓
Soft
0.075 / 0.05 / 0.025
1.1±0.6
63.3±3.0
98.7±0.3
Mid
0.30 / 0.20 / 0.10
0.5±0.3
39.7±5.6
98.5±0.3
Hard
1.0 / 1.0 / 1.0
0.1±0.2
3.2±0.8
84.0±0.7
Appendix
Table 2: Per-constraint violation rates (%, ↓ ) under varying pmax severity on the Floating Platform (SCoCaT, 3 seeds). The sweep varies pmax for docking velocity, angular rate and fuel only; effective values are shown in that order. Line-of-sight is not swept, so the LoS column reports the rate under each swept combination. These runs use 3 seeds and a constraint implementation that predates the one used for all other tables; only the trend across severity levels is comparable to Table 9 , where the default configuration gives 0.066 .
Threshold
pimax
ID
Constraint
FP
CS
FP
CS
Definition
Cbounds
Out-of-bounds
∥p∥2>8m
∥p∥2>8m
1.0
1.0
Exits workspace boundary.
Ccoll
Collision
Fcontact>0.1N
–
1.0
–
Contact force with target.
Cvd
Docking velocity
∥v∥2>0.05m/s
∥v∥2>0.1m/s
0.3
0.3
Linear speed within 0.5m of port.
Cω
Angular velocity
∣ω∣>0.1rad/s
∣ω∣>0.3rad/s
0.2
0.1
Body angular rate during approach.
Cfuel
Fuel consumption
∥u∥1>5.0
∥u∥1>6.0
0.1
0.1
ℓ1 -norm of thrust commands per step.
Appendix
Table 3: Active CaT constraints for both platforms. pimax is the maximum per-step termination probability at full violation; δt scales linearly with violation magnitude [ 5 ] . Hard constraints trigger immediate resets. FP = Floating Platform; CS = CubeSat. “–”: constraint inactive on that platform.
Table 7: Floating Platform success under the declared criterion, 5 seeds, dense reward. TCR and Compl. reproduce Tables 8 and 9 ; Hold and Decl. are the consecutive-dwell and full criteria of Section 3.4 . TCR credits PPO where neither baseline meets the criterion.
Method
Ent. ↑
TCR ↑
GCO ↑
Pos. ∗ (mm) ↓
Hdg. ∗ ( ∘ ) ↓
Dwell ∗ (mm) ↓
App. (m/s) ↓
PPO
0.908± 0.154
0.498± 0.488
0.129± 0.123
29.8± 19.7
0.735± 0.074
29.2± 18.5
0.039± 0.019
PPO (TP)
1.000
0.997± 0.003
0.304± 0.053
11.6± 1.5
0.709± 0.044
11.4± 1.3
0.026± 0.004
PPO (sparse)
0.000
0.000
0.000
–
–
–
–
CaT
0.538± 0.374
0.004± 0.008
0.014± 0.015
37.8± 12.2
0.843± 0.107
33.7± 9.2
0.016± 0.003
CaT (TP)
0.755± 0.155
0.006± 0.006
0.018± 0.009
26.5± 4.3
0.744± 0.069
24.1± 4.3
0.018± 0.003
CaT+Curr
0.132± 0.214
0.000
0.002± 0.003
42.2± 19.4
0.965± 0.293
28.1± 3.9
0.020± 0.006
Appendix
Table 8: Task-performance ablation on the Floating Platform (3-DoF, simulation). Each condition uses 5 independent seeds. Ent. : fraction of episodes that ever entered the goal corridor ( ↑ ). TCR : fraction of episodes in which the agent spent ≥ 10% of timesteps inside the goal corridor ( ↑ ). GCO : mean fraction of episode timesteps inside the goal corridor ( ↑ ). Pos. ∗ / Hdg. ∗ : mean final position (mm) and heading ( ∘ ) error for corridor-entering episodes ( ↓ ). Dwell ∗ : mean position error (mm) from first corridor entry to episode end, for corridor-entering episodes ( ↓ ). App. : mean planar approach speed (m/s) at the step before first corridor entry ( ↓ ). Bold : best per column among completed runs. Reward: dense (default), TP = terminal-proximity, sparse = end-of-episode binary. Signal: binary (default), cont. = continuous.
Method
Compl. ↑
Dock.Vel. ↓
Ang.Vel. ↓
LoS ↓
Fuel ↓
PPO
0.803± 0.011
0.070± 0.037
0.004± 0.001
0.832± 0.025
0.276± 0.048
PPO (TP)
0.798± 0.014
0.044± 0.003
0.007± 0.001
0.822± 0.094
0.340± 0.052
PPO (sparse)
0.646± 0.136
0.006± 0.005
0.657± 0.400
0.829± 0.007
0.492± 0.490
CaT
0.994
0.001
0.000
0.036± 0.003
0.001
CaT (TP)
0.994± 0.001
0.001
0.000
0.035± 0.003
0.001
CaT+Curr
0.986± 0.010
0.002
0.006± 0.001
0.055± 0.031
0.023± 0.036
Appendix
Table 9: Constraint-compliance ablation on the Floating Platform (3-DoF, simulation). Each condition uses 5 independent seeds. Compl. : 1−vˉ where vˉ is the mean per-timestep violation rate averaged over the active CaT constraints ( ↑ ). Remaining columns: individual per-timestep constraint violation rates ( ↓ ). The average runs over all six active constraints; out-of-bounds and collision are 0.000 throughout and are omitted from the columns, see Table 13 for collision results under the obstacle-avoidance setting. Bold : best per column among completed runs. Reward: dense (default), TP = terminal-proximity, sparse = end-of-episode binary. Signal: binary (default), cont. = continuous.
Method
Ent. ↑
TCR ↑
Hold ↑
Decl. ↑
GCO ↑
Pos. ∗ (mm) ↓
Ori. ∗ ( ∘ ) ↓
Dwell ∗ (mm) ↓
App. (m/s) ↓
PPO
1.000
1.000
1.000
0.931± 0.068
0.715± 0.011
3.6± 0.7
0.503± 0.058
3.8± 0.5
0.010± 0.001
CaT
0.037± 0.084
0.000
0.000
0.000
0.001± 0.001
48.9
2.245
49.0
0.013
SCoCaT (bin.)
0.794± 0.444
0.794± 0.444
0.694± 0.408
0.613± 0.359
0.397± 0.256
7.1± 0.8
4.577± 1.374
7.3± 0.6
0.012± 0.003
SCoCaT (cont.)
1.000
0.850± 0.192
0.469± 0.485
0.450± 0.489
0.456± 0.334
8.3± 3.7
1.395± 0.389
8.3± 3.6
0.016± 0.009
Appendix
Table 10: Task-performance results on the 6U CubeSat (6-DoF, simulation). 5 independent seeds; dense reward throughout. Scored at the goal corridor (position ≤ 10 mm, orientation ≤ 10 ∘ ) over the [0.01,4] m spawn range. Ent. : fraction of episodes that ever entered the goal corridor ( ↑ ). TCR : fraction of episodes in which the agent spent ≥ 10% of timesteps inside the goal corridor ( ↑ ). Hold : fraction of episodes with ≥ 100 consecutive in-corridor timesteps, pose only ( ↑ ). Decl. : Hold and the terminal-velocity gate ( ∥v∥<0.01 m/s, ∥ω∥<5∘ /s), the success criterion declared in Section 3.4 ( ↑ ). GCO : mean fraction of episode timesteps inside the goal corridor ( ↑ ). Pos. ∗ / Ori. ∗ : mean final position (mm) and orientation ( ∘ ) error for corridor-entering episodes ( ↓ ). Dwell ∗ : mean position error (mm) from first corridor entry to episode end ( ↓ ). App. : mean 3-D approach speed (m/s) at the step before first corridor entry ( ↓ ). PPO is unconstrained: it upper-bounds success and lower-bounds compliance (Table 11 ). Bold : best per column.
Method
Compl. ↑
Dock.Vel. ↓
Ang.Vel. ↓
LoS ↓
Fuel ↓
OoB ↓
PPO
0.839± 0.007
0.047± 0.003
0.118± 0.014
0.792± 0.039
0.008± 0.011
0.000
CaT
0.971± 0.007
0.004± 0.002
0.035± 0.023
0.135± 0.043
0.000± 0.001
0.000
SCoCaT (bin.)
0.945± 0.018
0.004± 0.005
0.039± 0.041
0.286± 0.082
0.000
0.000
SCoCaT (cont.)
0.941± 0.019
0.017± 0.019
0.072± 0.074
0.264± 0.020
0.001± 0.001
0.000
Appendix
Table 11: Constraint-compliance results on the 6U CubeSat (6-DoF, simulation). 5 independent seeds; dense reward throughout. Compl. : 1−vˉ where vˉ is the mean per-timestep violation rate over the six active CaT constraints ( ↑ ). Remaining columns: individual per-timestep constraint violation rates ( ↓ ). PPO is unconstrained and lower-bounds compliance; CaT’s high compliance is not a safety result, as a policy that never enters the corridor cannot violate in-corridor constraints (Table 10 ). Bold : best per column.
Method
Ent. ↑
SR ↑
GCO ↑
Pos. ∗ (mm) ↓
Hdg. ∗ ( ∘ ) ↓
Dwell ∗ (mm) ↓
App. (m/s) ↓
PPO
0.913± 0.147
0.237± 0.203
0.129± 0.122
29.2± 18.7
0.733± 0.074
25.7± 14.4
0.036± 0.010
CaT
0.458± 0.335
0.027± 0.028
0.012± 0.014
37.2± 11.7
0.881± 0.095
29.6± 8.9
0.017± 0.003
SCoCaT
0.779± 0.025
0.770± 0.026
0.538± 0.017
5.5± 0.6
0.303± 0.019
5.6± 0.5
0.039± 0.002
SCoCaT (cont.)
0.769± 0.041
0.590± 0.066
0.415± 0.037
7.5± 0.9
0.294± 0.031
7.7± 0.8
0.024± 0.003
Appendix
Table 12: Task-performance results on the Floating Platform with obstacle avoidance (GoToPoseObs, 3-DoF). A spherical obstacle is placed uniformly at random in the approach path each episode. Ent. : fraction of episodes that ever entered the goal corridor ( ↑ ). SR : episode success rate at the real-experiment criterion ( ≤ 0.01 m, ≤ 0.1 rad ≈ 5.7 ∘ ) ( ↑ ). Bold : best per column.
Method
Compl. ↑
Collision ↓
Dock.Vel. ↓
Ang.Vel. ↓
LoS ↓
Fuel ↓
PPO
0.804± 0.011
0.001
0.066± 0.034
0.004± 0.001
0.831± 0.025
0.274± 0.047
CaT
0.982± 0.003
0.000
0.001
0.000
0.107± 0.015
0.001
SCoCaT
0.965± 0.004
0.000
0.001
0.001
0.208± 0.022
0.002
SCoCaT (cont.)
0.966± 0.006
0.000
0.001
0.001± 0.001
0.203± 0.033
0.002
Appendix
Table 13: Constraint-compliance results on the Floating Platform with obstacle avoidance (GoToPoseObs, 3-DoF). The collision constraint Ccoll ( pmax=1.0 ) is active in addition to all constraints in Table 3 . Out-of-bounds violations are omitted (never violated). Bold : best per column.
Safe reinforcement learning (Safe RL) aims to maximize expected return while satisfying safety constraints, typically modeled as Constrained Markov Decision Processes (CMDPs). While primal-dual methods scale well to deep RL, they often suffer from delayed constraint correction, leading to oscillatory behavior and prolonged safety violations. In this paper, we propose Constraint-Sensitive Policy Optimization (CSPO), a first-order primal-dual method that incorporates local constraint sensitivity into policy updates. CSPO augments the primal objective with a constraint-sensitive correction derived from the shortest signed distance to the safety boundary, enabling smarter recovery steps back to safety, compensating for delayed Lagrange multiplier updates, reducing oscillations near the boundary, and preserving the KKT solutions of the original constrained problem. Experiments on navigation and locomotion benchmarks demonstrate that CSPO achieves faster safety recovery and high reward preservation, resulting in higher constrained returns compared to state-of-the-art primal-dual and penalty-based methods
Safe navigation for mobile robots demands policies that remain reliable under the high-consequence perception uncertainty of cluttered environments. Yet most existing safe reinforcement learning (RL) methods assess safety through average cumulative cost. Such metrics can mask dangerous tail-risk behaviors. To address this, we propose a framework that trains risk-sensitive policies through Conditional Value-at-Risk (CVaR) constrained optimization on an off-policy TD3 backbone and evaluates their safety margins post-training through neural network reachability verification. During training, the policy is optimized under CVaR constraints on cumulative costs, promoting sensitivity to high-cost tail outcomes rather than average behavior alone. After training, we compute action reachable sets under bounded observation uncertainty using Taylor Model analysis, yielding a safety rate metric that quantifies the proportion of evaluated states at which the policy's reachable action set remains within prescribed safety margins. A key finding is that policies trained with CVaR constraints maintain larger safety margins from obstacles across evaluated states. This makes them significantly more amenable to formal reachability verification. Experiments across ten navigation scenarios and six baselines show that our method achieves a 98.3% success rate, the highest safety verification rate among all compared methods, while revealing that average cost rankings and reachability-based safety rankings can diverge. This indicates that reachability verification captures risks which are missed by empirical cost metrics alone. We further validate our approach on a physical Clearpath Jackal robot, demonstrating successful sim-to-real transfer.
Qisong He, Xinmiao Huang, Jinwei Hu +4
University of Liverpool, Liverpool, UK. · Universit´e Grenoble Alpes, Grenoble, France.
Ensuring safe behavior in reinforcement learning (RL) is challenging when safety constraints are implicit and cannot be densely measured. In many settings, supervision is limited to coarse approvals or rejections of whole trajectories (e.g., whether a rollout remained within an unknown safety threshold). We propose TraCeS (Trajectory-based Constraint Estimation for Safety), a method for learning per-timestep violation credit from such sparse trajectory-level labels. TraCeS trains a sequential violation estimator whose per-step credits factorize the predicted probability that a trajectory has not yet violated the constraint, and integrates this learned signal into constrained policy optimization. The method requires neither a known cost function nor a known threshold, and remains compatible with standard continuous-control algorithms. We provide a theoretical analysis of the approximation gap introduced by the learning objective, and demonstrate empirically that TraCeS improves constraint satisfaction and feedback efficiency over baselines across multiple continuous-control benchmarks, including long-horizon tasks and settings with noisy or inconsistent labels.
Siow Meng Low, Ze Gong, Akshat Kumar
School of Computing and Information Systems, Singapore Management University, Singapore · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China