Termination-based constrained reinforcement learning is attractive for safety-critical robotic deployments: it avoids online optimization at inference, scales easily to many constraints via a single scalar per constraint, and is simpler to implement than commonly used Lagrangian methods. Instead of pricing violations through summed cost penalties, this approach makes violations structurally unprofitable by shortening the effective horizon for each violation. We identify a structural failure mode of this method class on terminal-navigation tasks: reaching a precise goal configuration while satisfying safety constraints that tighten along the final approach. When the goal sits inside the region close to where the constraints become active, the survival-weighted objective makes dwelling outside the goal region strictly preferable to entering, producing high constraint compliance with low task completion. We formalize this pathology and show that a minimal augmentation to off-the-shelf RL algorithms like PPO resolves this ``feasibility collapse''. We empirically demonstrate that adding a dense per-step success signal via an auxiliary value critic improves the task completion rate while maintaining safety-critical constraint compliance. Validation across two representative spacecraft platforms: a 6U-CubeSat spanning the mass and degree-of-freedom envelope of operational proximity operations, and a floating platform testbed for zero-shot sim-to-real transfer in our laboratory, supports the generality of these findings.
Figures & tables
Figure 1: Left: Ablation Pareto scatter Each point is one reward × success indicator configuration; For floating platform, results show that SCoCaT occupies the ideal region (compliance ≥0.97 , GCO >0.5 ); CaT is compliant but collapses to near-zero task performance. Stars are the corresponding 6U CubeSat runs under the dense reward. Right: Constraint compliance vs. task success on two platforms. Each radar axis shows docking-corridor entry rate or per-constraint compliance ( 1−violation rate ); larger area is better. For the Floating Platform, PPO enters the corridor but violates safety constraints. CaT satisfies constraints but fails to enter the corridor (Section 3.1 ). SCoCaT achieves high corridor-entry and high constraint compliance. The same pattern holds on the 6U CubeSat: CaT enters the corridor in 3.7% of episodes and never completes the task, while SCoCaT reaches 79.4% task completion rate (TCR; Section 3.4 ) at 0.945 compliance against CaT’s 0.971 , a gap carried entirely by line-of-sight, the one constraint our tiering classifies as soft. PPO is unconstrained and buys its higher success at the lowest compliance of any method ( 0.839 ), with 12× the docking-velocity and 3× the angular-rate violation rate of SCoCaT.
Figure 2: Sim-to-real docking evaluation (Section 4.1 ). Left: Hardware setup and target interface. Right: SCoCaT Simulation result for Floating Platform showing average corridor occupancy.
Task
Constraint violation rate ( ↓ )
In-corridor velocity ( ↓ )
Method
df (cm) ↓
Hold (ts) ↑
SR
LoS
AngVel
Actions
vG (m/s)
ωG (rad/s)
PPO
2.4±2.1
41.8±12.3
4/4
0.855±0.174
0.592±0.045
0.012±0.019
0.022±0.002
0.111±0.006
CaT
4.0±2.2
11.5±3.7
0/4
0.121±0.136
0.628±0.160
0.000±0.000
0.012†
0.039†
SCoCaT
1.0±0.9
23.0±9.4
3/4
0.153±0.151
0.682±0.092
0.000±0.000
0.024±0.001
0.142±0.020
† n=1 run entered the corridor; single-trial value.
Table 1: Sim-to-real results on the floating platform testbed (4 runs per method, zero-shot). Constraint violation rates are the fraction of timesteps in violation ( ↓ ). vG and ωG : mean linear and angular body velocity measured while inside the 10 mm goal corridor. PPO uses the same terminal-proximity reward as CaT and SCoCaT (fair comparison). Bold : best per column.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Regime
Eff. pmax (dock / ang / fuel)
Dock.Vel. ↓
Fuel ↓
LoS ↓
Soft
0.075 / 0.05 / 0.025
1.1±0.6
63.3±3.0
98.7±0.3
Mid
0.30 / 0.20 / 0.10
0.5±0.3
39.7±5.6
98.5±0.3
Hard
1.0 / 1.0 / 1.0
0.1±0.2
3.2±0.8
84.0±0.7
Appendix
Table 2: Per-constraint violation rates (%, ↓ ) under varying pmax severity on the Floating Platform (SCoCaT, 3 seeds). The sweep varies pmax for docking velocity, angular rate and fuel only; effective values are shown in that order. Line-of-sight is not swept, so the LoS column reports the rate under each swept combination. These runs use 3 seeds and a constraint implementation that predates the one used for all other tables; only the trend across severity levels is comparable to Table 9 , where the default configuration gives 0.066 .
Threshold
pimax
ID
Constraint
FP
CS
FP
CS
Definition
Cbounds
Out-of-bounds
∥p∥2>8m
∥p∥2>8m
1.0
1.0
Exits workspace boundary.
Ccoll
Collision
Fcontact>0.1N
–
1.0
–
Contact force with target.
Cvd
Docking velocity
∥v∥2>0.05m/s
∥v∥2>0.1m/s
0.3
0.3
Linear speed within 0.5m of port.
Cω
Angular velocity
∣ω∣>0.1rad/s
∣ω∣>0.3rad/s
0.2
0.1
Body angular rate during approach.
Cfuel
Fuel consumption
∥u∥1>5.0
∥u∥1>6.0
0.1
0.1
ℓ1 -norm of thrust commands per step.
Appendix
Table 3: Active CaT constraints for both platforms. pimax is the maximum per-step termination probability at full violation; δt scales linearly with violation magnitude [ 5 ] . Hard constraints trigger immediate resets. FP = Floating Platform; CS = CubeSat. “–”: constraint inactive on that platform.
Table 7: Floating Platform success under the declared criterion, 5 seeds, dense reward. TCR and Compl. reproduce Tables 8 and 9 ; Hold and Decl. are the consecutive-dwell and full criteria of Section 3.4 . TCR credits PPO where neither baseline meets the criterion.
Method
Ent. ↑
TCR ↑
GCO ↑
Pos. ∗ (mm) ↓
Hdg. ∗ ( ∘ ) ↓
Dwell ∗ (mm) ↓
App. (m/s) ↓
PPO
0.908± 0.154
0.498± 0.488
0.129± 0.123
29.8± 19.7
0.735± 0.074
29.2± 18.5
0.039± 0.019
PPO (TP)
1.000
0.997± 0.003
0.304± 0.053
11.6± 1.5
0.709± 0.044
11.4± 1.3
0.026± 0.004
PPO (sparse)
0.000
0.000
0.000
–
–
–
–
CaT
0.538± 0.374
0.004± 0.008
0.014± 0.015
37.8± 12.2
0.843± 0.107
33.7± 9.2
0.016± 0.003
CaT (TP)
0.755± 0.155
0.006± 0.006
0.018± 0.009
26.5± 4.3
0.744± 0.069
24.1± 4.3
0.018± 0.003
CaT+Curr
0.132± 0.214
0.000
0.002± 0.003
42.2± 19.4
0.965± 0.293
28.1± 3.9
0.020± 0.006
Appendix
Table 8: Task-performance ablation on the Floating Platform (3-DoF, simulation). Each condition uses 5 independent seeds. Ent. : fraction of episodes that ever entered the goal corridor ( ↑ ). TCR : fraction of episodes in which the agent spent ≥ 10% of timesteps inside the goal corridor ( ↑ ). GCO : mean fraction of episode timesteps inside the goal corridor ( ↑ ). Pos. ∗ / Hdg. ∗ : mean final position (mm) and heading ( ∘ ) error for corridor-entering episodes ( ↓ ). Dwell ∗ : mean position error (mm) from first corridor entry to episode end, for corridor-entering episodes ( ↓ ). App. : mean planar approach speed (m/s) at the step before first corridor entry ( ↓ ). Bold : best per column among completed runs. Reward: dense (default), TP = terminal-proximity, sparse = end-of-episode binary. Signal: binary (default), cont. = continuous.
Method
Compl. ↑
Dock.Vel. ↓
Ang.Vel. ↓
LoS ↓
Fuel ↓
PPO
0.803± 0.011
0.070± 0.037
0.004± 0.001
0.832± 0.025
0.276± 0.048
PPO (TP)
0.798± 0.014
0.044± 0.003
0.007± 0.001
0.822± 0.094
0.340± 0.052
PPO (sparse)
0.646± 0.136
0.006± 0.005
0.657± 0.400
0.829± 0.007
0.492± 0.490
CaT
0.994
0.001
0.000
0.036± 0.003
0.001
CaT (TP)
0.994± 0.001
0.001
0.000
0.035± 0.003
0.001
CaT+Curr
0.986± 0.010
0.002
0.006± 0.001
0.055± 0.031
0.023± 0.036
Appendix
Table 9: Constraint-compliance ablation on the Floating Platform (3-DoF, simulation). Each condition uses 5 independent seeds. Compl. : 1−vˉ where vˉ is the mean per-timestep violation rate averaged over the active CaT constraints ( ↑ ). Remaining columns: individual per-timestep constraint violation rates ( ↓ ). The average runs over all six active constraints; out-of-bounds and collision are 0.000 throughout and are omitted from the columns, see Table 13 for collision results under the obstacle-avoidance setting. Bold : best per column among completed runs. Reward: dense (default), TP = terminal-proximity, sparse = end-of-episode binary. Signal: binary (default), cont. = continuous.
Method
Ent. ↑
TCR ↑
Hold ↑
Decl. ↑
GCO ↑
Pos. ∗ (mm) ↓
Ori. ∗ ( ∘ ) ↓
Dwell ∗ (mm) ↓
App. (m/s) ↓
PPO
1.000
1.000
1.000
0.931± 0.068
0.715± 0.011
3.6± 0.7
0.503± 0.058
3.8± 0.5
0.010± 0.001
CaT
0.037± 0.084
0.000
0.000
0.000
0.001± 0.001
48.9
2.245
49.0
0.013
SCoCaT (bin.)
0.794± 0.444
0.794± 0.444
0.694± 0.408
0.613± 0.359
0.397± 0.256
7.1± 0.8
4.577± 1.374
7.3± 0.6
0.012± 0.003
SCoCaT (cont.)
1.000
0.850± 0.192
0.469± 0.485
0.450± 0.489
0.456± 0.334
8.3± 3.7
1.395± 0.389
8.3± 3.6
0.016± 0.009
Appendix
Table 10: Task-performance results on the 6U CubeSat (6-DoF, simulation). 5 independent seeds; dense reward throughout. Scored at the goal corridor (position ≤ 10 mm, orientation ≤ 10 ∘ ) over the [0.01,4] m spawn range. Ent. : fraction of episodes that ever entered the goal corridor ( ↑ ). TCR : fraction of episodes in which the agent spent ≥ 10% of timesteps inside the goal corridor ( ↑ ). Hold : fraction of episodes with ≥ 100 consecutive in-corridor timesteps, pose only ( ↑ ). Decl. : Hold and the terminal-velocity gate ( ∥v∥<0.01 m/s, ∥ω∥<5∘ /s), the success criterion declared in Section 3.4 ( ↑ ). GCO : mean fraction of episode timesteps inside the goal corridor ( ↑ ). Pos. ∗ / Ori. ∗ : mean final position (mm) and orientation ( ∘ ) error for corridor-entering episodes ( ↓ ). Dwell ∗ : mean position error (mm) from first corridor entry to episode end ( ↓ ). App. : mean 3-D approach speed (m/s) at the step before first corridor entry ( ↓ ). PPO is unconstrained: it upper-bounds success and lower-bounds compliance (Table 11 ). Bold : best per column.
Method
Compl. ↑
Dock.Vel. ↓
Ang.Vel. ↓
LoS ↓
Fuel ↓
OoB ↓
PPO
0.839± 0.007
0.047± 0.003
0.118± 0.014
0.792± 0.039
0.008± 0.011
0.000
CaT
0.971± 0.007
0.004± 0.002
0.035± 0.023
0.135± 0.043
0.000± 0.001
0.000
SCoCaT (bin.)
0.945± 0.018
0.004± 0.005
0.039± 0.041
0.286± 0.082
0.000
0.000
SCoCaT (cont.)
0.941± 0.019
0.017± 0.019
0.072± 0.074
0.264± 0.020
0.001± 0.001
0.000
Appendix
Table 11: Constraint-compliance results on the 6U CubeSat (6-DoF, simulation). 5 independent seeds; dense reward throughout. Compl. : 1−vˉ where vˉ is the mean per-timestep violation rate over the six active CaT constraints ( ↑ ). Remaining columns: individual per-timestep constraint violation rates ( ↓ ). PPO is unconstrained and lower-bounds compliance; CaT’s high compliance is not a safety result, as a policy that never enters the corridor cannot violate in-corridor constraints (Table 10 ). Bold : best per column.
Method
Ent. ↑
SR ↑
GCO ↑
Pos. ∗ (mm) ↓
Hdg. ∗ ( ∘ ) ↓
Dwell ∗ (mm) ↓
App. (m/s) ↓
PPO
0.913± 0.147
0.237± 0.203
0.129± 0.122
29.2± 18.7
0.733± 0.074
25.7± 14.4
0.036± 0.010
CaT
0.458± 0.335
0.027± 0.028
0.012± 0.014
37.2± 11.7
0.881± 0.095
29.6± 8.9
0.017± 0.003
SCoCaT
0.779± 0.025
0.770± 0.026
0.538± 0.017
5.5± 0.6
0.303± 0.019
5.6± 0.5
0.039± 0.002
SCoCaT (cont.)
0.769± 0.041
0.590± 0.066
0.415± 0.037
7.5± 0.9
0.294± 0.031
7.7± 0.8
0.024± 0.003
Appendix
Table 12: Task-performance results on the Floating Platform with obstacle avoidance (GoToPoseObs, 3-DoF). A spherical obstacle is placed uniformly at random in the approach path each episode. Ent. : fraction of episodes that ever entered the goal corridor ( ↑ ). SR : episode success rate at the real-experiment criterion ( ≤ 0.01 m, ≤ 0.1 rad ≈ 5.7 ∘ ) ( ↑ ). Bold : best per column.
Method
Compl. ↑
Collision ↓
Dock.Vel. ↓
Ang.Vel. ↓
LoS ↓
Fuel ↓
PPO
0.804± 0.011
0.001
0.066± 0.034
0.004± 0.001
0.831± 0.025
0.274± 0.047
CaT
0.982± 0.003
0.000
0.001
0.000
0.107± 0.015
0.001
SCoCaT
0.965± 0.004
0.000
0.001
0.001
0.208± 0.022
0.002
SCoCaT (cont.)
0.966± 0.006
0.000
0.001
0.001± 0.001
0.203± 0.033
0.002
Appendix
Table 13: Constraint-compliance results on the Floating Platform with obstacle avoidance (GoToPoseObs, 3-DoF). The collision constraint Ccoll ( pmax=1.0 ) is active in addition to all constraints in Table 3 . Out-of-bounds violations are omitted (never violated). Bold : best per column.
School of Computing and Information Systems, Singapore Management University, Singapore · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Shenzhen, China