Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression -- but has been applied only to reward maximization. We propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic; outside, a recovery branch biases denoising toward regions with lower worst-case violation. On quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM attains the best or near-best task performance with low false-safe rates, whereas the primal-dual baseline admits more unsafe behavior and reachability-based baselines tend to be more conservative; on Safety-Gymnasium velocity tasks, SSM attains the lowest cost with competitive reward.
Figures & tables
Method family
Online
Safety semantics
Policy class
Actor update
Primal–dual safe RL
✓
Expected cost
Gaussian / deterministic
Primal–dual PG
HJ / reachability safe RL
✓
State-wise / HJ
Gaussian / deterministic
Actor–critic / shielded
Reward-only diffusion RL
✓
Reward only
Diffusion / flow
Score / flow matching
Offline safe diffusion
—
State-wise / HJ
Diffusion
Offline guided regression
SSM (Ours)
✓
State-wise / HJ
Diffusion
HJ-gated score matching
Table 1: Positioning of SSM by design properties. SSM combines online actor–critic training, state-wise HJ safety, a diffusion-policy class, and an HJ-gated score-matching actor update.
Figure 1: SSM on a planar-quadrotor stabilize-and-avoid task. Trajectory colors indicate the two route modes learned by the policy rather than safe and unsafe outcomes: green trajectories pass to the left of the obstacles, and orange trajectories pass to the right. (a) Rollouts from randomly sampled initial states. (b) Rollouts from the nominal, symmetric initial state (blue triangle), perturbed within the small neighborhood shown by the light-blue box; from nearly the same initial state, the policy takes either route to the goal. The green star marks the goal center, and the surrounding green disk is the goal region, whose entry counts as a successful episode. Red crosses mark collisions. The background is a two-dimensional slice of the learned HJ value Vh at fixed z˙=−0.5 , and the dashed contour around the gray regions is its zero level set Vh=0 , the estimated boundary of the viable set {s:Vh(s)≤0} .
Figure 2: Benchmark environments. Renderings of the Quad2D, Quad3D, and F16 tasks.
Task
Metric
RESPO
CAL
RAC
ALGD
SSM (ours)
Swimmer
Reward
37±1
31±7
75±57
49±4
45±6
Cost
8.4±1.3
27.9±21.3
14.5±28.4
6.5±2.4
0.5±0.4
HalfCheetah
Reward
2323±323
2631±107
5014±3781
2695±99
2754±13
Cost
8.4±3.9
28.3±12.5
391.7±536.3
42.0±39.7
0.0±0.1
Tasks within budget
2/2
0/2
1/2
1/2
2/2
Table 2: Safety-Gymnasium velocity tasks. Final episodic reward and cost (mean ± SD over five seeds). The cost counts the steps whose velocity exceeds the task threshold; a method is within budget on a task when its mean cost over the five seeds is at most d=25 , the budget that CAL and ALGD train against. RESPO trains for 10M environment steps and the others for 1M. Red: mean cost above budget. Among entries within budget, bold blue and bold green mark the highest and second-highest reward. SSM uses the posterior-target implementation of Appendix F.6 .
Table 3: Shared SSM hyperparameters used on all three benchmarks (Quad2D, Quad3D, F16).
Task
Actor hidden
Training steps
Warmup
Episode horizon
γr
Quad2D tracking
(512,512)
2×106
5,000
360
0.99
Quad3D stabilization
(512,512,512)
1×106
5,000
500
0.99
F16 stabilization
(512,512,512)
2×106
20,000
640
0.995
Appendix
Table 4: Task-specific settings. The denoiser actor uses three hidden layers on the higher-dimensional Quad3D and F16 tasks; the critic architecture is shared across tasks (Appendix F.1 ).
Table 5: Full Quad3D ablation results ( n=500 evaluation episodes). The top block reports estimator-level and candidate-source baselines; the bottom block sweeps the local Gaussian-proposal grid K∈{8,16,32}×ση∈{0.1,0.3,0.8} . Shaded row: the canonical SSM configuration used in the main results. Boldface within the sampled- Qh block: the lowest terminal tracking ℓ1 (default K=8,ση=0.3 ) and the safety-clean cell (K,ση)=(32,0.3) , which is the only sampled-gate cell that achieves zero observed violation across all three safety metrics at the lowest accompanying ℓ1 .
Adam; actor 10−4 with global-norm clipping at 1.0 ; critics 3×10−4 ; constant rates
Batch size
critics 256 ; actor 64
Discounts γr , γh ; target EMA rate
0.99 , 0.99 ; 0.005
Gate candidates K ; proposal scale ση
8 ; 0.3
Appendix
Table 6: SSM settings on the velocity tasks. Settings not listed follow Algorithm 1 .
Method
Metric
25%
50%
75%
100%
HalfCheetah velocity
SSM
Reward
2675±25
2724±27
2744±25
2754±13
Cost
0.0±0.0
0.1±0.1
0.0±0.0
0.0±0.1
RAC
Reward
1949±2579
3356±2977
3720±3456
5014±3781
Cost
195.6±437.3
196.0±438.3
196.5±438.5
391.7±536.3
CAL
Reward
2451±175
2530±136
2534±193
2631±107
Appendix
Table 7: Safety-Gymnasium results at 25, 50, 75, and 100% of each method’s horizon (1M steps; 10M for RESPO): episodic reward and cost, mean ± SD over five seeds. Asterisks mark values interpolated per seed between the two neighboring evaluated checkpoints. The row RESPO (0.25–1M) reads the RESPO runs at the step counts of the other methods.
αr
β
Mq
Reward ↑
Cost ↓
Goals ↑
P(Cep=0)↑
1
5
20
36.10
1.00
18.65
93.5%
0.25
5
20
35.39
0.52
18.19
96.7%
4
5
20
36.03
1.05
18.57
93.3%
1
1.25
20
35.99
0.83
18.59
95.2%
1
20
20
35.99
0.97
18.39
92.7%
1
5
5
35.55
0.61
18.33
94.3%
Appendix
Table 8: Sensitivity of SSM to αr , β , and Mq on SafetyCarGoal1-v0 . The first row is the default; each other row changes one coefficient.
Method
Budget d
Reward ↑
Cost ↓
Goals ↑
QSM-Lag
5
36.25
1.58
18.73
SAC-Lag
5
30.89
8.25
15.83
QSM-Lag
0
36.34
1.27
18.90
SAC-Lag
0
29.01
5.67
14.67
SSM (ours)
—
36.10
1.00
18.65
Appendix
Table 9: Diffusion versus Gaussian actors under the same expected-cost formulation on SafetyCarGoal1-v0 , with the summary of Table 8 .