Online safe reinforcement learning (RL) seeks policies that maximize reward while satisfying safety constraints. A popular line of research in safe RL relaxes safety to a soft expected-cost constraint and solves the resulting Constrained Markov Decision Process via primal-dual Lagrangian updates that only enforce safety on average. To address this limitation, hard, state-wise constraints are introduced and often imposed through Hamilton-Jacobi (HJ) reachability. Yet such constraints require solving different objectives in the feasible and infeasible regions: reward maximization in the former, recovery toward the feasible regions in the latter. The resulting target action distributions are inherently multimodal, and this structure poses a fundamental challenge for the Gaussian or deterministic actors used in existing HJ-based safe RL, which often collapse onto suboptimal modes. Diffusion policies provide the expressiveness needed to represent such distributions, and recent work on Q-score matching offers a route to training them for online RL by score regression -- but has been applied only to reward maximization. We propose Safe Score Matching (SSM), an off-policy actor-critic method that adapts Q-score matching to hard-constrained safe RL by gating a two-branch score target with HJ reachability: inside the feasible set, the denoising process degenerates to Q-score matching on actions classified as viable by the HJ critic; outside, a recovery branch biases denoising toward regions with lower worst-case violation. On quadrotor and fixed-wing trajectory-tracking and stabilize-and-avoid benchmarks, SSM attains the best or near-best task performance with low false-safe rates, whereas the primal-dual baseline admits more unsafe behavior and reachability-based baselines tend to be more conservative; on Safety-Gymnasium velocity tasks, SSM attains the lowest cost with competitive reward.
Figures & tables
Method family
Online
Safety semantics
Policy class
Actor update
Primal–dual safe RL
✓
Expected cost
Gaussian / deterministic
Primal–dual PG
HJ / reachability safe RL
✓
State-wise / HJ
Gaussian / deterministic
Actor–critic / shielded
Reward-only diffusion RL
✓
Reward only
Diffusion / flow
Score / flow matching
Offline safe diffusion
—
State-wise / HJ
Diffusion
Offline guided regression
SSM (Ours)
✓
State-wise / HJ
Diffusion
HJ-gated score matching
Table 1: Positioning of SSM by design properties. SSM combines online actor–critic training, state-wise HJ safety, a diffusion-policy class, and an HJ-gated score-matching actor update.
Figure 1: SSM on a planar-quadrotor stabilize-and-avoid task. Trajectory colors indicate the two route modes learned by the policy rather than safe and unsafe outcomes: green trajectories pass to the left of the obstacles, and orange trajectories pass to the right. (a) Rollouts from randomly sampled initial states. (b) Rollouts from the nominal, symmetric initial state (blue triangle), perturbed within the small neighborhood shown by the light-blue box; from nearly the same initial state, the policy takes either route to the goal. The green star marks the goal center, and the surrounding green disk is the goal region, whose entry counts as a successful episode. Red crosses mark collisions. The background is a two-dimensional slice of the learned HJ value Vh at fixed z˙=−0.5 , and the dashed contour around the gray regions is its zero level set Vh=0 , the estimated boundary of the viable set {s:Vh(s)≤0} .
Figure 2: Benchmark environments. Renderings of the Quad2D, Quad3D, and F16 tasks.
Task
Metric
RESPO
CAL
RAC
ALGD
SSM (ours)
Swimmer
Reward
37±1
31±7
75±57
49±4
45±6
Cost
8.4±1.3
27.9±21.3
14.5±28.4
6.5±2.4
0.5±0.4
HalfCheetah
Reward
2323±323
2631±107
5014±3781
2695±99
2754±13
Cost
8.4±3.9
28.3±12.5
391.7±536.3
42.0±39.7
0.0±0.1
Tasks within budget
2/2
0/2
1/2
1/2
2/2
Table 2: Safety-Gymnasium velocity tasks. Final episodic reward and cost (mean ± SD over five seeds). The cost counts the steps whose velocity exceeds the task threshold; a method is within budget on a task when its mean cost over the five seeds is at most d=25 , the budget that CAL and ALGD train against. RESPO trains for 10M environment steps and the others for 1M. Red: mean cost above budget. Among entries within budget, bold blue and bold green mark the highest and second-highest reward. SSM uses the posterior-target implementation of Appendix F.6 .
Table 3: Shared SSM hyperparameters used on all three benchmarks (Quad2D, Quad3D, F16).
Task
Actor hidden
Training steps
Warmup
Episode horizon
γr
Quad2D tracking
(512,512)
2×106
5,000
360
0.99
Quad3D stabilization
(512,512,512)
1×106
5,000
500
0.99
F16 stabilization
(512,512,512)
2×106
20,000
640
0.995
Appendix
Table 4: Task-specific settings. The denoiser actor uses three hidden layers on the higher-dimensional Quad3D and F16 tasks; the critic architecture is shared across tasks (Appendix F.1 ).
Table 5: Full Quad3D ablation results ( n=500 evaluation episodes). The top block reports estimator-level and candidate-source baselines; the bottom block sweeps the local Gaussian-proposal grid K∈{8,16,32}×ση∈{0.1,0.3,0.8} . Shaded row: the canonical SSM configuration used in the main results. Boldface within the sampled- Qh block: the lowest terminal tracking ℓ1 (default K=8,ση=0.3 ) and the safety-clean cell (K,ση)=(32,0.3) , which is the only sampled-gate cell that achieves zero observed violation across all three safety metrics at the lowest accompanying ℓ1 .
Adam; actor 10−4 with global-norm clipping at 1.0 ; critics 3×10−4 ; constant rates
Batch size
critics 256 ; actor 64
Discounts γr , γh ; target EMA rate
0.99 , 0.99 ; 0.005
Gate candidates K ; proposal scale ση
8 ; 0.3
Appendix
Table 6: SSM settings on the velocity tasks. Settings not listed follow Algorithm 1 .
Method
Metric
25%
50%
75%
100%
HalfCheetah velocity
SSM
Reward
2675±25
2724±27
2744±25
2754±13
Cost
0.0±0.0
0.1±0.1
0.0±0.0
0.0±0.1
RAC
Reward
1949±2579
3356±2977
3720±3456
5014±3781
Cost
195.6±437.3
196.0±438.3
196.5±438.5
391.7±536.3
CAL
Reward
2451±175
2530±136
2534±193
2631±107
Appendix
Table 7: Safety-Gymnasium results at 25, 50, 75, and 100% of each method’s horizon (1M steps; 10M for RESPO): episodic reward and cost, mean ± SD over five seeds. Asterisks mark values interpolated per seed between the two neighboring evaluated checkpoints. The row RESPO (0.25–1M) reads the RESPO runs at the step counts of the other methods.
αr
β
Mq
Reward ↑
Cost ↓
Goals ↑
P(Cep=0)↑
1
5
20
36.10
1.00
18.65
93.5%
0.25
5
20
35.39
0.52
18.19
96.7%
4
5
20
36.03
1.05
18.57
93.3%
1
1.25
20
35.99
0.83
18.59
95.2%
1
20
20
35.99
0.97
18.39
92.7%
1
5
5
35.55
0.61
18.33
94.3%
Appendix
Table 8: Sensitivity of SSM to αr , β , and Mq on SafetyCarGoal1-v0 . The first row is the default; each other row changes one coefficient.
Method
Budget d
Reward ↑
Cost ↓
Goals ↑
QSM-Lag
5
36.25
1.58
18.73
SAC-Lag
5
30.89
8.25
15.83
QSM-Lag
0
36.34
1.27
18.90
SAC-Lag
0
29.01
5.67
14.67
SSM (ours)
—
36.10
1.00
18.65
Appendix
Table 9: Diffusion versus Gaussian actors under the same expected-cost formulation on SafetyCarGoal1-v0 , with the summary of Table 8 .
Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
Hikmet Simsir, Kutay Demiray, Ozgur S. Oguz
Department of Computer Engineering, Bilkent University
Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a single episode. While diffusion-based planners enable flexible trajectory generation, existing guidance schemes often treat reward improvement and constraint satisfaction as competing gradient objectives, which can lead to unreliable safety compliance under cost limits. We reinterpret adaptive safe trajectory generation as sampling from a constrained trajectory distribution, where the budget restricts the trajectory region, and reward shapes preferences within that region. This perspective motivates Safe Decoupled Guidance Diffusion (SDGD), which conditions classifier-free guidance on the cost limit to bias sampling toward trajectories satisfying the specified limit, while using reward-gradient guidance to refine trajectories for higher return. Because direct reward guidance can increase return while also steering samples toward trajectories with higher cumulative cost, we introduce Feasible Trajectory Relabeling (FTR) to reshape reward targets and discourage such directions. We further provide a first-order sampling-time analysis showing that FTR suppresses reward-induced cost drift under a prefix-restorative alignment condition. Extensive evaluations on the DSRL benchmark show that SDGD achieves the strongest safety compliance among baselines, satisfying the constraint on 94.7% of tasks (36/38), while obtaining the highest reward among safe methods on 21 tasks.
Rufeng Chen, Zhaofan Zhang, Zhejiang Yang +2
1The Hong Kong University of Science and Technology (Guangzhou) · 2Jilin University
Offline safe reinforcement learning (RL) seeks reward-maximizing policies from static datasets under strict safety constraints. Existing methods often rely on soft expected-cost objectives or iterative generative inference, which can be insufficient for safety-critical real-time control. We propose Safe Flow Q-Learning (SafeFQL), which extends FQL to safe offline RL by combining a Hamilton--Jacobi reachability-inspired safety value function with an efficient one-step flow policy. SafeFQL learns the safety value via a self-consistency Bellman recursion, trains a flow policy by behavioral cloning, and distills it into a one-step actor for reward-maximizing safe action selection without rejection sampling at deployment. Empirically, SafeFQL trades modestly higher offline training cost for substantially lower inference latency than diffusion-style safe generative baselines, which is advantageous for real-time safety-critical deployment. Across boat navigation, and Safety Gymnasium MuJoCo tasks, SafeFQL matches or exceeds prior offline safe RL performance while substantially reducing constraint violations.