Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
Figures & tables
Figure 1: (A) Evaluation paths obtained from fully trained policies on the same fixed PointGoal1 layout. See Figure 6 for clearance distribution of these paths. (B) Illustrative diagram for the RADPO algorithm.
Figure 2: Training performance across ten continuous-control safety tasks. Episodic return (top rows) and episodic cost (bottom rows) for RADPO and baseline methods. Curves show empirical means across six random seeds; shaded regions denote one standard deviation. Dashed horizontal lines indicate the safety budget d=25 .
Figure 3: Late-training return versus safety cost. Markers indicate mean performance across seeds, computed over the final ten logged training points per run. Error bars represent one standard deviation across seed averages. Optimal performance lies toward the top-left quadrant (high return, low cost). Vertical dashed lines denote the budget d=25 and the shaded region is infeasible.
Figure 4: Ablations on safety-signal formulation and penalty placement. Training return (top) and episodic cost (bottom). Curves depict means across evaluated seeds with one standard deviation shaded. Solid lines show smoothing over raw trajectories (faint curves). Inset axes provide magnified views around the cost budget d=25 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Ablations on hyperparameter values. Ablation results for reachability discount γc , shaping coefficient β and candidates per state K . Curves depict means across evaluated seeds with one standard deviation shaded.
Time per environment transition (ms)
Method
Sample acquisition
Critic optimization
Diffusion-policy optimization
Other learning
Overall
ALGD
3.298
16.057
78.892
41.524
139.771
Lagrangian-style
0.418
0.545
0.459
3.387
4.809
RADPO
0.616
0.423
0.406
2.697
4.142
Appendix
Table 1: Computational cost on PointGoal1. All entries are ms per environment step.
Figure 6: Distribution of states for trajectories in Figure 1 .A.
Family
Task
Horizon T
Env. steps
Eval. every
Velocity
SafetyAntVelocity-v1
1000
1M
25k
SafetyHalfCheetahVelocity-v1
1000
1M
25k
SafetyHopperVelocity-v1
1000
1M
25k
SafetyHumanoidVelocity-v1
1000
3M
25k
SafetySwimmerVelocity-v1
1000
1M
25k
SafetyWalker2dVelocity-v1
1000
1M
25k
Appendix
Table 2: Tasks, interaction budgets and evaluation cadence.
Group
Hyperparameter
Value
Networks
Critic / reachability-critic MLP
3×256 , Mish
Diffusion actor MLP
3×256 , Mish
Denoising steps
20
Noise schedule
cosine
Reward critics (twin)
2
Optimisation
Optimiser
Adam
Appendix
Table 3: RADPO hyperparameters held constant across all ten tasks.
Task
UTD
α0
σ0
β
ηλ
Ant
1.0
0.2
0.5
1.5
0.010
HalfCheetah
1.0
0.2
0.5
1.5
0.010
Hopper
0.6
0.2
0.5
0.8
0.002
Humanoid
1.0
0.2
0.1
2.4
0.120
Swimmer
0.2
0.2
0.5
0.025
0.100
Walker2d
0.2
0.2
0.5
1.5
0.010
Appendix
Table 4: Per-task RADPO hyperparameters. UTD is the number of optimiser updates per collected transition, α0 the initial exploration temperature ( σexp=σ0α0 at initialisation), σ0 the exploration scale, β the shaping coefficient and ηλ the controller step size. Everything else is as in Table 3 .
Hyperparameter
Velocity / Circle
Goal
Critic / actor MLP
3×256 , GELU
2×256 , ReLU
Actor learning rate
3×10−4
5×10−6
Critic learning rate
3×10−4
1×10−3
Entropy temperature α
learned, α0=0.2
fixed 10−5
α learning rate
3×10−4
—
Policy delay
1
2
Appendix
Table 5: SAC-Lagrangian. “Velocity” covers the six velocity tasks and the two Circle tasks; “Goal” covers CarGoal1 and PointGoal1.
Task
k
c
Multiplier LR
Internal d
Extra
Ant
0.0
0.01
1×10−5
25
—
HalfCheetah
0.0
0.3
1×10−4
25
Qc≥0
Hopper
0.5
0.2
1×10−5
25
done-mask, Qc≥0
Humanoid
0.0
0.01
1×10−5
25
done-mask
Swimmer
0.0
0.01
1×10−5
25
—
Walker2d
0.0
0.2
1×10−4
25
done-mask
Appendix
Table 6: CAL per-task settings. k is the cost-critic UCB coefficient, c the rectification coefficient. “done-mask” indicates that termination is masked in the cost backup; “ Qc≥0 ” that the cost critic is floored at zero.
Hyperparameter
Velocity
Humanoid
Navigation
Policy / value MLP
2×64 , tanh
2×64 , tanh
2×64 , tanh
Learning rate
3×10−4
3×10−4
3×10−4
Steps per epoch
2048
6000
20000
Epochs
500
500
150
Total environment steps
1.024M
3M
3M
Update iterations / epoch
40
40
40
Appendix
Table 7: PPO-Lagrangian.
Task
Budget
Rollout
Actor LR
σ0
Target KL
Ant
1M
10,000
6×10−4
0.2
0.03
HalfCheetah
1M
10,000
6×10−4
0.2
0.03
Hopper
1M
10,000
6×10−4
0.2
0.03
Humanoid
3M
5,000
1×10−4
0.2
0.03
Swimmer
1M
10,000
6×10−4
0.2
0.03
Walker2d
1M
10,000
3×10−4
0.2
0.03
Appendix
Table 8: RESPO per-task settings. Budgets and rollout sizes count total environment transitions across four workers. Actor LR and σ0 are the initial learning rate and policy standard deviation. Observation normalization is used only for Humanoid.
Released code
This paper
Environments
Safety-Gym ( Safexp-* ), MuJoCo
Safety Gymnasium (all ten tasks)
Navigation layout
fixed for every episode and seed
resampled every episode
Episode horizon
400 (navigation)
1000
Goal reached
episode terminates
new goal spawned, episode continues
Cost budget
10 (navigation) / raw velocity
d=25 on all tasks
Evaluation
1 episode
20 deterministic episodes
Appendix
Table 9: ALGD: released training protocol versus the protocol used here.
Group
Hyperparameter
Value
Networks
Reward critics (twin)
2×256 , ReLU
Cost-critic ensemble
4 members, 2×256 , SiLU
Score / noise network
3×128 , ReLU, time embedding
Denoising steps K
5
Diffusion
DDPM noise schedule
linear β∈[10−4,0.02]
VE-SDE noise scale
geometric σ∈[0.01,1.0]
Appendix
Table 10: ALGD hyperparameters (released defaults, shared by the DDPM and VE-SDE agents).
Task
Agent
UTD
Ant
DDPM
10
HalfCheetah
DDPM
10
Hopper
DDPM
10
Humanoid
DDPM
10
Swimmer
DDPM
10
Walker2d
DDPM
10
Appendix
Table 11: ALGD per-task settings. UTD is gradient updates per environment step.
Task
Arm
Cost critic
Shaping target
β
CarGoal1
RADPO
reachability
Vreach(s′)
0.05
discounted critic
discounted
Vc(s′)
0.01
cost-only shaping
-
ct
0.05
lag-style shaping
discounted
none (cost in weights)
—
PointCircle1
RADPO
reachability
Vreach(s′)
0.05
discounted critic
discounted
Vc(s′)
0.01
Appendix
Table 12: Ablation arms. Every arm shares its task’s ηλ , σ0 , α0 , UTD and interaction budget with the RADPO arm beside it, and all settings of Table 3 .
Figure 7: Deterministic-evaluation return and cost across the ten tasks. Evaluation return (upper row of each block) and evaluation cost (lower row), measured over 20 deterministic episodes per checkpoint; the training-rollout counterpart is Figure 2 . Curves are means over the runs of an arm with one standard deviation shaded, and dashed horizontal lines mark the budget d=25 .
Offline reinforcement learning enables policy learning from fixed datasets without additional environment interaction, making it appealing for safety-critical applications where online exploration is costly or unsafe. Diffusion-based decision-making methods have recently achieved strong performance in offline RL by modeling rich, multimodal trajectory distributions. However, existing diffusion planners are typically risk-neutral and therefore may overlook rare but catastrophic outcomes that are crucial in real-world deployment. In this work, we propose RS-Diffuser, a risk-sensitive offline diffusion planning framework that combines diffusion-based trajectory generation with distributional value critics. RS-Diffuser learns a diffusion planner over future state trajectories, a separate inverse dynamics model for action decoding, and a Monte Carlo distributional critic that estimates the full return distribution of candidate plans through quantile regression. At sampling time, we incorporate a risk-sensitive guidance signal into the denoising process, using gradients computed from tail-aware objectives such as Conditional Value at Risk to steer generation toward desired risk profiles. As a result, a single trained model can flexibly produce risk-averse, risk-neutral, or risk-seeking behaviors by changing only the inference-time risk parameter. Extensive experiments on risk-sensitive D4RL and risky robot navigation benchmarks demonstrate that RS-Diffuser achieves state-of-the-art performance, improving both overall return and worst-case robustness while reducing safety violations.
Offline safe reinforcement learning often requires policies to adapt at deployment time to safety budgets that vary across episodes or change within a single episode. While diffusion-based planners enable flexible trajectory generation, existing guidance schemes often treat reward improvement and constraint satisfaction as competing gradient objectives, which can lead to unreliable safety compliance under cost limits. We reinterpret adaptive safe trajectory generation as sampling from a constrained trajectory distribution, where the budget restricts the trajectory region, and reward shapes preferences within that region. This perspective motivates Safe Decoupled Guidance Diffusion (SDGD), which conditions classifier-free guidance on the cost limit to bias sampling toward trajectories satisfying the specified limit, while using reward-gradient guidance to refine trajectories for higher return. Because direct reward guidance can increase return while also steering samples toward trajectories with higher cumulative cost, we introduce Feasible Trajectory Relabeling (FTR) to reshape reward targets and discourage such directions. We further provide a first-order sampling-time analysis showing that FTR suppresses reward-induced cost drift under a prefix-restorative alignment condition. Extensive evaluations on the DSRL benchmark show that SDGD achieves the strongest safety compliance among baselines, satisfying the constraint on 94.7% of tasks (36/38), while obtaining the highest reward among safe methods on 21 tasks.
Rufeng Chen, Zhaofan Zhang, Zhejiang Yang +2
1The Hong Kong University of Science and Technology (Guangzhou) · 2Jilin University
Diffusion models have emerged as powerful tools for planning and control by learning multimodal distributions over actions and trajectories. Yet reliable inference-time safety enforcement remains a key barrier to their deployment in safety-critical tasks. Existing approaches typically project each denoising iterate onto the feasible set, even though constraints are defined only on the final clean trajectory. Enforcing feasibility on noisy intermediate samples can therefore overconstrain the sampling dynamics, substantially degrading sample quality. To address this limitation, we introduce DiRecT (Diffusion-based planning via Receding-horizon denoising with Terminal constraints), a training-free algorithm for constrained sampling from diffusion models via stochastic optimal control (SOC). DiRecT enforces constraints only on the final clean sample, avoiding unnecessary restrictions on the intermediate denoising dynamics. Inspired by model predictive control, we derive a principled receding-horizon surrogate for the otherwise intractable constrained SOC formulation, yielding an efficient algorithm that cleanly separates stochastic denoising from constraint satisfaction, progressively steering samples toward feasible final trajectories without distorting the learned diffusion dynamics. Furthermore, DiRecT is highly flexible: it can leverage off-the-shelf or domain-specific optimizers, incorporate priors over environment dynamics, and optimize additional soft rewards. Extensive experiments on safe planning benchmarks demonstrate that DiRecT substantially improves deployment safety and task performance over existing diffusion-based planning baselines.