Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.
Figures & tables
Figure 1: (A) Evaluation paths obtained from fully trained policies on the same fixed PointGoal1 layout. See Figure 6 for clearance distribution of these paths. (B) Illustrative diagram for the RADPO algorithm.
Figure 2: Training performance across ten continuous-control safety tasks. Episodic return (top rows) and episodic cost (bottom rows) for RADPO and baseline methods. Curves show empirical means across six random seeds; shaded regions denote one standard deviation. Dashed horizontal lines indicate the safety budget d=25 .
Figure 3: Late-training return versus safety cost. Markers indicate mean performance across seeds, computed over the final ten logged training points per run. Error bars represent one standard deviation across seed averages. Optimal performance lies toward the top-left quadrant (high return, low cost). Vertical dashed lines denote the budget d=25 and the shaded region is infeasible.
Figure 4: Ablations on safety-signal formulation and penalty placement. Training return (top) and episodic cost (bottom). Curves depict means across evaluated seeds with one standard deviation shaded. Solid lines show smoothing over raw trajectories (faint curves). Inset axes provide magnified views around the cost budget d=25 .
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: Ablations on hyperparameter values. Ablation results for reachability discount γc , shaping coefficient β and candidates per state K . Curves depict means across evaluated seeds with one standard deviation shaded.
Time per environment transition (ms)
Method
Sample acquisition
Critic optimization
Diffusion-policy optimization
Other learning
Overall
ALGD
3.298
16.057
78.892
41.524
139.771
Lagrangian-style
0.418
0.545
0.459
3.387
4.809
RADPO
0.616
0.423
0.406
2.697
4.142
Appendix
Table 1: Computational cost on PointGoal1. All entries are ms per environment step.
Figure 6: Distribution of states for trajectories in Figure 1 .A.
Family
Task
Horizon T
Env. steps
Eval. every
Velocity
SafetyAntVelocity-v1
1000
1M
25k
SafetyHalfCheetahVelocity-v1
1000
1M
25k
SafetyHopperVelocity-v1
1000
1M
25k
SafetyHumanoidVelocity-v1
1000
3M
25k
SafetySwimmerVelocity-v1
1000
1M
25k
SafetyWalker2dVelocity-v1
1000
1M
25k
Appendix
Table 2: Tasks, interaction budgets and evaluation cadence.
Group
Hyperparameter
Value
Networks
Critic / reachability-critic MLP
3×256 , Mish
Diffusion actor MLP
3×256 , Mish
Denoising steps
20
Noise schedule
cosine
Reward critics (twin)
2
Optimisation
Optimiser
Adam
Appendix
Table 3: RADPO hyperparameters held constant across all ten tasks.
Task
UTD
α0
σ0
β
ηλ
Ant
1.0
0.2
0.5
1.5
0.010
HalfCheetah
1.0
0.2
0.5
1.5
0.010
Hopper
0.6
0.2
0.5
0.8
0.002
Humanoid
1.0
0.2
0.1
2.4
0.120
Swimmer
0.2
0.2
0.5
0.025
0.100
Walker2d
0.2
0.2
0.5
1.5
0.010
Appendix
Table 4: Per-task RADPO hyperparameters. UTD is the number of optimiser updates per collected transition, α0 the initial exploration temperature ( σexp=σ0α0 at initialisation), σ0 the exploration scale, β the shaping coefficient and ηλ the controller step size. Everything else is as in Table 3 .
Hyperparameter
Velocity / Circle
Goal
Critic / actor MLP
3×256 , GELU
2×256 , ReLU
Actor learning rate
3×10−4
5×10−6
Critic learning rate
3×10−4
1×10−3
Entropy temperature α
learned, α0=0.2
fixed 10−5
α learning rate
3×10−4
—
Policy delay
1
2
Appendix
Table 5: SAC-Lagrangian. “Velocity” covers the six velocity tasks and the two Circle tasks; “Goal” covers CarGoal1 and PointGoal1.
Task
k
c
Multiplier LR
Internal d
Extra
Ant
0.0
0.01
1×10−5
25
—
HalfCheetah
0.0
0.3
1×10−4
25
Qc≥0
Hopper
0.5
0.2
1×10−5
25
done-mask, Qc≥0
Humanoid
0.0
0.01
1×10−5
25
done-mask
Swimmer
0.0
0.01
1×10−5
25
—
Walker2d
0.0
0.2
1×10−4
25
done-mask
Appendix
Table 6: CAL per-task settings. k is the cost-critic UCB coefficient, c the rectification coefficient. “done-mask” indicates that termination is masked in the cost backup; “ Qc≥0 ” that the cost critic is floored at zero.
Hyperparameter
Velocity
Humanoid
Navigation
Policy / value MLP
2×64 , tanh
2×64 , tanh
2×64 , tanh
Learning rate
3×10−4
3×10−4
3×10−4
Steps per epoch
2048
6000
20000
Epochs
500
500
150
Total environment steps
1.024M
3M
3M
Update iterations / epoch
40
40
40
Appendix
Table 7: PPO-Lagrangian.
Task
Budget
Rollout
Actor LR
σ0
Target KL
Ant
1M
10,000
6×10−4
0.2
0.03
HalfCheetah
1M
10,000
6×10−4
0.2
0.03
Hopper
1M
10,000
6×10−4
0.2
0.03
Humanoid
3M
5,000
1×10−4
0.2
0.03
Swimmer
1M
10,000
6×10−4
0.2
0.03
Walker2d
1M
10,000
3×10−4
0.2
0.03
Appendix
Table 8: RESPO per-task settings. Budgets and rollout sizes count total environment transitions across four workers. Actor LR and σ0 are the initial learning rate and policy standard deviation. Observation normalization is used only for Humanoid.
Released code
This paper
Environments
Safety-Gym ( Safexp-* ), MuJoCo
Safety Gymnasium (all ten tasks)
Navigation layout
fixed for every episode and seed
resampled every episode
Episode horizon
400 (navigation)
1000
Goal reached
episode terminates
new goal spawned, episode continues
Cost budget
10 (navigation) / raw velocity
d=25 on all tasks
Evaluation
1 episode
20 deterministic episodes
Appendix
Table 9: ALGD: released training protocol versus the protocol used here.
Group
Hyperparameter
Value
Networks
Reward critics (twin)
2×256 , ReLU
Cost-critic ensemble
4 members, 2×256 , SiLU
Score / noise network
3×128 , ReLU, time embedding
Denoising steps K
5
Diffusion
DDPM noise schedule
linear β∈[10−4,0.02]
VE-SDE noise scale
geometric σ∈[0.01,1.0]
Appendix
Table 10: ALGD hyperparameters (released defaults, shared by the DDPM and VE-SDE agents).
Task
Agent
UTD
Ant
DDPM
10
HalfCheetah
DDPM
10
Hopper
DDPM
10
Humanoid
DDPM
10
Swimmer
DDPM
10
Walker2d
DDPM
10
Appendix
Table 11: ALGD per-task settings. UTD is gradient updates per environment step.
Task
Arm
Cost critic
Shaping target
β
CarGoal1
RADPO
reachability
Vreach(s′)
0.05
discounted critic
discounted
Vc(s′)
0.01
cost-only shaping
-
ct
0.05
lag-style shaping
discounted
none (cost in weights)
—
PointCircle1
RADPO
reachability
Vreach(s′)
0.05
discounted critic
discounted
Vc(s′)
0.01
Appendix
Table 12: Ablation arms. Every arm shares its task’s ηλ , σ0 , α0 , UTD and interaction budget with the RADPO arm beside it, and all settings of Table 3 .
Figure 7: Deterministic-evaluation return and cost across the ten tasks. Evaluation return (upper row of each block) and evaluation cost (lower row), measured over 20 deterministic episodes per checkpoint; the training-rollout counterpart is Figure 2 . Curves are means over the runs of an arm with one standard deviation shaded, and dashed horizontal lines mark the budget d=25 .