Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
Authors: Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He
Organizations: State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · ByteDance Seed
Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic Itô integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe improves over prior diffusion-RL baselines.
Figures & tables
Method
vbase ( ηt )
∇Vt
s(t)
w^1(A,t)
w^2(t)
Sampler
Flow-GRPO [ 24 , 41 ]
vold ( ηt )
s(t)⋅A\upepsilon
2ηttΔt1−t
A⋅4ηt(1+ηt)2t1−t
1+ηt
Flow-SDE
GRPO-Guard [ 37 ]
vold ( ηt )
s(t)⋅A\upepsilon
2ηttΔt1−t
0
s(t)Δt(1+ηt)
Flow-SDE
TempFlow-GRPO [ 14 ]
vold ( ηt )
s(t)⋅A\upepsilon
2ηttΔt1−t
A⋅4(1+ηt)2ηtt2(1−t)Δt
(1+ηt)/s(t)
Branched-SDE
AWM [ 40 ]
E[v∣xt] ( 1 )
−s(t)⋅A(v−vold)
t1−t
A/Δt
s(t)Δt2
Flow-SDE / ODE
DiffusionNFT [ 45 ]
E[v∣xt] ( 1 )
−s(t)⋅A(v−vold)
t1−t
1/Δt
βs(t)Δt2
DPM-ODE [ 26 ]
Ours (Sec. 3.3 )
vold ( ηt )
∇VtKDE (Eq. ( 19 ))
t1−t
(1−t)α1/Δt
s(t)Δttα2
Flow-SDE
Table 1: Diffusion RL methods as rows in the five-knob design space of Eq. ( 18 ). We use the FM parameterisation throughout; v and \upepsilon denote the FM conditional velocity and the per-step Brownian increment. See Appx. D for detailed derivations of each row.
Figure 1 : (a) Comparison between different objectives on Pickscore reward. Original objective Eq. ( 12 ) with ηt=1 (Orange line); one-sample estimator with weighting in Eq. ( 17 ) (Blue line); reweighted version of Eq. ( 17 ) (Green line). (b,c) Empirical variance and squared bias of the value-gradient estimators. We set ηt=1 and evaluate SD3.5-Medium with PickScore reward on four prompts, using 128 trajectories per prompt and 40 denoising steps. At every intermediate state, the reference value gradient is approximated from all 128 trajectories; each estimator is computed with G=24 then compared with this common reference. The KDE training configuration uses h=32.768 .
Figure 2 : On-policy weight (row 1) and off-policy (row 2) weight ablations for the stochastic, deterministic, and our KDE estimator. The results for the same estimator are in the same column.
Figure 3 : (a,b,c). Training time comparison between our method and several baselines. (d). Comparison between Flow-GRPO, DiffusionNFT, and our method on Qwen-Image, HPSv3 reward. * means using the official implementation with different training configs compared to other baselines.
Figure 4 : Ablation studies on the components of our final recipe.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Sampling noise schedule ablations for the stochastic one-sample estimator, deterministic one-sample estimator, and our KDE estimator. For each estimator, we run experiments for sampler noise schedule ηt∈{0.005,0.225} on the Pickscore reward.
Figure 6 : Visualization of different weights
Hyperparameter
Value
Group size G
24 for SD 3.5, 16 for Qwen-Image
Train denoising steps T
10
Inference denoising steps
40 for SD 3.5, 50 for Qwen-Image
Rollout SDE noise schedule
0.225 for stochastic estimator, 0.005 for deterministic and KDE estimator
Optimizer
AdamW, lr=3×10−4 , β1=0.9,β2=0.999
KL coefficient β
10−4
Appendix
Table 2: Hyperparameters used throughout Sec. 4 .
Training time (GPU hours)
18.58
36.44
54.28
72.12
DiffusionNFT
22.85±0.02886
23.06±0.01618
23.16±0.02153
23.22±0.01274
Ours
22.95±0.05240
23.28±0.02239
23.48±0.005547
23.56±0.01461
Appendix
Table 3: PickScore across three random seeds (mean ± standard deviation). Our method achieves a higher mean at every checkpoint with small run-to-run variation.
Method
Training time per epoch (seconds)
DiffusionNFT
115
KDE estimator (ours)
97
Appendix
Table 4 : Measured wall-clock training time per epoch under matched SD3.5-Medium settings.
Figure 7 : Qualitative comparison across GRPO-Guard, DiffusionNFT and our method.
Figure 8 : Qualitative comparison across GRPO-Guard, DiffusionNFT and our method.
Figure 9 : Qualitative comparison across GRPO-Guard, DiffusionNFT and our method.