Designing Reinforcement Learning for Diffusion Models: A Unified Path-Space View
Authors: Yixian Xu, Yuanrui Zhang, Shengjie Luo, Liwei Wang, Di He
Organizations: State Key Laboratory of General Artificial Intelligence, School of Intelligence Science and Technology, Peking University · ByteDance Seed
Reinforcement learning (RL) post-training provides a direct way to align diffusion models with human preferences and task-specific rewards. However, current RL algorithms for diffusion models remain fragmented: reverse-trajectory methods rely on discretized likelihood ratios, whereas forward-matching methods train on reward-labeled noising versions of the rollout samples. This paper shows that these seemingly different losses arise from a single path-space principle. Starting from the regularized diffusion-RL objective, we use importance sampling between sampling SDEs to obtain an explicit policy-gradient estimator on trajectory space. The estimator contains the stochastic Itô integral underlying Flow-GRPO-type updates; we derive an equivalent variance-reduced value-gradient form that recovers the forward-matching structure of AWM and DiffusionNFT. This identifies the empirical gap between these method families as a variance-reduction effect rather than a difference in RL principle. The derivation yields a unified design space organized by value-gradient estimation, weight functions, and sampling choices. Within this space, we propose a multi-sample KDE value-gradient estimator that reuses rollout groups, together with scale-bounded weight families that retain stable existing recipes while excluding singular ones. Experiments on SD3.5-M and Qwen-Image models validate the variance-reduction explanation and show that the resulting recipe improves over prior diffusion-RL baselines.
Figures & tables
Method
vbase ( ηt )
∇Vt
s(t)
w^1(A,t)
w^2(t)
Sampler
Flow-GRPO [ 24 , 41 ]
vold ( ηt )
s(t)⋅A\upepsilon
2ηttΔt1−t
A⋅4ηt(1+ηt)2t1−t
1+ηt
Flow-SDE
GRPO-Guard [ 37 ]
vold ( ηt )
s(t)⋅A\upepsilon
2ηttΔt1−t
0
s(t)Δt(1+ηt)
Flow-SDE
TempFlow-GRPO [ 14 ]
vold ( ηt )
s(t)⋅A\upepsilon
2ηttΔt1−t
A⋅4(1+ηt)2ηtt2(1−t)Δt
(1+ηt)/s(t)
Branched-SDE
AWM [ 40 ]
E[v∣xt] ( 1 )
−s(t)⋅A(v−vold)
t1−t
A/Δt
s(t)Δt2
Flow-SDE / ODE
DiffusionNFT [ 45 ]
E[v∣xt] ( 1 )
−s(t)⋅A(v−vold)
t1−t
1/Δt
βs(t)Δt2
DPM-ODE [ 26 ]
Ours (Sec. 3.3 )
vold ( ηt )
∇VtKDE (Eq. ( 19 ))
t1−t
(1−t)α1/Δt
s(t)Δttα2
Flow-SDE
Table 1: Diffusion RL methods as rows in the five-knob design space of Eq. ( 18 ). We use the FM parameterisation throughout; v and \upepsilon denote the FM conditional velocity and the per-step Brownian increment. See Appx. D for detailed derivations of each row.
Figure 1 : (a) Comparison between different objectives on Pickscore reward. Original objective Eq. ( 12 ) with ηt=1 (Orange line); one-sample estimator with weighting in Eq. ( 17 ) (Blue line); reweighted version of Eq. ( 17 ) (Green line). (b,c) Empirical variance and squared bias of the value-gradient estimators. We set ηt=1 and evaluate SD3.5-Medium with PickScore reward on four prompts, using 128 trajectories per prompt and 40 denoising steps. At every intermediate state, the reference value gradient is approximated from all 128 trajectories; each estimator is computed with G=24 then compared with this common reference. The KDE training configuration uses h=32.768 .
Figure 2 : On-policy weight (row 1) and off-policy (row 2) weight ablations for the stochastic, deterministic, and our KDE estimator. The results for the same estimator are in the same column.
Figure 3 : (a,b,c). Training time comparison between our method and several baselines. (d). Comparison between Flow-GRPO, DiffusionNFT, and our method on Qwen-Image, HPSv3 reward. * means using the official implementation with different training configs compared to other baselines.
Figure 4 : Ablation studies on the components of our final recipe.
Appendix figures & tables8 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5 : Sampling noise schedule ablations for the stochastic one-sample estimator, deterministic one-sample estimator, and our KDE estimator. For each estimator, we run experiments for sampler noise schedule ηt∈{0.005,0.225} on the Pickscore reward.
Figure 6 : Visualization of different weights
Hyperparameter
Value
Group size G
24 for SD 3.5, 16 for Qwen-Image
Train denoising steps T
10
Inference denoising steps
40 for SD 3.5, 50 for Qwen-Image
Rollout SDE noise schedule
0.225 for stochastic estimator, 0.005 for deterministic and KDE estimator
Optimizer
AdamW, lr=3×10−4 , β1=0.9,β2=0.999
KL coefficient β
10−4
Appendix
Table 2: Hyperparameters used throughout Sec. 4 .
Training time (GPU hours)
18.58
36.44
54.28
72.12
DiffusionNFT
22.85±0.02886
23.06±0.01618
23.16±0.02153
23.22±0.01274
Ours
22.95±0.05240
23.28±0.02239
23.48±0.005547
23.56±0.01461
Appendix
Table 3: PickScore across three random seeds (mean ± standard deviation). Our method achieves a higher mean at every checkpoint with small run-to-run variation.
Method
Training time per epoch (seconds)
DiffusionNFT
115
KDE estimator (ours)
97
Appendix
Table 4 : Measured wall-clock training time per epoch under matched SD3.5-Medium settings.
Figure 7 : Qualitative comparison across GRPO-Guard, DiffusionNFT and our method.
Figure 8 : Qualitative comparison across GRPO-Guard, DiffusionNFT and our method.
Figure 9 : Qualitative comparison across GRPO-Guard, DiffusionNFT and our method.
Reinforcement learning has been widely applied to diffusion and flow models for visual tasks such as text-to-image generation. However, these tasks remain challenging because diffusion models have intractable likelihoods, which creates a barrier for directly applying popular policy-gradient type methods. Existing approaches primarily focus on crafting new objectives built on already heavily engineered LLM objectives, using ad hoc estimators for likelihood, without a thorough investigation into how such estimation affects overall algorithmic performance. In this work, we provide a systematic analysis of the RL design space by disentangling three factors: i) policy-gradient objectives, ii) likelihood estimators, and iii) rollout sampling schemes. We show that adopting an evidence lower bound (ELBO) based model likelihood estimator, computed only from the final generated sample, is the dominant factor enabling effective, efficient, and stable RL optimization, outweighing the impact of the specific policy-gradient loss functional. We validate our findings across multiple reward benchmarks using SD 3.5 Medium, and observe consistent trends across all tasks. Our method improves the GenEval score from 0.24 to 0.95 in 90 GPU hours, which is 4.6 times more efficient than FlowGRPO and 2× more efficient than the SOTA method without reward hacking.
Jaemoo Choi, Yuchen Zhu, Wei Guo +6
Georgia Institute of Technology · National University of Singapore · Nanjing university.
Reinforcement Learning (RL) has emerged as a central paradigm for advancing Large Language Models (LLMs), where both pre-training and RL post-training stages are grounded in the same log-likelihood formulation. In contrast, recent RL approaches for diffusion models, most notably Denoising Diffusion Policy Optimization (DDPO), optimize an objective different from the pretraining objectives--score/flow matching loss. In this work, we establish a novel theoretical analysis: DDPO is an implicit form of score/flow matching with noisy targets, which increases variance and slows convergence. Building on this analysis, we introduce Advantage Weighted Matching (AWM), a policy-gradient method for diffusion. It uses the score/flow-matching loss and reweights each sample by its advantage. In effect, AWM raises the influence of high-reward samples and suppresses low-reward ones while keeping the modeling objective identical to pretraining. This simple yet effective design yields substantial benefits: on the GenEval, OCR, and PickScore benchmarks, AWM delivers up to a 34× speedup over Flow-GRPO (which builds on DDPO), when applied to Stable Diffusion 3.5 Medium and FLUX, without compromising generation quality. Code is available at https://github.com/scxue/advantage_weighted_matching
Reinforcement learning from human feedback (RLHF) has emerged as a powerful paradigm for aligning generative models with human preferences. However, applying RLHF to diffusion models remains highly feedback inefficient, as existing approaches typically require large amounts of human or reward model evaluations. This limitation reduces the practicality of diffusion RLHF in realworld settings where feedback is the primary bottleneck. In this paper, we propose two complementary strategies that substantially improve the feedback efficiency of diffusion RLHF while preserving generalization to unseen prompts. Our key observation is that reward information in diffusion trajectories is unevenly distributed: not all denoising timesteps or trajectories contribute equally to learning from a reward signal. By emphasizing informative timesteps and trajectories during optimization, we obtain more effective gradient updates. First, we introduce a per-timestep weighting scheme that reweights denoising steps during policy optimization. We theoretically connect this weighting to the optimal convergence properties of proximal policy optimization (PPO) and approximate the resulting weighting trend empirically. Second, we introduce a replay mechanism that prioritizes informative trajectories, enabling the model to reuse past samples instead of repeatedly querying new rewards. Together, these strategies significantly improve the feedback efficiency of diffusion RLHF. Under identical hyperparameter settings, our approach achieves up to a 6× improvement in sample efficiency compared to widely used diffusion RLHF baselines.
Eric Zhu, Abhinav Shrivastava, Soumik Mukhopadhyay
Carnegie Mellon University · University of Maryland, College Park