Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoffs through careful theoretical considerations and method design. We analyze the theoretical framework and mathematically demonstrate that \emph{only-latter timestep} updates of diffusion model may be harmful for diversity contrary to the conclusions presented in a previous work. Additionally, we propose an incremental Feynman-Kac training based on strong theoretical foundations in order to achieve the best-yet alignment-diversity tradeoffs. We perform extensive experiments and compare our method against related diffusion policy optimization approaches in three different tasks and also provide strong ablations for each component, thus validating strong performance gains in both alignment and diversity.
Figures & tables
Figure 1 : Performance in alignment and diversity. (a) Our method, iADD, achieves a higher CLIP Score at any given diversity level compared to other methods. Notably, as reward increases, competing methods (DDPO and B2 -DiffuRL) produce increasingly cartoonish, over-saturated outputs, whereas our method preserves photorealistic quality throughout training. The visualized images are for the prompt “a dog washing dishes”. (b) Alignment performance on common and rare prompts, evaluated at constant Inception Score(IS) for each method. Rare prompts represent conditions that inherently yield low-reward outcomes under the pretrained model.
Figure 2 : Versatility across diverse generation tasks.
Components
Alignment
Diversity
Rarity
Training Eff.
Timesteps
Early timesteps
↑
↑ , Prop. 1
↑
↑
Timestep sampling
↑
↑ , Prop. 1 , 2
↑
↑
Incremental timesteps
↑
↑ , Remark. 2
↑
↑
.
FK
FK Guidance
↑ , Sec. 4.3
↑
↑↑
↓
FK Branching
↑
↑↑ , Sec. 4.3
↑
↓
Table 1 : Disassembly of our Method’s components and their roles. Up arrow ( ↑ ) denotes improvement and down arrow ( ↓ ) denotes worsening. Additional reference to sections or theoretical justifications are mentioned in the most relevant column.
Method
Training
Inference
Formulation
Incr.
FK
Branch.
No Test-time Scaling
DDPM/DDIM
DPOK [ 14 ]
✗
✗
✗
✓
DDPM
Score-as-Action [ 74 ]
✗
✓
✗
✓
DDPM
DDPO [ 4 ]
✗
✗
✗
✓
DDIM
B 2 -DiffuRL [ 28 ]
✗
✗
✓
✓
DDIM
FK Steering & FK Correctors [ 59 , 60 ]
✗
✗
✗
✗
DDIM
Table 2 : Summary of methods and techniques across training, inference, and diffusion formulation. Our method is designed for the DDIM formulation.
Figure 3 : Overview of FK Sampling and Incremental Timestep Training. (A) FK Sampling: Our approach employs a Feynman-Kac (FK) particle filter for sampling. Starting from Gaussian noise at t=T , trajectories ( p1,…,p4 ) are evaluated at discrete checkpoints tn based on their potentials g . The FKBranching mechanism duplicates high-potential particles while pruning low-potential ones, effectively redirecting the stochastic denoising paths toward regions of higher expected reward. (B) Incremental timestep updates: During training, we utilize a progressive stage-wise strategy where the cardinality of the updated timesteps S increases.
Figure 4 : Qualitative comparison on common and rare prompts. Our method better preserves prompt alignment while maintaining realistic visual quality, with larger gains on rare prompts.
Figure 5 : Qualitative results. (a) 3D scene synthesis. Top: collision reward. Bottom: spatial reward (“TV stand facing bed”). DDPO and B 2 -DiffuRL satisfy constraints but introduce collisions; our method avoids them. (b) Vanishing-point correction. Red: off-vanishing intersections; green: vanishing point. DDPO and B 2 -DiffuRL shows weaker convergence and desaturation; our method preserves quality while enforcing perspective. (Better viewed zoomed.)
Method
Image Generation
3D Indoor Scene
Vanishing Point Correction
Reward ↑
IS ↑
Rarity ↑
AUC ↑
Collision ↓
CKL ↓
Reward ↑
IS ↑
Rarity ↑
Pretrained (SD)
0.3272
1.3514
-
-
53.87
30.48%
0.6105
1.2735
-
DDPO
0.3417
1.2786
24.81%
53.8%
63.39%
7.726%
0.6815
1.2600
21.33%
B 2 -DiffuRL
0.3452
1.2728
24.22%
58.7%
64.06%
6.278%
0.6486
1.2550
31.11%
Ours
0.3553
1.3241
66.85%
77.7%
59.20%
6.13%
0.7763
1.2881
95.11%
Table 3 : Performance comparison across metrics and modalities. Evaluated at their best operating points, our method (iADD) consistently achieves the best alignment-diversity trade-off across all three tasks compared to DDPO and B2 -DiffuRL. Notably, iADD maintains higher diversity (better IS and lower collision rates) at comparable or higher reward levels, while also demonstrating state-of-the-art robustness on rare prompts (Rarity). Arrows indicate the desired direction.
Method
CLIP ↑
IS ↑
Rarity ↑
AUC ↑
LPIPS ↑
AIG (inference)
0.3526
1.3147
17.50%
–
0.7531
iADD
0.3830
1.3035
78.33%
83.7%
0.7693
DanceGRPO
0.3886
1.2750
61.67%
39.3%
0.7624
BranchGRPO
0.3854
1.2917
56.67%
43.73%
0.7447
iADD+GRPO
0.4126
1.2942
88.33%
73.3%
0.7778
Table 4 : Comparison with recent guidance and GRPO methods on Template 1 (animal activities). Results use the same prompt set, with trainable methods evaluated at matched checkpoints. iADD provides the strongest standalone diversity and reward–diversity AUC, while iADD+GRPO obtains the highest alignment, rarity, and intra-prompt LPIPS.
Figure 6 : Empirical validation of the timestep analysis. Left: cumulative ∥δx0∥ advantage of uniform over late-only updates across training stages (Prop. 1 ). Right: full-Jacobian log-determinant for dense and sparse updates; the less-negative value indicates better volume preservation under sparse optimization (Prop. 2 ).
Figure 7 : Additional quantitative results. Our method achieves fewer collisions at matched spatial reward in 3D scene synthesis, improves diversity in the ablation study, and is preferred by human raters over DDPO and B 2 -DiffuRL.
Figure S8 : Pseudo code of iADD
Method
Inception Score
Ours ( λFK=10 )
1.2830
Ours ( λFK=2 )
1.3035
Table S5 : Effect of λFK on Inception Score.
Template
Representative Prompts
T1: Animal-Action
1. a cat washing dishes 2. a cat playing chess 3. a dog riding a bike 4. a dog playing chess 5. a horse riding a bike 6. a horse playing chess 7. a dolphin riding a bike 8. a gorilla playing chess 9. a fox riding a bike 10. a chicken playing chess
T2: Rare Binding
1. red apple 2. yellow banana 3. orange orange 4. purple grape 5. red watermelon 6. yellow apple 7. brown banana 8. green strawberry 9. black grape 10. brown kiwi
T3: Spatial Relations
1. vase on table 2. shirt on person 3. watch on person 4. jacket on person 5. motorcycle on road 6. motorcycle behind person 7. person behind person 8. building behind trees 9. hydrant behind motorcycle 10. trees behind grass
T4: Geometric
1. a straight railroad track leading to a vanishing point at dawn 2. a long marble hallway in a palace with ornate ceiling 3. a straight canal through tulip fields in the Netherlands 4. a bamboo forest path lined with tall bamboo stalks 5. an infinity pool edge overlooking a mountain valley 6. railway tracks stretching to the horizon at sunset 7. an empty straight road disappearing into the distance 8. a long tunnel with circular arches receding into darkness 9. a straight highway through a flat desert landscape 10. a tree-lined avenue with branches forming an arch overhead
Table S6 : Representative prompts from each template category for text-to-image evaluation.
Figure S9 : Pretrained Reward Distribution across Templates. We visualize the CLIP reward (T1-T3) and Vanishing Point reward (T4) of the pretrained Stable Diffusion model across sorted evaluation prompts. For each template, we define Rare prompts as those in the bottom 25% (red) and Common prompts as those in the top 25% (green). The characteristic curve highlights significant performance variance in the base model, particularly for challenging prompts in T1-T3.
Figure S10 : Quartile Performance Analysis. Comparison of model performance across prompt difficulty quartiles (Q1: Common, Q2: Intermediate, Q3: Rare). The horizontal line represents the global pretrained SD baseline. Our method ( Ours ) consistently outperforms SD, DDPO, and B 2 -DiffuRL, with the largest relative gain observed in the most challenging Rare quartile ( +16.7% ), while maintaining significant leads in intermediate ( +12.0% ) and common ( +10.5% ) scenarios.
Method
Reward ↑
Diversity ↑
Rarity ↑
AUC ↑
SD [ 53 ]
0.3565
1.3067
-
-
DPOK [ 14 ]
0.3582
1.2855
28.33
-
B 2 Diff-RL [ 28 ] + KL regularization
0.3646
1.2784
27.5
-
Ours + KL regularization
0.3800
1.3002
55.83
-
Table S7 : Comparison of different methods under KL-divergence regularization. Diversity is measured while fixing the reward to 0.38 , and reward is measured while fixing diversity to 1.285 . While KL regularization helps reduce reward hacking as suggested by DPOK, our method still achieves higher diversity and reward than existing approaches.
Method
Alignment ↑
Diversity ↑
Rarity ↑
AUC ↑
SD
-
-
-
DDPO
-
-
-
B 2 -DiffuRL
-
-
-
Ours
-
-
-
Table S8 : Quantitative results for Template T1.
Method
Alignment ↑
Diversity ↑
Rarity ↑
AUC ↑
SD
-
-
-
DDPO
-
-
-
B 2 -DiffuRL
-
-
-
Ours
-
-
-
Table S9 : Quantitative results for Template T2.
Method
Alignment ↑
Diversity ↑
Rarity ↑
AUC ↑
SD
-
-
-
DDPO
-
-
-
B 2 -DiffuRL
-
-
-
Ours
-
-
-
Table S10 : Quantitative results for Template T3.
Figure S11 : Inception Score vs Mean Reward plot for the vanishing point task. Our method achieves higher alignment while the inception score also increases, demonstrating an improvement in both metrics without a strict tradeoff.
Figure S12 : User study results for the vanishing point task, showing global preference, per-template preference, and per-prompt preference. Human evaluators consistently preferred our method.
Figure S13 : User study results for general image synthesis broken down per template and per prompt.
Figure S14 : Quartile reward plots for Template T1 and T2 scenarios.
Figure S15 : Quartile reward plots for Template T3 and the vanishing point geometric task.
Method
GPU Hours (Training Time)
DDPO
2.55
DDPO + incremental
2.46
B 2 -DiffuRL
3.42
B 2 -DiffuRL + incremental
2.76
FK + incremental (ours)
9.55
Branch-FK + incremental (iADD)
7.43
Table S11 : Time taken to reach the same mean reward point (0.388 CLIP score).
RL post-training has become increasingly pivotal for improving diffusion policies, but existing diffusion policy-gradient methods are often unstable and cannot achieve reliable policy improvement. We identify the cause as the double-drift phenomenon: optimizing a variational surrogate can let the ELBO separate from the true log-likelihood, which then makes the resulting proxy policy gradient misaligned with the true policy gradient of expected return. We propose \textbf{DiPOD}, a diffusion policy optimization framework that maintains tight-bound behavior throughout training by interleaving self-distillation with policy-improving gradient updates. This leads to a simple and practical algorithm: augmenting each diffusion policy-gradient update with an on-policy ELBO regularizer. Across diffusion language model post-training and continuous-control diffusion policies, DiPOD substantially stabilizes training and reaches higher rewards than previous methods.
Reinforcement learning (RL) has shown extraordinary potential in aligning diffusion models to downstream tasks, yet most of them still suffer from significant reward hacking, which degrades generative diversity and quality by inducing visual mode collapse and amplifying unreliable rewards. We identify the root cause as the mode-seeking nature of these methods, which maximize expected reward without effectively constraining probability distribution over acceptable trajectories, causing concentration on a few high-reward paths. In contrast, we propose Trajectory Matching Policy Optimization (TMPO), which replaces scalar reward maximization with trajectory-level reward distribution matching. Specifically, TMPO introduces a Softmax Trajectory Balance (Softmax-TB) objective to match the policy probabilities of K trajectories to a reward-induced Boltzmann distribution. We prove that this objective inherits the mode-covering property of forward KL divergence, preserving coverage over all acceptable trajectories while optimizing reward. To further reduce multi-trajectory training time on large-scale flow-matching models, TMPO incorporates Dynamic Stochastic Tree Sampling, where trajectories share denoising prefixes and branch at dynamically scheduled steps, reducing redundant computation while improving training effectiveness. Extensive results across diverse alignment tasks such as human preference, compositional generation and text rendering show that TMPO improves generative diversity over state-of-the-art methods by 9.1%, and achieves competitive performance in all downstream and efficiency metrics, attaining the optimal trade-off between reward and diversity.
Jiaming Li, Chenyu Zhu, Nanxi Yi +9
1MAIR Lab, Huazhong University of Science and Technology · 2Kuaishou Technology · 3Nanyang Technological University +2
Online reinforcement learning is becoming increasingly important for aligning diffusion models with non-differentiable objectives. However, existing methods still face limitations in assigning fine-grained credit along denoising trajectories and in realizing stable value-based optimization. We propose a state-aligned latent actor-critic framework for diffusion post-training, in which the diffusion model serves as its own timestep-conditioned value function and predicts values directly on noisy latent states. This enables trajectory-level PPO training, supports stable actor-critic optimization with simple conditioning and value pretraining strategies, and naturally allows the learned critic to be reused for inference-time steering. We further extend the framework to multi-reward optimization, where joint training with complementary rewards helps alleviate reward hacking. Across both UNet- and DiT-based backbones, our method consistently outperforms prior group-relative RL and actor-critic baselines on single-reward and multi-reward benchmarks, while test-time steering provides additional gains in generation quality.
Zhengyang Liang, Qihang Zhang, Ceyuan Yang
University of Toronto · Vector Institute · The Chinese University of Hong Kong