cs.LGSep 27, 2026

Elucidating the Design Space of Regression-based Diffusion Reinforcement Learning

Authors: Toyota Li, David Zhao, Alan Zhao

Organizations: Tencent

Abstract

A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.

Figures & tables

Appendix figures & tables3 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Rethinking the Design Space of Reinforcement Learning for Diffusion Models: On the Importance of Likelihood Estimation Beyond Loss Design

    Feb 4, 2026Jaemoo Choi, Yuchen Zhu, Wei Guo +6Diffusion-Based Reinforcement Learning MethodsReward Gradients

  2. FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

    Jun 29, 2026Zheming Fu, Ruizhe He, Wei Shang +4Generative Flow NetworksVelocity Field

  3. Advantage Weighted Matching: Aligning RL with Pretraining in Diffusion Models

    Sep 29, 2025Shuchen Xue, Chongjian Ge, Shilong Zhang +2Diffusion ModelsDiffusion Policies