A nascent family of methods that forgoes the policy gradient and reweights a supervised regression instead has garnered momentum in reinforcement learning for diffusion and flow models. DiffusionNFT, FlowAWR, and RAM are representative regimes with contrasting motivations. It is yet opaque what, if anything, they share. We substantiate that each is the solution of one divergence-constrained reward-maximization problem, and they are differentiated only by the convex generator that defines the constraint. Under the unified modeling framework, we unravel the relaxations that prior art made during building the advantage-embedded regression target: approximating the KKT condition and posterior normalizer for the linear and exponential tilt shapes DiffusionNFT and FlowAWR respectively, while preserving the exact sparsemax projection onto the probability simplex for linear tilt leads to another superior model type in this work. Beyond the theoretical underpinnings, we further empirically investigate the design space and shed light on the training recipe for regression-style diffusion RL. Retaining the merits discovered during our exploration gives rise to DiffusionRFT, our paradigm that converges faster, trains more stably, and attains the top performance.
Figures & tables
Figure 1: Tilt × baseline grid. Three tilts (exponential, dense linear, sparsemax) crossed with two baselines (canonical constant b , variance-minimizing control variate b∗ ). The sparsemax tilt reduces to the dense linear tilt when the truncation is inactive, i.e., miniA>−1 . Yet, this criterion fails on 86.8% of the dense run’s steps. The optimal control variate b∗ stabilizes and strengthens every tilt.
Figure 2: Loss geometry: the v -space update is misdirected, not too small. More advantage signal (b) and largest movement budget (c), but an order-of-magnitude larger residual (d), and correspondingly less reward (a). The dashed line marks the matched horizon among these runs.
Figure 3: Anchor-lock is a displacement failure, not a signal failure. (a) spans the full budget, where both the official and reimplemented RAM escape the lock with the aid of a 100× advantage scale but then collapse. (b)–(d) are restricted to the matched window in which the frozen-anchor runs coexist: dashed, non-zero step updates; solid, no effective accumulation (b) and no prompt group is ever solved (c), even though the advantage is larger than under a rolling anchor (d). The three frozen-anchor curves overlap in every panel, which pinpoints the anchor as the root cause.
Figure 4: Decoupling the renoising timestep makes the batch heavy-tailed, and stratification makes it worse. (a) Training reward. (b) Within-batch spread of the policy-EMA gap ∥fθ−fold∥2 on a logarithmic axis: lowest for the trajectory-coupled schedule, ∼2× once the timestep is decoupled, and ∼8× once the draws are stratified. (c) Advantage magnitude: the trajectory-coupled run declines as it solves its prompt groups, while the decoupled runs remain high. (d) Fraction of prompt groups whose rewards have zero variance, which carry no advantage. A uniform distribution spends its evaluations on uninformative states the sampler hardly visits instead of the ones it resolves.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 5: How far the implemented target is from the exact one. By equation 45 the rescaling rt of equation 14 has mean one at every noise level and variance bounded by E[A2] , a property of the reward group alone. (a) That bound, measured on four runs and plotted as a bound on the standard deviation, so every quantity is shown as a square root. No reward model is presumed. For any centered A , (E∣A∣)2≤E[A2]≤maxi∣Ai∣E∣A∣ . The band is that bracket and the line its upper end. (b) The decay of sd(rt) with t , computed in a closed form for a Gaussian-mixture endpoint law, which is exact because both exponential tilting and Gaussian smoothing preserve the family. Dots mark the first nine timesteps a T=10 trajectory-coupled buffer supplies under the shift=3 schedule. When averaged over them, the dispersion is 35% of its worst case. It is a surrogate law and not the model, so this panel is illustrative whereas (a) and (c) are measured. (c) The same bound against the reward it was measured at. As the policy converges the group becomes homogeneous, the advantages shrink and the approximation becomes exact in the limit that matters.
Figure 6: Generalization to Pickscore reward. DiffusionRFT achieves consistent superiority over DiffusionNFT, as displayed by training reward (a) and validation reward (b).
Figure 7: The trajectory schedule covers only part of the noise axis. t is the noise level, t=0 being clean data. (a) Pushing the uniform grid through the SD3.5 shift κ=3 gives a density κ2=9× heavier at t=1 than at t=0 , and the nine loss timesteps inherit it. (b) Those nine nodes never fall below t=0.43 , so the shaded band is off-support for the loss : drawing t∼U[0,1] sends 43% of the loss evaluations there, and stratification guarantees it. (c) The same statement as a CDF. This is the mechanism behind figure 4 : decoupling the renoising timestep does not merely re-weight the noise axis, it regresses at noise levels the policy is never asked to act on.