Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
Figures & tables
Figure 1: MEND stays ahead of ReFL and DiffusionNFT at every evaluated update. Top: SD3.5-M above, MEND below, same prompt and initial noise. Bottom, equal budget: (a) training reward and (b) held-out PickScore against ReFL and DiffusionNFT under the same protocol, with Flow-GRPO (about 4k updates) at its own setting as a level; (c) the HPSv2.1 run’s trained reward.
Method
KL / frozen reference
Reward weights
∇R
Target check
Updates
Flow-GRPO
yes
yes
no
no
∼ 4k
ReFL
no
no
yes
no
100
DiffusionNFT
yes
yes
no
no
1.7k
MEND
no
no
yes
yes
100
Table 1: Objective components. Blue matches MEND ; red differs. The check compares reward gain with displacement cost. MEND uses a lagged behavior adapter, not a frozen reference. Updates refer to the compared runs.
Figure 2: (a) The base flow. (b) Reward reweighting concentrates on the peaks and loses the hard mode. (c) Gradient ascent moves every sample and pushes some of the mass off the data. (d) MEND : capped samples stay, certified moves are short, and the hard mode keeps its mass (Appendix F.3 ).
Figure 3: Rows share a prompt and seed across SD3.5-M, Flow-GRPO, DiffusionNFT, and MEND . PickScore on each tile.
Figure 4: MEND overview. (1) A group for prompt c ; samples at or above the cap κ are kept. (2) Each sample below the cap gets 3 proposals along its reward gradient; the verdict subtracts a distance price from the capped reward and keeps the best candidate, here the shortest move y1 . (3) The trained adapter regresses at the stored state onto the behavior prediction plus the accepted displacement d , a velocity target by equation 5 ; the behavior model then follows it by EMA.
Method
Rewards
Updates
PickScore
HPSv2.1
HPSv3
Image Reward
CLIPScore
Aesthetic
Dist. to base ↓
SD3.5-M
none
0
22.35
0.280
2.72
0.83
0.283
5.39
0
Flow-GRPO
PickScore
∼ 4k
23.52
0.316
7.04
1.27
0.280
5.90
0.313
DiffusionNFT
five
1.7k
23.82
0.331
7.50
1.49
0.292
6.02
0.538
MEND (ours)
PickScore
100
23.70
0.319
7.15
1.32
0.291
5.88
0.313
MEND (ours)
three
300
23.89
0.344
5.87
1.28
0.300
5.92
0.488
Table 2: MEND outperforms Flow-GRPO on five of six evaluators at the same base distance. Its three-reward run surpasses DiffusionNFT on those three rewards. DrawBench 200×5 .
Figure 5: DrawBench prompts (text under each column), SD3.5-M above and MEND below from the same initial noise, PickScore on each tile.
Figure 6: (a) PickScore on 64 held-out Pick-a-Pic prompts at the protocol’s setting, against DiffusionNFT under the same protocol, with Flow-GRPO and DiffusionNFT (about 4k and 1.7k updates) as levels; (b) held-out evaluators and (c) distance to the base on DrawBench at guidance 4.5.
Figure 7: SD3.5-M and MEND after 25, 50, 75 and 100 updates, with the same initial noise along each row and PickScore on each tile. Top: the grasshopper appears by update 50 and the moustache by update 75. Bottom: the wifi symbol appears at update 50 and is sharp by update 75.
Training reward
PickScore
HPSv2.1
ImageReward
CLIPScore
ReFL
DiffusionNFT
none (SD3.5-M)
20.58
0.207
−0.52
0.239
PickScore
24.03
0.301
1.13
0.273
23.92
23.43
ImageReward
22.33
0.303
1.51
0.266
1.28
1.46
HPSv2.1
23.06
0.360
1.21
0.268
0.358
0.336
CLIPScore
22.27
0.268
0.98
0.314
0.308
0.298
Table 3: MEND surpasses ReFL and DiffusionNFT on every training reward. Each MEND run uses SD3.5-M, 100 updates and the protocol of Zhou et al. (2026) ; the base uses the same evaluation setting. The last two columns give ReFL and DiffusionNFT under this protocol on the row’s training reward. Bold: MEND ’s trained metric; underline: second best among the three methods on that metric.
Figure 11
Figure 10: More SD3.5-M and MEND results.
Figure 11: Matched comparisons with Flow-GRPO and DiffusionNFT. Same prompt and seed per row; PickScore on each tile, best in bold.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Guidance trades reward components against diversity. Val64 ( 64×2 , 512 pixels) for the baselines and MEND at update 100; the marked setting is guidance 4.5.
Method
Upd.
Setting
Reg.
PickScore
HPSv2.1
HPSv3
ImageReward
CLIPScore
Aesthetic
Div.
Dist.
Flow-GRPO, PickScore adapter
∼ 4k
4.5/4.5
KL
23.52
0.316
7.04
1.27
0.280
5.90
0.202
0.313
MEND , trained with guidance (Protocol F)
100
4.5/4.5
none
23.56
0.304
5.08
1.07
0.277
5.94
0.214
0.280
Appendix
Table 4: Training with guidance. Protocol F and the Flow-GRPO PickScore adapter, DrawBench 200×5 ; both train and sample at guidance 4.5, with different budgets and regularizers.
Figure 13: Guidance sweep in images. Guidance 1, 2, 3 and 4.5 (columns) for SD3.5-M, Flow-GRPO, DiffusionNFT and MEND at update 100 (rows); two val64 prompts, seed 43, 512 pixels, selected by eye. DiffusionNFT over-saturates as guidance grows.
Figure 17
Table 5: Full SD3.5-M results. DrawBench 200×5 ; setting is training/evaluation guidance, † marks rows from Zhou et al. (2026) (n/r: not reported) and ∗ the default Adam ϵ . Grey sub-rows give 95% prompt-bootstrap intervals; bold and underline mark the best and second best trained headline rows.
We present AdvantageFlow, a forward-process reinforcement learning (RL) algorithm for rectified flow models. The algorithm minimizes an advantage-weighted prediction loss, which maximizes reward, regularized by the rollout policy, which convexifies the objective and makes its optimization stable. Our objective can be viewed as fitting a local reward-improving target distribution. The rollout regularization arises as a variance reduction step. We evaluate AdvantageFlow empirically on text-to-image generation with Stable Diffusion 3.5 Medium and FLUX.1, and compare it to both forward- and reverse-process RL algorithms.
Diffusion and flow-matching models scale because pretraining is supervised regression: a clean sample is noised analytically, and a model regresses against a closed-form target. RL post-training aligns the model with a reward. In image generation, this makes samples compose objects correctly, render text legibly, and match human preferences. Existing methods rely on costly SDE rollouts, reward gradients, or surrogate losses, sacrificing pretraining's regression structure. We show that the structure extends to RL post-training. Under KL-regularized reward maximization, the optimal generative process tilts the clean-endpoint distribution towards samples with higher reward and leaves the noising law unchanged. Combining this with the adjoint-matching optimality condition and a REINFORCE identity, we derive Reinforce Adjoint Matching (RAM): a consistency loss that corrects the pretraining target with the reward. At each step, we draw a clean endpoint from the current model, evaluate its reward, noise it as in pretraining, and regress. No SDE rollouts, backward adjoint sweeps, or reward gradients are required. Like the pretraining objective, RAM is simple and scales. On Stable Diffusion 3.5M, RAM achieves the highest reward on composability, text rendering, and human preference, reaching Flow-GRPO's peak reward in up to 50× fewer training steps.
Andreas Bergmeister, Stefanie Jegelka, Nikolas Nüsken +2
1TU Munich, MCML · 2MIT CSAIL · 3King’s College London +2
Recent work has demonstrated that online reinforcement learning (RL) can substantially improve the quality and alignment of flow matching models for image and video generation. Methods such as Flow-GRPO and CPS cast the denoising process as a Markov Decision Process and apply PPO-style ratio clipping to enforce a trust region. However, we argue that ratio clipping is structurally ill-suited for flow models: the probability ratio between new and old policies is a noisy, single-sample estimate of the true policy divergence, leading to over-constraining in some regions of the trajectory and under-constraining in others. We propose Flow-DPPO (Flow Divergence Proximal Policy Optimization), which replaces ratio clipping with a divergence proximal constraint. A key observation is that the per-step policy in flow models is Gaussian, enabling exact and cheap computation of the KL divergence between old and new policies. Flow-DPPO employs an asymmetric divergence mask that blocks gradient updates only when they simultaneously move away from the trusted region and violate the divergence threshold. Experiments show that Flow-DPPO achieves higher rewards with better KL-proximal efficiency, alleviates catastrophic forgetting, promotes balanced multi-objective optimization, and enables stable multi-epoch training where ratio clipping degrades. Code and models are available at https://github.com/Tencent-Hunyuan/UniRL/tree/main/FlowDPPO.
Bowen Ping, Xiangxin Zhou, Penghui Qi +3
Xi’an Jiaotong University · Tencent Hunyuan · National University of Singapore