Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or backpropagates the reward and moves every sample without checking that the move is worth its size. We introduce MEND, a reinforcement learning method built on proximal velocity matching. MEND caps rewards within each prompt group, so samples that already score well receive no move. Below the cap, it proposes moves along the reward gradient and accepts one only when its capped reward gain exceeds a quadratic displacement price. The model then regresses onto the resulting velocity targets, with no KL term, frozen reference model, or advantage weights. In 100 updates, MEND outperforms Flow-GRPO (about 4k updates) on five of six evaluators at the same distance to base-model images. Under an equal-budget protocol, it surpasses ReFL and DiffusionNFT at every evaluated update across four training rewards, reaching PickScore 24.03 versus 23.92 and 23.43, respectively. A 300-update three-reward run also surpasses the five-reward DiffusionNFT model on all three rewards it trains on. MEND is general and easy to adopt: it applies to any flow backbone with a differentiable reward.
Figures & tables
Figure 1: MEND stays ahead of ReFL and DiffusionNFT at every evaluated update. Top: SD3.5-M above, MEND below, same prompt and initial noise. Bottom, equal budget: (a) training reward and (b) held-out PickScore against ReFL and DiffusionNFT under the same protocol, with Flow-GRPO (about 4k updates) at its own setting as a level; (c) the HPSv2.1 run’s trained reward.
Method
KL / frozen reference
Reward weights
∇R
Target check
Updates
Flow-GRPO
yes
yes
no
no
∼ 4k
ReFL
no
no
yes
no
100
DiffusionNFT
yes
yes
no
no
1.7k
MEND
no
no
yes
yes
100
Table 1: Objective components. Blue matches MEND ; red differs. The check compares reward gain with displacement cost. MEND uses a lagged behavior adapter, not a frozen reference. Updates refer to the compared runs.
Figure 2: (a) The base flow. (b) Reward reweighting concentrates on the peaks and loses the hard mode. (c) Gradient ascent moves every sample and pushes some of the mass off the data. (d) MEND : capped samples stay, certified moves are short, and the hard mode keeps its mass (Appendix F.3 ).
Figure 3: Rows share a prompt and seed across SD3.5-M, Flow-GRPO, DiffusionNFT, and MEND . PickScore on each tile.
Figure 4: MEND overview. (1) A group for prompt c ; samples at or above the cap κ are kept. (2) Each sample below the cap gets 3 proposals along its reward gradient; the verdict subtracts a distance price from the capped reward and keeps the best candidate, here the shortest move y1 . (3) The trained adapter regresses at the stored state onto the behavior prediction plus the accepted displacement d , a velocity target by equation 5 ; the behavior model then follows it by EMA.
Method
Rewards
Updates
PickScore
HPSv2.1
HPSv3
Image Reward
CLIPScore
Aesthetic
Dist. to base ↓
SD3.5-M
none
0
22.35
0.280
2.72
0.83
0.283
5.39
0
Flow-GRPO
PickScore
∼ 4k
23.52
0.316
7.04
1.27
0.280
5.90
0.313
DiffusionNFT
five
1.7k
23.82
0.331
7.50
1.49
0.292
6.02
0.538
MEND (ours)
PickScore
100
23.70
0.319
7.15
1.32
0.291
5.88
0.313
MEND (ours)
three
300
23.89
0.344
5.87
1.28
0.300
5.92
0.488
Table 2: MEND outperforms Flow-GRPO on five of six evaluators at the same base distance. Its three-reward run surpasses DiffusionNFT on those three rewards. DrawBench 200×5 .
Figure 5: DrawBench prompts (text under each column), SD3.5-M above and MEND below from the same initial noise, PickScore on each tile.
Figure 6: (a) PickScore on 64 held-out Pick-a-Pic prompts at the protocol’s setting, against DiffusionNFT under the same protocol, with Flow-GRPO and DiffusionNFT (about 4k and 1.7k updates) as levels; (b) held-out evaluators and (c) distance to the base on DrawBench at guidance 4.5.
Figure 7: SD3.5-M and MEND after 25, 50, 75 and 100 updates, with the same initial noise along each row and PickScore on each tile. Top: the grasshopper appears by update 50 and the moustache by update 75. Bottom: the wifi symbol appears at update 50 and is sharp by update 75.
Training reward
PickScore
HPSv2.1
ImageReward
CLIPScore
ReFL
DiffusionNFT
none (SD3.5-M)
20.58
0.207
−0.52
0.239
PickScore
24.03
0.301
1.13
0.273
23.92
23.43
ImageReward
22.33
0.303
1.51
0.266
1.28
1.46
HPSv2.1
23.06
0.360
1.21
0.268
0.358
0.336
CLIPScore
22.27
0.268
0.98
0.314
0.308
0.298
Table 3: MEND surpasses ReFL and DiffusionNFT on every training reward. Each MEND run uses SD3.5-M, 100 updates and the protocol of Zhou et al. (2026) ; the base uses the same evaluation setting. The last two columns give ReFL and DiffusionNFT under this protocol on the row’s training reward. Bold: MEND ’s trained metric; underline: second best among the three methods on that metric.
Figure 11
Figure 10: More SD3.5-M and MEND results.
Figure 11: Matched comparisons with Flow-GRPO and DiffusionNFT. Same prompt and seed per row; PickScore on each tile, best in bold.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: Guidance trades reward components against diversity. Val64 ( 64×2 , 512 pixels) for the baselines and MEND at update 100; the marked setting is guidance 4.5.
Method
Upd.
Setting
Reg.
PickScore
HPSv2.1
HPSv3
ImageReward
CLIPScore
Aesthetic
Div.
Dist.
Flow-GRPO, PickScore adapter
∼ 4k
4.5/4.5
KL
23.52
0.316
7.04
1.27
0.280
5.90
0.202
0.313
MEND , trained with guidance (Protocol F)
100
4.5/4.5
none
23.56
0.304
5.08
1.07
0.277
5.94
0.214
0.280
Appendix
Table 4: Training with guidance. Protocol F and the Flow-GRPO PickScore adapter, DrawBench 200×5 ; both train and sample at guidance 4.5, with different budgets and regularizers.
Figure 13: Guidance sweep in images. Guidance 1, 2, 3 and 4.5 (columns) for SD3.5-M, Flow-GRPO, DiffusionNFT and MEND at update 100 (rows); two val64 prompts, seed 43, 512 pixels, selected by eye. DiffusionNFT over-saturates as guidance grows.
Figure 17
Table 5: Full SD3.5-M results. DrawBench 200×5 ; setting is training/evaluation guidance, † marks rows from Zhou et al. (2026) (n/r: not reported) and ∗ the default Adam ϵ . Grey sub-rows give 95% prompt-bootstrap intervals; bold and underline mark the best and second best trained headline rows.