MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators
Authors: Haocheng Tang, Tianchi Xie, Xingqiao Lin
Organizations: Northeastern University Boston, MA 02115, USA · Tsinghua University Beijing, 100084, PRC · Carnegie Mellon University Pittsburgh, PA 15213, USA
MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent x0-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.
Figures & tables
Method
ImageReward ↑
CLIPScore ↑
Aesthetic ↑
PickScore ↑
HPSv2 ↑
HPSv3 ↑
GenEval2 ↑
OCR ↑
Multi-step models (40 steps)
SD3.5-M †
−0.463
0.240
5.178
20.77
0.208
2.769
0.104
0.134
+ DiffusionNFT † ( Zheng et al., 2026 )
1.426
0.299
5.492
23.55
0.330
13.893
0.224
0.639
Few-step models (4 steps)
DMD ( Yin et al., 2024 )
0.924
0.284
5.506
22.28
0.287
11.716
0.204
0.400
CDM ( Liu et al., 2026 )
1.031
0.282
5.572
22.42
0.298
12.519
0.202
0.323
Table 1: Text-to-image results on SD3.5-Medium. All rows without a mark are copied from Table 1 of MeanFlowNFT ( Huang et al., 2026b ) , which evaluates at 1024×1024 following DiffusionNFT. † Evaluated with our pipeline (same prompts, resolution, seeds, and scorers as MFA). Among few-step models, bold marks the best value and underline the second best. PickScore is the raw logit.
Figure 1: Qualitative comparison on prompts from GenEval, OCR, and DrawBench.
Figure 3
Figure 5: Training reward and held-out score of the matched MFA ablations. Boundary-only training collapses late; the positive-only and A≡1 runs collapse within the first few dozen steps and are stopped at step 300.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
MeanFlowNFT
MeanFlowAdvantage
Init
AnyFlow + fresh LoRA
AnyFlow + fresh LoRA
LoRA rank / α
32 / 64
32 / 64
Image size (train)
512
512
Rollout NFE / CFG
4 / none
4 / none
Prompt groups L per update
48
48
Optimizer
AdamW
AdamW
Appendix
Table S1: SD3.5 Medium RL (Stage 3 protocol of MeanFlowNFT). Sampling and reward choices follow the baseline; the objectives and regularization coefficients differ.
MeanFlowNFT
MeanFlowAdvantage
Init
RMF ckpts/0120
RMF ckpts/0120
Signals per update
8
8
Group size K
8
8
Rollout NFE
1
1
(s,t) mix (boundary, s=1 , general)
(0.5,0.25,0.25)
(0.5,0.25,0.25)
t sampler
logit-normal (−0.4,1)
logit-normal (−0.4,1)
Appendix
Table S2: FANTOM5 on-policy DNA RL (no teacher). Both methods share every row above the line; only the loss coefficients differ.
MeanFlowNFT
MeanFlowAdvantage
Seed
Sei MSE ↓
6 -mer ↑
Sei MSE ↓
6 -mer ↑
Δ [ 95% CI]
123
0.0445
0.927
0.0437
0.952
0.0008[−0.0031,0.0046]
7
0.0526
0.940
0.0537
0.954
−0.0012[−0.0051,0.0026]
42
0.0441
0.944
0.0472
0.953
−0.0031[−0.0069,0.0005]
Mean
0.0471
0.937
0.0482
0.953
—
Std.
0.0048
0.009
0.0051
0.001
—
Appendix
Table S3: Per-seed on-policy DNA results behind Table S6 . Checkpoints are selected on valid ( n=256 ) and reported on test ( n=512 ), both with eval seed 7 . Δ is the paired NFT − MFA test Sei MSE with a 95% sequence bootstrap interval; positive favors MFA.
MeanFlowNFT
MeanFlowAdvantage
Init
RMF ckpts/0120
RMF ckpts/0120
Sequence length
1024 bp
1024 bp
Train / valid / test chr.
not 8 – 10 / 10 / 8 – 9
not 8 – 10 / 10 / 8 – 9
Batch size
8
8
Student NFE
1 (unguided)
1 (unguided)
Teacher NFE / guidance
10 / 10
10 / 10
Appendix
Table S4: FANTOM5 reward-filtered promoter distillation. Shared rows are identical for both losses; method columns list only the knobs that differ.
Figure S1: More Qualitative Comparison Cases. The prompts are taken from GenEval, OCR and DrawBench respectively,where we compare the corresponding MeanFlowNFT model with our model.
Variant
Train step
Reward ↑
Held-out Sum ↑
Grad. norm ↓
s/step ↓
Shared induced V (control)
2000
1.5493
8.9841
0.97
73.3
Direct u (no induced correction)
2000
1.5071
8.8915
2.10
64.8
s=t diffusion-only
2000
0.8734
2.9590
445.79
65.4
Separate derivative
2000
1.5486
9.0382
1.00
89.9
No adaptive w
2000
1.4959
8.6952
11.97
87.4
γ=5
2000
1.5424
8.9936
0.84
72.9
Appendix
Table S5: Matched image-model ablations. “Reward” is the mean training reward over the final 50 recorded steps of each run; “Held-out Sum” is the fixed DrawBench online aggregate used for the ablation study. The early-stopped positive-only runs are not equal-budget endpoints.
Figure S2: Optimization stability across MeanFlowAdvantage ablations, highlighting gradient norms, loss dynamics, and collapse behavior.
Method
Sei MSE ↓
6 -mer corr. ↑
Δ [ 95% CI]
Pretrained RMF
0.0714
0.928
—
On-policy RL, no teacher (3 seeds)
MeanFlowNFT
0.0471±0.0048
0.937±0.009
−0.0012[−0.0051,0.0026]
MeanFlowAdvantage
0.0482±0.0051[0.0437]
0.953±0.001
Reward-graded distillation from a 10 -step Sei-guided teacher
MeanFlowNFT
0.0570
0.955
0.0103[0.0064,0.0143]
Appendix
Table S6: One-step promoter generation on the test chromosomes ( n=512 ), from the same RMF initialization and the same protocol within each block. Sei MSE is the training reward; the 6 -mer correlation is never optimized. On-policy rows are mean ± s.d. over three training seeds, with the best seed in brackets; distillation rows are the single reported run. Δ is the paired NFT − MFA Sei MSE with a 95% sequence bootstrap interval, so a positive Δ favors MFA.
Figure S3: Reward-graded MeanFlow distillation. (a) A frozen rollout and a Sei-guided teacher define the reward gap Δr and the advantage A . (b)–(c) Validation curves ( n=256 ); bands are sequence-bootstrap 95% intervals, and stars mark the checkpoints selected by Sei MSE.
MeanFlowNFT
MeanFlowAdvantage
Seed
Sei MSE ↓
6 -mer ↑
Sei MSE ↓
6 -mer ↑
Δ [ 95% CI]
Red.
123
0.0570
0.955
0.0467
0.953
0.0103[0.0064,0.0143]
18.1%
42
0.0585
0.957
0.0467
0.953
0.0118[0.0075,0.0162]
20.2%
456
0.0586
0.958
0.0477
0.947
0.0110[0.0068,0.0152]
18.7%
789
0.0594
0.958
0.0465
0.949
0.0129[0.0084,0.0173]
21.7%
Mean
0.0584
0.957
0.0469
0.951
0.0115
19.7%
Appendix
Table S7: Reward-graded distillation repeated over four training seeds. Checkpoints are selected on valid ( n=256 ) and reported on test ( n=512 , eval seed 7 ). Δ is the paired NFT − MFA Sei MSE, with a 95% bootstrap over sequences; positive favors MFA. Seed 123 is the run reported in Table S6 .
lr
κ
NFT MSE ↓
MFA MSE ↓
Δ [ 95% CI]
10−5
1.25
0.0550
0.0545
0.0005[−0.0007,0.0018]
10−5
1.5
0.0490
0.0490
0.0000[−0.0010,0.0010]
2.5×10−5
1.25
0.0525
0.0518
0.0007[−0.0013,0.0028]
2.5×10−5
1.5
0.0474
0.0482
−0.0009[−0.0025,0.0008]
Appendix
Table S8: A≡1 special case under matched learning rate and nominal extrapolation. Checkpoints are selected on valid ( n=256 ) and evaluated on test ( n=512 ). Δ is paired NFT − MFA Sei MSE with a 95% bootstrap interval.
Method
step
Sei MSE ↓
6 -mer ↑
Pretrained RMF
0
0.0754
0.924
MeanFlowNFT
70
0.0553
0.951
MeanFlowAdvantage
55
0.0501
0.950
Appendix
Table S9: Individually tuned A≡1 special case on valid 1 -NFE FANTOM5 ( n=512 , chr10). Online selection used n=256 .
step
Δ MSE
95% CI
0
0.000
[0.000,0.000]
10
0.0103
[0.0065,0.0141]
50
0.0065
[0.0031,0.0101]
60
0.0058
[0.0025,0.0092]
70
0.0052
[0.0022,0.0083]
100
0.0081
[0.0044,0.0116]
Appendix
Table S10: Paired valid Sei Δ MSE (NFT − MFA) for the individually tuned A≡1 special case, n=512 , 2000 bootstrap draws. The interval is the 95% percentile CI.
MeanFlowNFT
MeanFlowAdvantage
Seed
step
Sei MSE ↓
6 -mer ↑
step
Sei MSE ↓
6 -mer ↑
123†
70
0.0542
0.956
55
0.0486
0.958
42
60
0.0550
0.955
55
0.0492
0.955
456
100
0.0550
0.961
45
0.0496
0.951
789
90
0.0551
0.955
35
0.0479
0.953
Mean
—
0.0548
0.957
—
0.0488
0.954
Appendix
Table S11: FANTOM5 training-seed ablation for the individually tuned A≡1 special case. All runs start from the same RMF ckpts/0120 snapshot. Checkpoints are selected on valid ( n=256 , eval seed 7 ) and reported on test ( n=512 , eval seed 7 ). Only the training seed changes.
Figure S4: Promoter analyses that are not in Figure S3 . (a)–(b) Test 1 -NFE Sei MSE and 6 -mer versus pretrain epoch ( n=256 ). The star is the RL init ( ckpts/0120 ); the shaded region is the collapsed tail. (c) Paired valid Δ MSE (NFT − MFA) with 95% sequence-level bootstrap CIs ( n=512 ). The star is the on-grid NFT pick (step 70 ). (d) Test Sei MSE at the tuned A≡1 checkpoints versus CAGE-sorted eval n∈{256,512,1024,2048} .
Figure S5: Training-pair Sei MSE of accepted distillation pairs (length- 15 rolling mean). Solid: 1 -NFE student. Dashed: 10 -NFE teacher. These scores are computed for the accept filter ( ΔrSei≥0.002 ); Table S6 instead uses hard-decoded held-out sequences scored by Sei.
Aligning generative flow models on continuous spaces via online reinforcement learning is constrained by intractable trajectory likelihoods. Existing density-approximated policy gradient methods rely on stochastic SDE samplers to construct tractable transition kernels, which introduce training-inference inconsistencies and necessitates Classifier-Free Guidance (CFG). While implicit frameworks such as DiffusionNFT directly optimize forward-process velocity fields, its heuristic fixed-magnitude corrections prevent optimization strength from relative intra-group quality. We propose \textit{Flow Advantage-Weighted Rectification} (\textbf{FlowAWR}), a paradigm that recasts continuous generative policy optimization as supervised regression toward a theoretically optimal velocity field. Starting from the optimal policy of a KL-constrained reward maximization, FlowAWR derives the optimal velocity field that admits a magnitude-aware, advantage-weighted rectification form, yielding SDE-free optimization and CFG-free generation. In comparative evaluations on SD3.5-Medium, FlowAWR achieves improved alignment performance alongside a 2× to 5× convergence acceleration over DiffusionNFT (e.g., reaching a 24.12 PickScore in 1.2k steps, versus 23.82 in 2.0k steps for DiffusionNFT and 23.50 in >4k steps for FlowGRPO). Under multi-reward constraints, FlowAWR sustains generation quality, satisfying structural rules while maintaining stable out-of-domain performance.
Zheming Fu, Ruizhe He, Wei Shang +4
Beihang University · Joy Future Academy · Zhongguancun Academy +2
We present AdvantageFlow, a forward-process reinforcement learning (RL) algorithm for rectified flow models. The algorithm minimizes an advantage-weighted prediction loss, which maximizes reward, regularized by the rollout policy, which convexifies the objective and makes its optimization stable. Our objective can be viewed as fitting a local reward-improving target distribution. The rollout regularization arises as a variance reduction step. We evaluate AdvantageFlow empirically on text-to-image generation with Stable Diffusion 3.5 Medium and FLUX.1, and compare it to both forward- and reverse-process RL algorithms.
Diffusion and flow matching have emerged as expressive policy classes in reinforcement learning, but their reliance on multi-step denoising imposes substantial computational overhead at inference time, which is particularly problematic in online RL. MeanFlow offers a promising alternative by learning an average velocity field that maps noise to data in a single network evaluation. However, MeanFlow typically requires samples from the target distribution to construct its target velocity field, which are unavailable in online RL. We propose Score-Based One-step MeanFlow Policy Optimization (SOM), an actor-critic algorithm that resolves this by constructing the target velocity field directly from the Q-function via score estimation and a probability flow ODE, thereby concentrating probability mass on high-value modes. In the fully online RL setting, SOM achieves state-of-the-art performance on locomotion tasks with a single generation step, while substantially reducing both training and inference time compared to prior diffusion- and flow-matching-based policies.
Kyungyoon Kim, Donghyeon Ki, Hee-Jun Ahn +1
Korea University, Decision Making Lab · Gauss Labs Inc.