MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators
Authors: Haocheng Tang, Tianchi Xie, Xingqiao Lin
Organizations: Northeastern University Boston, MA 02115, USA · Tsinghua University Beijing, 100084, PRC · Carnegie Mellon University Pittsburgh, PA 15213, USA
MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent x0-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.
Figures & tables
Method
ImageReward ↑
CLIPScore ↑
Aesthetic ↑
PickScore ↑
HPSv2 ↑
HPSv3 ↑
GenEval2 ↑
OCR ↑
Multi-step models (40 steps)
SD3.5-M †
−0.463
0.240
5.178
20.77
0.208
2.769
0.104
0.134
+ DiffusionNFT † ( Zheng et al., 2026 )
1.426
0.299
5.492
23.55
0.330
13.893
0.224
0.639
Few-step models (4 steps)
DMD ( Yin et al., 2024 )
0.924
0.284
5.506
22.28
0.287
11.716
0.204
0.400
CDM ( Liu et al., 2026 )
1.031
0.282
5.572
22.42
0.298
12.519
0.202
0.323
Table 1: Text-to-image results on SD3.5-Medium. All rows without a mark are copied from Table 1 of MeanFlowNFT ( Huang et al., 2026b ) , which evaluates at 1024×1024 following DiffusionNFT. † Evaluated with our pipeline (same prompts, resolution, seeds, and scorers as MFA). Among few-step models, bold marks the best value and underline the second best. PickScore is the raw logit.
Figure 1: Qualitative comparison on prompts from GenEval, OCR, and DrawBench.
Figure 3
Figure 5: Training reward and held-out score of the matched MFA ablations. Boundary-only training collapses late; the positive-only and A≡1 runs collapse within the first few dozen steps and are stopped at step 300.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
MeanFlowNFT
MeanFlowAdvantage
Init
AnyFlow + fresh LoRA
AnyFlow + fresh LoRA
LoRA rank / α
32 / 64
32 / 64
Image size (train)
512
512
Rollout NFE / CFG
4 / none
4 / none
Prompt groups L per update
48
48
Optimizer
AdamW
AdamW
Appendix
Table S1: SD3.5 Medium RL (Stage 3 protocol of MeanFlowNFT). Sampling and reward choices follow the baseline; the objectives and regularization coefficients differ.
MeanFlowNFT
MeanFlowAdvantage
Init
RMF ckpts/0120
RMF ckpts/0120
Signals per update
8
8
Group size K
8
8
Rollout NFE
1
1
(s,t) mix (boundary, s=1 , general)
(0.5,0.25,0.25)
(0.5,0.25,0.25)
t sampler
logit-normal (−0.4,1)
logit-normal (−0.4,1)
Appendix
Table S2: FANTOM5 on-policy DNA RL (no teacher). Both methods share every row above the line; only the loss coefficients differ.
MeanFlowNFT
MeanFlowAdvantage
Seed
Sei MSE ↓
6 -mer ↑
Sei MSE ↓
6 -mer ↑
Δ [ 95% CI]
123
0.0445
0.927
0.0437
0.952
0.0008[−0.0031,0.0046]
7
0.0526
0.940
0.0537
0.954
−0.0012[−0.0051,0.0026]
42
0.0441
0.944
0.0472
0.953
−0.0031[−0.0069,0.0005]
Mean
0.0471
0.937
0.0482
0.953
—
Std.
0.0048
0.009
0.0051
0.001
—
Appendix
Table S3: Per-seed on-policy DNA results behind Table S6 . Checkpoints are selected on valid ( n=256 ) and reported on test ( n=512 ), both with eval seed 7 . Δ is the paired NFT − MFA test Sei MSE with a 95% sequence bootstrap interval; positive favors MFA.
MeanFlowNFT
MeanFlowAdvantage
Init
RMF ckpts/0120
RMF ckpts/0120
Sequence length
1024 bp
1024 bp
Train / valid / test chr.
not 8 – 10 / 10 / 8 – 9
not 8 – 10 / 10 / 8 – 9
Batch size
8
8
Student NFE
1 (unguided)
1 (unguided)
Teacher NFE / guidance
10 / 10
10 / 10
Appendix
Table S4: FANTOM5 reward-filtered promoter distillation. Shared rows are identical for both losses; method columns list only the knobs that differ.
Figure S1: More Qualitative Comparison Cases. The prompts are taken from GenEval, OCR and DrawBench respectively,where we compare the corresponding MeanFlowNFT model with our model.
Variant
Train step
Reward ↑
Held-out Sum ↑
Grad. norm ↓
s/step ↓
Shared induced V (control)
2000
1.5493
8.9841
0.97
73.3
Direct u (no induced correction)
2000
1.5071
8.8915
2.10
64.8
s=t diffusion-only
2000
0.8734
2.9590
445.79
65.4
Separate derivative
2000
1.5486
9.0382
1.00
89.9
No adaptive w
2000
1.4959
8.6952
11.97
87.4
γ=5
2000
1.5424
8.9936
0.84
72.9
Appendix
Table S5: Matched image-model ablations. “Reward” is the mean training reward over the final 50 recorded steps of each run; “Held-out Sum” is the fixed DrawBench online aggregate used for the ablation study. The early-stopped positive-only runs are not equal-budget endpoints.
Figure S2: Optimization stability across MeanFlowAdvantage ablations, highlighting gradient norms, loss dynamics, and collapse behavior.
Method
Sei MSE ↓
6 -mer corr. ↑
Δ [ 95% CI]
Pretrained RMF
0.0714
0.928
—
On-policy RL, no teacher (3 seeds)
MeanFlowNFT
0.0471±0.0048
0.937±0.009
−0.0012[−0.0051,0.0026]
MeanFlowAdvantage
0.0482±0.0051[0.0437]
0.953±0.001
Reward-graded distillation from a 10 -step Sei-guided teacher
MeanFlowNFT
0.0570
0.955
0.0103[0.0064,0.0143]
Appendix
Table S6: One-step promoter generation on the test chromosomes ( n=512 ), from the same RMF initialization and the same protocol within each block. Sei MSE is the training reward; the 6 -mer correlation is never optimized. On-policy rows are mean ± s.d. over three training seeds, with the best seed in brackets; distillation rows are the single reported run. Δ is the paired NFT − MFA Sei MSE with a 95% sequence bootstrap interval, so a positive Δ favors MFA.
Figure S3: Reward-graded MeanFlow distillation. (a) A frozen rollout and a Sei-guided teacher define the reward gap Δr and the advantage A . (b)–(c) Validation curves ( n=256 ); bands are sequence-bootstrap 95% intervals, and stars mark the checkpoints selected by Sei MSE.
MeanFlowNFT
MeanFlowAdvantage
Seed
Sei MSE ↓
6 -mer ↑
Sei MSE ↓
6 -mer ↑
Δ [ 95% CI]
Red.
123
0.0570
0.955
0.0467
0.953
0.0103[0.0064,0.0143]
18.1%
42
0.0585
0.957
0.0467
0.953
0.0118[0.0075,0.0162]
20.2%
456
0.0586
0.958
0.0477
0.947
0.0110[0.0068,0.0152]
18.7%
789
0.0594
0.958
0.0465
0.949
0.0129[0.0084,0.0173]
21.7%
Mean
0.0584
0.957
0.0469
0.951
0.0115
19.7%
Appendix
Table S7: Reward-graded distillation repeated over four training seeds. Checkpoints are selected on valid ( n=256 ) and reported on test ( n=512 , eval seed 7 ). Δ is the paired NFT − MFA Sei MSE, with a 95% bootstrap over sequences; positive favors MFA. Seed 123 is the run reported in Table S6 .
lr
κ
NFT MSE ↓
MFA MSE ↓
Δ [ 95% CI]
10−5
1.25
0.0550
0.0545
0.0005[−0.0007,0.0018]
10−5
1.5
0.0490
0.0490
0.0000[−0.0010,0.0010]
2.5×10−5
1.25
0.0525
0.0518
0.0007[−0.0013,0.0028]
2.5×10−5
1.5
0.0474
0.0482
−0.0009[−0.0025,0.0008]
Appendix
Table S8: A≡1 special case under matched learning rate and nominal extrapolation. Checkpoints are selected on valid ( n=256 ) and evaluated on test ( n=512 ). Δ is paired NFT − MFA Sei MSE with a 95% bootstrap interval.
Method
step
Sei MSE ↓
6 -mer ↑
Pretrained RMF
0
0.0754
0.924
MeanFlowNFT
70
0.0553
0.951
MeanFlowAdvantage
55
0.0501
0.950
Appendix
Table S9: Individually tuned A≡1 special case on valid 1 -NFE FANTOM5 ( n=512 , chr10). Online selection used n=256 .
step
Δ MSE
95% CI
0
0.000
[0.000,0.000]
10
0.0103
[0.0065,0.0141]
50
0.0065
[0.0031,0.0101]
60
0.0058
[0.0025,0.0092]
70
0.0052
[0.0022,0.0083]
100
0.0081
[0.0044,0.0116]
Appendix
Table S10: Paired valid Sei Δ MSE (NFT − MFA) for the individually tuned A≡1 special case, n=512 , 2000 bootstrap draws. The interval is the 95% percentile CI.
MeanFlowNFT
MeanFlowAdvantage
Seed
step
Sei MSE ↓
6 -mer ↑
step
Sei MSE ↓
6 -mer ↑
123†
70
0.0542
0.956
55
0.0486
0.958
42
60
0.0550
0.955
55
0.0492
0.955
456
100
0.0550
0.961
45
0.0496
0.951
789
90
0.0551
0.955
35
0.0479
0.953
Mean
—
0.0548
0.957
—
0.0488
0.954
Appendix
Table S11: FANTOM5 training-seed ablation for the individually tuned A≡1 special case. All runs start from the same RMF ckpts/0120 snapshot. Checkpoints are selected on valid ( n=256 , eval seed 7 ) and reported on test ( n=512 , eval seed 7 ). Only the training seed changes.
Figure S4: Promoter analyses that are not in Figure S3 . (a)–(b) Test 1 -NFE Sei MSE and 6 -mer versus pretrain epoch ( n=256 ). The star is the RL init ( ckpts/0120 ); the shaded region is the collapsed tail. (c) Paired valid Δ MSE (NFT − MFA) with 95% sequence-level bootstrap CIs ( n=512 ). The star is the on-grid NFT pick (step 70 ). (d) Test Sei MSE at the tuned A≡1 checkpoints versus CAGE-sorted eval n∈{256,512,1024,2048} .
Figure S5: Training-pair Sei MSE of accepted distillation pairs (length- 15 rolling mean). Solid: 1 -NFE student. Dashed: 10 -NFE teacher. These scores are computed for the accept filter ( ΔrSei≥0.002 ); Table S6 instead uses hard-decoded held-out sequences scored by Sei.