cs.AISep 29, 2026

MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators

Authors: Haocheng Tang, Tianchi Xie, Xingqiao Lin

Organizations: Northeastern University Boston, MA 02115, USA · Tsinghua University Beijing, 100084, PRC · Carnegie Mellon University Pittsburgh, PA 15213, USA

Abstract

MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent x0x_0-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves MeanFlow's native few-step sampler and provides a direct mechanism for transferring reward improvements to the deployed flow map. On SD3.5-Medium, MeanFlowAdvantage improves all eight reported metrics over the matched four-step MeanFlowNFT baseline and, with only four NFEs, matches or exceeds the 40-step DiffusionNFT baseline on six of eight metrics. The same objective also transfers to DNA promoter design, where it supports both teacher-free on-policy RL for a generator defined on a manifold and teacher-guided reward-graded distillation, with the latter yielding the lowest one-step Sei profile MSE among the compared configurations.

Figures & tables

Appendix figures & tables16 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. FlowAWR: Online Adaptive Flow Reinforcement via Advantage-Weighted Rectification

    Jun 29, 2026Zheming Fu, Ruizhe He, Wei Shang +4Generative Flow NetworksVelocity Field

  2. AdvantageFlow: Regularized Advantage-Weighted RL in Flow Models

    May 25, 2026Branislav Kveton, Anup Rao, Subhojyoti Mukherjee +2Flow ModelsRectified Flow

  3. Score-Based One-step MeanFlow Policy Optimization

    May 22, 2026Kyungyoon Kim, Donghyeon Ki, Hee-Jun Ahn +1Flow PoliciesMeanflow