Flow-based policies provide an expressive framework for continuous robot control, but their iterative ODE integration incurs substantial inference cost. Naively reducing the integration budget can severely degrade control, since policies optimized under full-step execution are not explicitly constrained to remain reliable under coarse numerical integration. We refer to this mismatch as the few-step discretization gap. To address this problem, we introduce RFPO, a flow-policy optimization framework for reliable few-step execution. Reward-aware online Reflow rectifies student-induced transport paths during on-policy learning, making the resulting policy more robust to coarse integration. A frozen Gaussian PPO controller supplies complementary action-space supervision at full and intermediate integration budgets, while the deployed policy remains a single flow student executed with one Euler step. Across Unitree Go2, Boston Dynamics Spot, Unitree H1, and Unitree G1, RFPO consistently preserves full-step control performance under one-step execution, with one-step returns remaining within 2.4% of their corresponding 64-step values across both zero and random initialization. On Unitree Go2, one-step execution retains 98.5% of the 64-step reward while reducing onboard mean inference latency from 4.39 ms to 0.08 ms, yielding a 54.9x speedup. Real-robot experiments further validate stable one-step locomotion. Code: https://github.com/AIGeeksGroup/RFPO. Website: https://aigeeksgroup.github.io/RFPO.
Figures & tables
Fig. 1: Few-step execution of reward-optimized flow policies. Naive step reduction can create a discretization gap between full-step and few-step actions. RFPO rectifies student-induced transport paths and uses complementary action-space supervision to enable reliable few-step execution.
Fig. 2: Overview of RFPO. During on-policy training, student-induced full-step endpoints are rectified by reward-aware online Reflow, while multi-budget student actions are distilled toward a frozen PPO controller. An adaptive budget objective further regularizes the fidelity–compute trade-off. At deployment, all auxiliary components are removed and only the flow student is retained for one-step Euler execution.
Embodiment
Method
Zero Init.
Random Init.
64-step
1-step
64-step
1-step
Go2
PPO [ 24 ]
35.58
–
–
–
FPO++ [ 32 ]
21.29
-187.70
17.66
-220.10
KD-only
29.95
-4919.70
27.65
-5121.70
RFPO (Ours)
34.78
34.25
31.56
31.50
Spot
PPO [ 24 ]
30.84
–
–
–
TABLE I: Cross-embodiment few-step control. Average episodic reward under 64-step and one-step Euler execution across two quadruped and two humanoid embodiments. PPO uses a single deterministic forward pass and therefore has no integration budget. Reward scales differ across tasks; comparisons should be made within each embodiment. Higher is better.
Fig. 3: Representative one-step velocity tracking on Unitree Go2. Responses of FPO++, KD-only, RA-Reflow, and RFPO to the same forward command vxcmd=0.8 m/s. FPO++ and KD-only fail to track the target under one-step execution, whereas RA-Reflow remains dynamically stable but does not follow the commanded velocity. RFPO maintains stable command tracking.
Fig. 4: Qualitative sim-to-real deployment on Unitree Go2. Representative simulation and real-world executions of forward locomotion, lateral locomotion, and yaw rotation using the fixed one-step RFPO policy. Frames are sampled at comparable normalized phases of each motion sequence.
Method
Euler Steps
64
32
16
8
4
1
FPO++
21.29
19.06
18.78
18.56
15.43
-187.70
KD-only
29.95
29.59
29.91
26.63
-398.10
-4919.70
Reflow
33.57
33.50
33.50
33.45
33.18
32.84
RA-Reflow
32.78
32.98
33.12
32.59
33.00
33.40
Reflow+KD
34.65
34.44
34.77
34.77
34.92
34.39
TABLE II: Few-step robustness on Unitree Go2. Average episodic reward under zero initialization across Euler integration budgets. Higher is better.
Fig. 5: Action-flow geometry of RFPO on Unitree Go2. Normalized action-state densities over flow time for representative hip and thigh joints across all four legs. The learned policy exhibits concentrated and smoothly evolving transport patterns across joint dimensions.
Method
Tracking Error ↓
Fall Frac. @1 ↓
64-step
1-step
PPO
0.055
–
–
FPO++
0.792
0.883
0.680
KD-only
0.178
0.810
0.334
Reflow
0.080
0.070
0.000
RA-Reflow
0.086
0.805
0.000
TABLE III: Command tracking and stability on Unitree Go2. Mean absolute forward-velocity tracking error ∣vx−vxcmd∣ and one-step fall fraction. Lower is better. All evaluated 64-step flow policies have zero fall fraction.
Euler Steps
Mean (ms) ↓
p99 (ms) ↓
Speedup ↑
64
4.39
20.06
1.0 ×
32
1.50
1.56
2.9 ×
16
0.79
1.07
5.6 ×
8
0.41
0.44
10.7 ×
4
0.23
0.26
19.1 ×
1
0.08
0.10
54.9 ×
TABLE IV: Onboard inference efficiency on Unitree Go2. Single-thread ONNX Runtime latency on the onboard CPU.
Motion
Joint Pos. Corr. ↑
Joint Pos. MAE ↓
Forward ( vx )
0.992
0.085
Lateral ( vy )
0.995
0.079
Yaw Turn ( ωz )
0.997
0.055
TABLE V: Matched-command sim-to-real consistency on Unitree Go2. Joint-position agreement between simulation and hardware using the deployed one-step RFPO policy.
Real-time robot control demands fast action generation. Diffusion and flow matching policies for robot control require multi-step sampling, limiting their deployment in real-time scenarios. Natively reducing the sampling steps to one sacrifices representation quality and task performance, creating a trilemma among speed, fidelity, and performance. We present One-Step Generative Policy Optimization (OGPO), a systematic framework to resolve this trilemma. OGPO first pairs a lightweight architecture with the interval velocity principle for distillation-free one-step inference, while representation spreading prevents representation quality degradation. It then performs on-policy reinforcement learning (RL) fine-tuning on this fast, stable policy to break the imitation learning ceiling. Experiments on RoboMimic and OpenAI Gym benchmarks show that OGPO matches or exceeds multi-step baselines while achieving a 5-20 times inference speedup and over 120Hz control frequency. Physical deployment on a Franka-Emika-Panda robot validates real-world applicability. Project page: https://ogpo-project.github.io/
Guowei Zou, Haitao Wang, Hejun Wu +3
School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China · Guangdong Key Laboratory of Big Data Analysis and Processing, Guangzhou, China
We present Reflow-regularized Flow Matching Policy Gradients (ReFPO), a simple online RL method that adds explicit Reflow regularization to FPO for efficient flow-based control. We uncover a key structural property: the gradient updates in Flow Matching Policy Gradients (FPO) can be interpreted as an implicit advantage-weighted Reflow process, providing a new geometric perspective on flow-based policy gradients. Building on this insight, ReFPO introduces an explicit geometric regularizer that can be implemented with a single line of code change without incurring additional computational overhead or auxiliary distillation stages. By synergizing advantage-guided updates with path rectification, our method reduces CFM proxy-ratio spikes, stabilizes PPO-style training, and enables high-fidelity one-step inference that often matches or exceeds multi-step performance. We experimentally demonstrate that ReFPO improves average performance and discretization robustness across GridWorld, MuJoCo Playground, and high-dimensional Humanoid Control tasks, providing a scalable and stable approach for generative policies in complex physical simulations.
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates (≥92% across all tasks), a per-episode inference time of 31.40ms (up to 175× faster than diffusion policies and 18× faster than prior flow matching policies), up to 4× faster training convergence than ACT, and 5× to 7× reduction in controller tracking error compared to discrete-action baselines.
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
MARS Lab, Nanyang Technological University, Singapore