Organizations: School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China · Guangdong Key Laboratory of Big Data Analysis and Processing, Guangzhou, China
Real-time robot control demands fast action generation. Diffusion and flow matching policies for robot control require multi-step sampling, limiting their deployment in real-time scenarios. Natively reducing the sampling steps to one sacrifices representation quality and task performance, creating a trilemma among speed, fidelity, and performance. We present One-Step Generative Policy Optimization (OGPO), a systematic framework to resolve this trilemma. OGPO first pairs a lightweight architecture with the interval velocity principle for distillation-free one-step inference, while representation spreading prevents representation quality degradation. It then performs on-policy reinforcement learning (RL) fine-tuning on this fast, stable policy to break the imitation learning ceiling. Experiments on RoboMimic and OpenAI Gym benchmarks show that OGPO matches or exceeds multi-step baselines while achieving a 5-20 times inference speedup and over 120Hz control frequency. Physical deployment on a Franka-Emika-Panda robot validates real-world applicability. Project page: https://ogpo-project.github.io/
Figures & tables
Figure 1 : OGPO Framework Overview. Stage 1 (Top & Middle): Pre-training uses a Compact Velocity Field (CVF) for one-step inference, combined with representation spreading to prevent representation quality degradation. Stage 2 (Bottom): On-policy RL fine-tuning, formulated as a two-layer policy factorization: the outer layer is the environment Markov decision process (MDP), while the inner latent chain reparameterizes the action distribution.
Figure 2
Figure 2 : Stage 1 (Pre-training): Success rate vs. denoising steps on a 4 × 3 grid. Rows correspond to Lift, Can, Square, and Transport tasks with increasing difficulty, and columns correspond to αspre∈{0.1,0.5,0.9} . CVF variants, CVF and CVF+Spread, achieve near-saturated performance at 1–5 steps, while ReFlow and ShortCut require 32–128 steps. Representation spreading reduces variance on complex tasks.
Figure 3 : Stage 2 (PPO Fine-tuning): Performance on RoboMimic manipulation tasks. CVF in blue is compared with DPPO in yellow, Gaussian in green, ReinFlow-S in red, and ReinFlow-R in purple. Lift is omitted as CVF achieves 100% success during Stage 1 pretraining.
Figure 4 : Stage 2 (PPO Fine-tuning): Performance on OpenAI Gym locomotion tasks in panels a–d and Kitchen manipulation tasks in panels e–g. CVF in blue is compared with DPPO in yellow, ReinFlow-S in red, ReinFlow-R in purple, and FQL in light blue.
Data
Method
Venue
NFE
Dist.
Lift
Can
Square
Transport
Avg.
Full Dataset (300 trajectories)
Full
LSTM-GMM
CoRL’21 [ 34 ]
-
✗
0.93
0.81
0.59
0.20
0.63
Full
IBC
CoRL’21 [ 38 ]
-
✗
0.02
0.01
0.00
0.00
0.01
Full
BET
NeurIPS’22 [ 42 ]
-
✗
0.99
0.90
0.43
0.06
0.60
Full
DP-C
RSS’23 [ 11 ]
100
✗
0.97
0.96
0.82
0.46
0.80
Full
DP-T
RSS’23 [ 11 ]
100
✗
1.00
0.94
0.81
0.35
0.78
Table 1 : Stage 2 (RL Fine-tuning): Performance comparison on RoboMimic (Multi-Human data). Full dataset: 300 trajectories. Simplified: 100 trajectories. OGPO results are after Stage 2 PPO fine-tuning.
NVIDIA RTX 4090
NVIDIA RTX 2080
Model
Vision
Action
Params
Size
Steps
Time
Freq
Speedup
Time
Freq
Speedup
DP
ResNet-18 × 2
UNet
281.19M
∼ 4.4GB
100 (DDPM)
391.1ms
2.6Hz
1 ×
2007.5ms
0.5Hz
1 ×
16 (DDIM)
63.7ms
15.7Hz
6 ×
385.3ms
2.6Hz
5 ×
10 (DDIM)
40.3ms
24.8Hz
10 ×
220.4ms
4.5Hz
9 ×
CP
ResNet-18 × 2
UNet
284.86M
∼ 4.4GB
1
5.4ms
187Hz
73 ×
35.5ms
28Hz
56 ×
ReFlow
light ViT
MLP
1.78M
∼ 28MB
20
8.4ms
119Hz
46 ×
47.7ms
21Hz
42 ×
Table 2 : Stage 2 (RL Fine-tuning): Model Architecture and Efficiency Comparison. Speedup is relative to DP-DDPM (100 steps). All measurements use batch size 1.
Figure 5 : Real-world deployment of OGPO on a Franka Panda robot. Left: Hardware setup with Intel RealSense D435i camera for visual observation. Right: Qualitative comparison between MP1 baseline (top row) and our OGPO method (rows 2-3) across four manipulation tasks. MP1 fails on Lift and Can tasks due to imprecise grasping, while OGPO successfully completes all four tasks (Lift, Can, Square, Transport), demonstrating superior sim-to-real transfer capability. Green checkmarks indicate successful task completion; red crosses indicate failure.
Generative control policies (GCPs), such as diffusion- and flow-based control policies, have emerged as effective parameterizations for robot learning. This work introduces Off-policy Generative Policy Optimization (OGPO), a sample-efficient algorithm for finetuning GCPs that maintains off-policy critic networks to maximize data reuse and propagate policy gradients through the full generative process of the policy via a modified PPO objective, using critics as the terminal reward. OGPO achieves state-of-the-art performance on manipulation tasks spanning multi-task settings, high-precision insertion, and dexterous control. To our knowledge, it is also the only method that can fine-tune poorly-initialized behavior cloning policies to near full task-success with no expert data in the online replay buffer, and does so with few task-specific hyperparameter tuning. Through extensive empirical investigations, we demonstrate that OGPO drastically outperforms methods alternatives on policy steering and learning residual corrections, and identify the key mechanisms behind its performance. We further introduce practical stabilization tricks, including success-buffer regularization, two-sided conservative advantages, and Q-variance reduction, to mitigate critic over-exploitation across state- and pixel-based settings. Beyond proposing OGPO, we conduct a systematic empirical study of GCP finetuning, identifying the stabilizing mechanisms and failure modes that govern successful off-policy full-policy improvement.
Generative models such as diffusion and flow matching have advanced robotic visuomotor policies by modeling multimodal action distributions, but their multi-step sampling or ODE solving introduces inference latency. Existing one-step acceleration methods often compress the whole generation process into a single large update, leading to spatial deviation, frequency distortion, and mode averaging. This paper proposes a high-fidelity one-step generative visuomotor policy framework that addresses these issues with three complementary mechanisms. Recursive Consistent Action Flow (RCAF) uses recursive correction to compensate for spatial truncation errors and align one-step predictions with refined flow trajectories. Dual-Timestep Frequency Consistency (DTFC) preserves high-frequency manipulation details through adaptive spectral consistency across flow timesteps. Contrastive Flow Matching (CFM) separates entangled action flows with a margin-based repulsive objective, reducing ambiguous actions in multimodal manipulation. Experiments on RoboTwin, RoboTwin 2.0, Adroit, DexArt, and real-world robot platforms show that the proposed method achieves competitive or superior performance compared with strong 10-step generative policy baselines while requiring only one forward pass (1 NFE), enabling low-latency visuomotor control.
Yuran Chen, Xinye Cai, Zhonglin Gong +1
School of Safety Science and Engineering, Anhui University of Science and Technology, Huainan 232001, China · State Key Laboratory of Digital Intelligent Technology for Unmanned Coal Mining, Anhui University of Science and Technology, Huainan 232001, China · School of Artificial Intelligence, Anhui University of Science and Technology, Hefei 231131, China
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates (≥92% across all tasks), a per-episode inference time of 31.40ms (up to 175× faster than diffusion policies and 18× faster than prior flow matching policies), up to 4× faster training convergence than ACT, and 5× to 7× reduction in controller tracking error compared to discrete-action baselines.
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
MARS Lab, Nanyang Technological University, Singapore