OGPO: One-Step Generative Policy Optimization for Real-Time Robot Control
Organizations: School of Computer Science and Engineering, Sun Yat-sen University, Guangzhou, China · Guangdong Key Laboratory of Big Data Analysis and Processing, Guangzhou, China
Abstract
Real-time robot control demands fast action generation. Diffusion and flow matching policies for robot control require multi-step sampling, limiting their deployment in real-time scenarios. Natively reducing the sampling steps to one sacrifices representation quality and task performance, creating a trilemma among speed, fidelity, and performance. We present One-Step Generative Policy Optimization (OGPO), a systematic framework to resolve this trilemma. OGPO first pairs a lightweight architecture with the interval velocity principle for distillation-free one-step inference, while representation spreading prevents representation quality degradation. It then performs on-policy reinforcement learning (RL) fine-tuning on this fast, stable policy to break the imitation learning ceiling. Experiments on RoboMimic and OpenAI Gym benchmarks show that OGPO matches or exceeds multi-step baselines while achieving a 5-20 times inference speedup and over 120Hz control frequency. Physical deployment on a Franka-Emika-Panda robot validates real-world applicability. Project page: https://ogpo-project.github.io/
Figures & tables
| Data | Method | Venue | NFE | Dist. | Lift | Can | Square | Transport | Avg. |
| Full Dataset (300 trajectories) | |||||||||
| Full | LSTM-GMM | CoRL’21 [ 34 ] | - | ✗ | 0.93 | 0.81 | 0.59 | 0.20 | 0.63 |
| Full | IBC | CoRL’21 [ 38 ] | - | ✗ | 0.02 | 0.01 | 0.00 | 0.00 | 0.01 |
| Full | BET | NeurIPS’22 [ 42 ] | - | ✗ | 0.99 | 0.90 | 0.43 | 0.06 | 0.60 |
| Full | DP-C | RSS’23 [ 11 ] | 100 | ✗ | 0.97 | 0.96 | 0.82 | 0.46 | 0.80 |
| Full | DP-T | RSS’23 [ 11 ] | 100 | ✗ | 1.00 | 0.94 | 0.81 | 0.35 | 0.78 |
| NVIDIA RTX 4090 | NVIDIA RTX 2080 | ||||||||||
| Model | Vision | Action | Params | Size | Steps | Time | Freq | Speedup | Time | Freq | Speedup |
| DP | ResNet-18 2 | UNet | 281.19M | 4.4GB | 100 (DDPM) | 391.1ms | 2.6Hz | 1 | 2007.5ms | 0.5Hz | 1 |
| 16 (DDIM) | 63.7ms | 15.7Hz | 6 | 385.3ms | 2.6Hz | 5 | |||||
| 10 (DDIM) | 40.3ms | 24.8Hz | 10 | 220.4ms | 4.5Hz | 9 | |||||
| CP | ResNet-18 2 | UNet | 284.86M | 4.4GB | 1 | 5.4ms | 187Hz | 73 | 35.5ms | 28Hz | 56 |
| ReFlow | light ViT | MLP | 1.78M | 28MB | 20 | 8.4ms | 119Hz | 46 | 47.7ms | 21Hz | 42 |