Diffusion- and flow-based robot policies have recently become widespread in robotic Imitation Learning (IL) due to their high performance and ability to model continuous and multimodal distributions. However, the iterative denoising procedure used by these models introduces significant prediction latency, hindering high-frequency closed-loop robot control and leading to jittery, unstable motion when frequent updates to the robot's action predictions are used. Therefore, it is common practice to train models to predict chunks of actions that can be executed sequentially without feedback, even when this reduces responsiveness and may mean the most recent state information is not used. In this article, we present Adaptive Mean Flow (AMF), a flow-based IL method that enables smooth and responsive, fully closed-loop robot control. AMF uses Mean Flow, which is an accelerated form of Flow Matching (FM), to minimize prediction latency. To ensure smoothness and consistency across predictions, AMF uses a corrupted version of the trajectory from the previous step when predicting new robot actions, with the signal-to-noise ratio increasing over the time parameter of the trajectory. This discourages large changes in the prediction from one step to the next, while allowing freedom to adapt the predictions for future steps. We evaluate AMF across a wide range of simulated and real robot tasks and demonstrate significantly improved performance compared with baselines. Code: https://github.com/akselva/Adaptive-mean-flow-RoboticIL.
Figures & tables
Fig. 1 : Trajectory generation with Adaptive Mean Flow visualized for a real-world conveyor belt picking task. a-b) Noisy points are sampled in the vicinity of the trajectory computed at the previous step. The noise increases for points along the path, corresponding to predictions farther into the future. c-d) The AMF denoising network computes a new trajectory by modifying the noisy sequence.
Fig. 2 : AMF inference. From the left: The previous action sequence prediction At− is corrupted with Gaussian noise according to Eq. 12 , producing AtS . Actions further in the future are more uncertain and, therefore, more corrupted. The Mean Flow network uθ(⋅) predicts the next action sequence At , using AtS and the most recent observation sequence Ot . The action at corresponds to the current time-step sent to the robot for execution, and At is shifted (Eq. 9 ) and reused as At− for the next control step.
Fig. 3 : Flow interval sampling for an example action sequence with To=2 and Tp=6 . Instead of using the same values s,r across the action sequence (left), AMF uses independent intervals si,ri . During training, we use a mix of independently sampled intervals (middle) and monotonically increasing intervals (right).
TABLE I : Main simulation results. For Push-T and Metaworld, the models are trained for 500 epochs, while for Dyn-BP and D3IL (Stack-1 and Stack-2) the models are trained for 200 and 90 epochs, respectively. We use a batch size of 256, and model checkpoints are collected every 10th epoch. For Push-T and Dyn-BP, the score is the maximum goal overlap ratio achieved during an episode, whereas for Metaworld and D3IL, the score is 1 for success and 0 otherwise. Each model rollout consists of 50 tests, and we report averaged metrics over the last three model checkpoints, across three seeds. Standard deviation is shown in parentheses.
Method
Mean
Max
FM-1
1.16
26.67
MF-1
2.19
29.69
MF-WS
1.33
23.63
AMF (Ours)
0.98
15.96
TABLE II : Trajectory smoothness analysis on Metaworld. We quantify action trajectory smoothness using the Euclidean norm of the third finite difference of consecutive Cartesian actions. We report its mean and maximum over evaluation trajectories, where lower values indicate smoother action sequences. The mean and maximum is computed for each task, then averaged across all tasks.
Fig. 4 : Success rate on 2 Metaworld tasks with increasing perturbations added to the end-effector. Each policy obtains a close to 100% success rate without perturbation; however, the performance of AMF (blue curve) consistently degrades the least as the perturbation magnitude σ (and thus the task difficulty) increases.
Ablation
Push-T (Overlap)
Metaworld Avg. (SR)
F-AdaLN-Z / pm=0.5
0.92 (0.02)
0.95
Concat / pm=0.5
0.86 (0.01)
0.90
Add / pm=0.5
0.85 (0.00)
0.92
F-AdaLN-Z / pm=0.0
0.89 (0.02)
0.93
F-AdaLN-Z / pm=1.0
0.88 (0.01)
0.77
S -conditioning only
0.9 (0.02)
0.93
TABLE III : Performance impact from selected design variations. The type of conditioning used for the flow intervals (S,R) is denoted in bold. pm is the probability of sampling interval sequences (S,R) using the monotone schedule (Eq. 12 ), leaving (1−pm) probability of sampling each (si,ri) independently.
Fig. 5 : Task score on Push-T for AMF and MF-WS with varying levels of noise corruption. For AMF, the value on the x-axis corresponds to the starting noise sa , while sb=1.0 and γ=1.5 . For MF-WS, the value corresponds to the constant noise corruption level.
Fig. 6 : Real-world manipulation tasks.
Method
Real Push-T
Real Box-close
Inference latency
FM-8
0.35
0.48
54 ms
FM-RTC
0.35
0.43
79 ms
AMF (Ours)
0.60
0.58
14 ms
TABLE IV : Results on static real-world tasks. Each policy is trained for 600 epochs and evaluated on 20 random initial configurations for each task. For Push-T, the policies receive a score of 1 for successful alignment and 0 otherwise. For Box-close, correctly picking the lid yields 0.5 score, while correctly placing it yields an additional 0.5.
Method
11.5 cm/s
12.5 cm/s
13.6 cm/s
15.3 cm/s
Avg
FM-8
0.70
0.70
0.20
0.30
0.48
FM-RTC
0.50
0.50
0.10
0.50
0.40
AMF (Ours)
0.80
0.50
0.30
0.60
0.55
TABLE V : Success rate on conveyor belt picking task. Each policy is trained for 600 epochs. We conduct experiments at 4 different belt speeds, corresponding to DC belt motor actuation with 10-13V. We perform 10 trials for each method at each speed setting, with slight variations in the initial block position on the belt.
Fig. 7 : Rollout comparison on real-world conveyor belt task, with the maximum belt speed of 15.3 cm/s. Top: FM-8 initially misjudges the trajectory needed to intercept the block, and is unable to recover, eventually ending up behind the object and failing the task. Middle: FM with real-time chunking adjusts more aggressively, but closes the gripper prematurely and misses the object. Bottom: AMF tracks slightly behind initially, but performs the necessary adjustments to catch up and successfully grasp the block.
Learning long-horizon robotic manipulation requires jointly achieving expressive behavior modeling, real-time inference, and stable execution, which remains challenging for existing generative policies. Diffusion-based approaches offer strong modeling capacity but incur high inference latency, while flow matching enables fast, near-single-step generation yet often suffers from unstable execution when operating directly in the raw action space. We propose Continuous Latent Action Flow Policy (CoLA-Flow Policy), a trajectory-level imitation learning framework that performs flow matching in a continuous latent action space. By encoding action sequences into temporally coherent latent trajectories and learning an explicit latent-space flow, CoLA-Flow Policy decouples global motion structure from low-level control noise, enabling smooth and reliable long-horizon execution. The framework further integrates geometry-aware point cloud conditioning and execution-time multimodal modulation, using visual cues as a representative modality to enhance real-world robustness. Experiments in simulation and on real robots show that CoLA-Flow Policy achieves near-single-step inference, improves trajectory smoothness by up to 93.7% and task success by up to 25 percentage points over raw action-space flow baselines, while remaining significantly faster than diffusion-based policies.
Wu Songwei, Jiang Zhiduo, Sun Wandong +4
Harbin Institute of Technology, Harbin 150001, China · Honor Device Co., Ltd., Shenzhen, China
In this paper, we study how to achieve one-step action generation in Robotic Foundation Models (RFMs), aiming to overcome the high inference latency of multi-step flow matching. MeanFlow provides a promising framework for this goal, yet its direct application leads to performance collapse. We discover that this stems from two distinctive dynamics exhibited in the RFM velocity field: (1) the ``local acceleration" exhibits stability early on, but surges sharply towards the end of the denoising process, and (2) the spread of its magnitudes across samples widens as denoising progresses. To address these issues, we introduce Kinematic MeanFlow (K-MF), a novel one-step action policy tailored for RFMs. Specifically, grounded in a kinematic identity, K-MF decouples the time derivative term in the MeanFlow formulation into two sub-interval terms separated by an intermediate point. This decoupled formulation enables the two terms to capture early-stage and late-stage denoising dynamics, respectively, while mitigating the error amplification across the process. As a result, our K-MF empowers RFMs to achieve one-step action generation in both training from scratch and fine-tuning paradigms across diverse tasks, while outperforming multi-step flow matching in most settings. In terms of inference efficiency, K-MF reduces action-head latency of GR00T-N1.6 by 67.5%~74.4% across L40 and Jetson Orin in eager and compiled modes, yielding end-to-end latency reductions of 30.3%~54.9%. Code will be available at https://github.com/IntelChina-AI/K-MF.
Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot's action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.
S. Talha Bukhari, Austin Garrett, Yi Wei +3
Department of Computer Science, Purdue University, West Lafayette, IN 47907, USA