Visuomotor policies learn a direct map from raw sensory observations to robot action sequences. Policies based on Diffusion and Flow Matching capture the multimodal distribution over action sequences in an end-to-end manner. This expressivity comes at the cost of multi-step numerical integration of the learned vector field for action generation, which can be expensive and time-consuming, impeding fast control rates required in robotics applications. Furthermore, robot action sequences are usually defined on a smooth, differentiable manifold, requiring that the learned policy respects the intrinsic geometry of the robot's action space. Here, we present Riemannian MeanFlow Policy (RMFP), which learns the conditioned flow map of the probability path on the robot action manifold. Our formulation employs a flow map consistency objective grounded in the data by a Riemannian Conditional Flow Matching anchor. The flow map consistency condition is stable to train and constrains the learned model to finite-time transport, which yields on-manifold action sequence generation with as few as one network function evaluation. We present results on the spherical LASA and Push-T benchmarks, on the Tool Hang and Transport tasks of the Robomimic suite, and on the Franka Kitchen task with manifold-constrained action generation, and demonstrate that RMFP attains performance competitive with prior work at a lower sampling cost. We also employ RMFP on a real-world robotic manipulation task to demonstrate fast action generation under imperfect sensor measurements in the physical world.
Figures & tables
Figure 1 : We present Riemannian MeanFlow Policy (RMFP) for fast action sequence generation on manifolds. Here we show 1-NFE trajectory generation for the LASA shape W projected onto S2 [ 8 ] . Demonstrations from the dataset are shown in black and the executed trajectory is colorized by time-parameterization. ⋆ denotes the north pole. As shown, RMFP traces a smoother path while abiding by the demonstrations. Median DTW is 0.023 for RMFP against 0.079 for RFMP.
Figure 2 : The Riemannian flowmap consistency objective ( 9 ) and Riemannian conditional flow matching anchor ( 10 ) are schematically shown. The horizontal axis represents time. A point as is drawn on the conditional geodesic (dashed) from a prior sample a0 to a demonstrated chunk a1 . The two flow maps through the intermediate time r are rolled out without gradient (grey arcs), and their endpoint a~t defines the chord target u^ of ( 6 ). Tangent vectors at as are denoted by straight arrows: the frozen chord (t−s)u^=Logas(a~t) ( dashed grey ), the single jump (t−s)uθ(as,s,t) ( teal ) that the gradient reaches, and the geodesic velocity a˙s ( gold ) against which uθ(as,s,s) is anchored. Lsemi penalizes the residual between the jump and the chord.
Figure 3 : Simulated environments of the evaluation benchmarks. (a) Spherical Push-T: the Push-T canvas under the stereographic map onto S2 , with the agent (blue), the block (grey-blue) and the goal pose (green). (b) Robomimic Tool Hang: a single arm inserts a hook into a base and hangs a wrench on it, from a state observation. (c) Robomimic Transport: two arms transfer a hammer from a covered container to a target bin on the opposite shelf, from camera observations. (d) Franka Kitchen: the policy commands an end-effector pose on R3×S3×R2 , and an episode counts as a success when at least four of the seven subtasks are completed.
Method
NFE 1
NFE 2
NFE 5
NFE 10
DTW
Jerk
DTW
Jerk
DTW
Jerk
DTW
Jerk
RFM
0.090
857
0.020
153
0.016
83
0.016
59
RMF- v
0.033
317
0.017
96
0.019
86
0.020
82
RMF- x1
0.022
153
0.022
146
0.026
143
0.024
113
Table I : LASA on S2 , ten characters. Median DTW to the demonstration, and the executed path’s jerk (the demonstration equals 1.0 ); lower is better. ( λsemi=5 )
Method
Coverage Score (%)↑
1 NFE
2 NFE
5 NFE
10 NFE
RDP- ϵ
15.0
51.8
83.9
81.6
RDP- x
73.1
74.0
76.0
74.4
RFM
67.9
77.2
84.7
83.9
RMF- v
78.5
81.0
85.7
84.1
RMF- x1
76.0
78.3
81.7
85.0
Table II : Spherical Push-T coverage score against the sampling budget
Method
Success Rate (%)↑
1 NFE
2 NFE
5 NFE
10 NFE
RDP- ϵ
3.7
19.7
36.9
47.1
RDP- x
23.7
41.7
43.1
39.1
RFM-uni
14.9
40.3
44.3
41.1
RFM-hemi
20.6
41.7
47.1
48.6
RMF- v -uni
34.9
41.7
44.3
45.1
Table III : Franka Kitchen success rate against the sampling budget
Method
Success Rate (%)↑
1 NFE
2 NFE
5 NFE
10 NFE
DP
0.0
46.0
64.0
74.0
RFM
60.0
60.0
60.0
64.0
RMF- v
64.0
76.0
72.0
74.0
RMF- x1
48.0
52.0
60.0
74.0
Table IV : Robomimic Tool Hang success rate against the sampling budget
Method
Success Rate (%)↑
1 NFE
2 NFE
5 NFE
10 NFE
DP
0.0
82.0
96.0
82.0
RFM
94.0
86.0
88.0
86.0
RMF- v
90.0
92.0
92.0
92.0
RMF- x1
94.0
90.0
92.0
90.0
Table V : Robomimic Transport success rate against the sampling budget
Figure 4 : Riemannian MeanFlow Policy executed in a real-world manipulation task. The policy is trained on 60 real-world demonstrations and deployed on the YAM robotic platform. Representative stages of the rollout show the robot approaching and grasping the object, transporting it to a goal state, and placing it into the target container.
Flow matching policies learn continuous velocity fields that transport noise to actions, enabling fast deterministic inference for robot manipulation. However, standard training optimizes a pointwise velocity objective while inference requires numerical integration of that field -- a mismatch that causes compounding trajectory errors. We propose four complementary remedies: (1) auxiliary rectified flow velocity regression that provides uniform temporal supervision across the full time interval; (2) multi-step trajectory consistency training that supervises the integrated displacement of the velocity field over trajectory segments, directly closing the train-inference gap; (3) velocity field regularization that enforces temporal smoothness, preventing oscillations that destabilize integration; and (4) fourth-order Runge-Kutta (RK4) inference that reduces global discretization error by orders of magnitude over Euler methods. Critically, these components are not independently sufficient -- RK4 without a smooth velocity field fails, and smoothness without trajectory-level supervision still drifts, as our ablation study confirms. We further pair these with a dual-view 3D point cloud encoder using two independent PointNet encoders for complementary spatial perception. On four real-robot tasks across a Franka arm and a Boston Dynamics Spot, our method achieves 70% and 60% overall success on two long-horizon multi-phase tasks where both baselines score 0%, and reaches 100% on precision tool placement. Three MetaWorld simulation tasks confirm consistent improvements, validating that trajectory-level supervision is essential for reliable policy execution.
Flow matching has recently become a new standard for behavior cloning in robotic manipulation. However, state-of-the-art flow matching policies suffer from a systematic structural mismatch: they rely on a globally fixed isotropic source distribution despite the strongly fragmented and heteroscedastic structure of robotic action spaces. This agnostic initialization forces the model to learn highly entangled vector fields, bottlenecking training efficiency and limiting overall policy performance. To address this limitation, we introduce Latent Action Guided Flow Matching (LAFM), a novel framework that replaces the monolithic Gaussian with an adaptive library of learned prior distributions. By grounding these distributions using a latent action model, LAFM maps current observations to discrete motion primitives, selecting a specialized base distribution that provides an informed, structurally aligned initialization for the denoising process. This dynamic adaptivity naturally accommodates heteroscedasticity in human demonstrations and makes transport trajectories shorter and less entangled. Empirically, LAFM substantially outperforms standard flow matching formulations, increasing task success rates by 23.4% in real-world robotic deployments and by 10.4% on the LIBERO-90 benchmark. Furthermore, we demonstrate that LAFM achieves state-of-the-art results, surpassing massively pre-trained vision-language-action models while utilizing significantly smaller architectures.
Bruno Machado, Alexandre Chapin, Emmanuel Dellandrea +1
Generative models such as diffusion and flow matching have become dominant paradigms for visuomotor policy learning, yet their reliance on iterative denoising incurs high inference latency incompatible with real-time robotic control. We present Fast Legendre-polynomial Action policy via Sparse History-anchored flow (FLASH Policy), which replaces discrete action-chunk generation with continuous Legendre polynomial trajectory representation. Specifically, by fitting expert demonstrations under sparse temporal sampling, FLASH enables a single inference to cover a significantly extended action horizon. To further accelerate generation, FLASH initiates the flow matching process from history polynomial coefficients rather than uninformative Gaussian noise, shortening the transport distance and enabling accurate single-step inference. Moreover, analytic polynomial differentiation directly provides desired velocity feed-forward signals to the torque controller without numerical approximation. Extensive experiments on five simulated and two real-world manipulation tasks demonstrate that FLASH achieves state-of-the-art success rates (≥92% across all tasks), a per-episode inference time of 31.40ms (up to 175× faster than diffusion policies and 18× faster than prior flow matching policies), up to 4× faster training convergence than ACT, and 5× to 7× reduction in controller tracking error compared to discrete-action baselines.
Jiaqi Bai, Jindou Jia, Yuxuan Hu +5
MARS Lab, Nanyang Technological University, Singapore