Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences · State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences · Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences · Computer Science and Engineering, the Faculty of Innovation Engineering, Macau University of Science and Technology
Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.
Figures & tables
Figure 1: Viewer-centered versus object-centric representations. (a) Existing visuomotor policies rely on entangled viewer-centered representations and historical motion priors, leading to failures under spatial and visual perturbations. (b) Object-centric representations decouple semantic (“what”) features and image-plane spatial (“where”) cues for more robust visuomotor reasoning.
Figure 2: SlotFlow Architecture. The proposed CFM decouples observations into semantic (“what”) features and lightweight image-plane spatial (“where”) cues.
Figure 3: CFM Overview. (a) Object-centric slot extraction with foveated feature decoding. (b) Coarse-to-fine refinement of the spatial anchor Pglobal and semantic representation SlotFfine through localized high-resolution crops.
Method
Steps
Level 0
Level 1
Level 2
Level 3
Pos Pert.
DDPM-DiT [ 15 ]
100
86%
0%
0%
0%
56%
FM-DiT [ 24 ]
10
98%
2%
2%
2%
60%
Score-Unet [ 32 ]
100
84%
2%
0%
0%
58%
VITA [ 11 ]
6
100%
4%
2%
4%
74%
A2A [ 17 ]
6
100%
12%
18%
10%
70%
SlotFlow (Ours)
6
100%
50%
52%
52%
80%
Table 1: Performance on Close Box. Success rates under visual, kinematic, and spatial perturbations. All models are trained for 100 epochs on 99 demonstrations.
Figure 4: Qualitative Results. Top: Simulation trajectories under Level 3 perturbations and spatial displacements. Bottom: Real-world UR3 deployment under tablecloth replacement and object position perturbations.
Method
ID
Tablecloth Shift
Pos Shift
A2A [ 17 ]
100%
55%
40%
SlotFlow (Ours)
100%
85%
75%
Table 2: Real-World Close Drawer Performance. Success rates under visual and spatial perturbations on a physical UR3 manipulation setup. All results are averaged over 20 real-world evaluation trials.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Steps
Latency (ms)
VITA
6
4.40
A2A
6
5.23
SlotFlow (Ours)
6
11.03
FM-DiT
10
21.95
DDPM-DiT
100
108.09
Score-Unet
100
116.85
Appendix
Table A1: Inference latency comparison. Average policy inference time measured on the Close Box task.
Method
Level 1
Level 2
Level 3
Pos. Pert.
A2A
12%
18%
10%
70%
SlotFlow + Gaussian init.
36%
32%
24%
76%
Mask-pooled DINO
4%
2%
4%
54%
Spatial-only
22%
22%
18%
64%
No-crop SlotFlow
26%
24%
30%
74%
Full SlotFlow
50%
52%
52%
80%
Appendix
Table A2: Controlled ablations on Close Box. Levels 1–3 correspond to the visual and kinematic perturbations defined in the main paper.
Method
5 cm
10 cm
15 cm
A2A
94%
80%
44%
SlotFlow
96%
84%
56%
Improvement
+2%
+4%
+12%
Appendix
Table A3: Recovery after mid-rollout object displacement. The object is moved after the action-history buffer has been accumulated, leaving historical actions unchanged.
Method
Steps
Level 0
Level 1
Level 2
Level 3
Pos Pert.
DDPM-DiT [ 15 ]
100
68%
2%
0%
0%
32%
FM-DiT [ 24 ]
10
74%
0%
0%
0%
32%
Score-Unet [ 32 ]
100
68%
0%
0%
0%
28%
VITA [ 11 ]
6
76%
0%
0%
0%
24%
A2A [ 17 ]
6
64%
12%
16%
16%
26%
SlotFlow (Ours)
6
76%
20%
26%
24%
38%
Appendix
Table A4: Performance on Pick Cube. Success rates across visual, kinematic, and spatial perturbation settings. All policies are trained for 100 epochs under identical training configurations.
Method
Steps
Level 0
Level 1
Level 2
Level 3
Pos Pert.
DDPM-DiT [ 15 ]
100
90%
24%
26%
22%
82%
FM-DiT [ 24 ]
10
92%
0%
4%
4%
76%
Score-Unet [ 32 ]
100
96%
24%
16%
24%
80%
VITA [ 11 ]
6
92%
18%
10%
6%
80%
A2A [ 17 ]
6
92%
72%
72%
74%
80%
SlotFlow (Ours)
6
94%
72%
74%
74%
86%
Appendix
Table A5: Performance on Push Cube. Success rates across visual, kinematic, and spatial perturbation settings. All policies are trained for 100 epochs under identical training configurations.
Developing efficient and accurate visuomotor policies poses a central challenge in robotic imitation learning. While recent rectified flow approaches have advanced visuomotor policy learning, they suffer from a key limitation: After iterative distillation, generated actions may deviate from the ground-truth actions corresponding to the current visual observation, leading to accumulated error as the reflow process repeats and unstable task execution. We present Selective Flow Alignment (SeFA), an efficient and accurate visuomotor policy learning framework. SeFA resolves this challenge by a selective flow alignment strategy, which leverages expert demonstrations to selectively correct generated actions and restore consistency with observations, while preserving multimodality. This design introduces a consistency correction mechanism that ensures generated actions remain observation-aligned without sacrificing the efficiency of one-step flow inference. Extensive experiments across both simulated and real-world manipulation tasks show that SeFA Policy surpasses state-of-the-art diffusion-based and flow-based policies, achieving superior accuracy and robustness while reducing inference latency by over 98%. By unifying rectified flow efficiency with observation-consistent action generation, SeFA provides a scalable and dependable solution for real-time visuomotor policy learning. Code is available on https://github.com/RongXueZoe/SeFA.
Flow matching policies learn continuous velocity fields that transport noise to actions, enabling fast deterministic inference for robot manipulation. However, standard training optimizes a pointwise velocity objective while inference requires numerical integration of that field -- a mismatch that causes compounding trajectory errors. We propose four complementary remedies: (1) auxiliary rectified flow velocity regression that provides uniform temporal supervision across the full time interval; (2) multi-step trajectory consistency training that supervises the integrated displacement of the velocity field over trajectory segments, directly closing the train-inference gap; (3) velocity field regularization that enforces temporal smoothness, preventing oscillations that destabilize integration; and (4) fourth-order Runge-Kutta (RK4) inference that reduces global discretization error by orders of magnitude over Euler methods. Critically, these components are not independently sufficient -- RK4 without a smooth velocity field fails, and smoothness without trajectory-level supervision still drifts, as our ablation study confirms. We further pair these with a dual-view 3D point cloud encoder using two independent PointNet encoders for complementary spatial perception. On four real-robot tasks across a Franka arm and a Boston Dynamics Spot, our method achieves 70% and 60% overall success on two long-horizon multi-phase tasks where both baselines score 0%, and reaches 100% on precision tool placement. Three MetaWorld simulation tasks confirm consistent improvements, validating that trajectory-level supervision is essential for reliable policy execution.
Visual signals play a crucial role in policy learning by enabling models to capture object motion and interaction dynamics. Just as humans reason about actions using both past experience and anticipated outcomes, effective policies should integrate past interactions with future predictions. However, existing visuomotor policies typically model either historical context or future dynamics in isolation, lacking a unified temporal representation of interaction dynamics. In this work, we introduce ChronoFlow, a temporally unified representation that captures past, current, and future interaction dynamics through sparse 3D keypoints of both objects and the gripper. Based on this representation, we propose ChronoFlow-Policy, a diffusion-based visuomotor policy that jointly learns ChronoFlow and action sequences through a co-training objective. Experiments on 14 simulated tasks and 5 real-world manipulation tasks demonstrate that ChronoFlow-Policy consistently outperforms strong diffusion-policy baselines and improves robustness in long-horizon and non-Markovian manipulation scenarios. Our project page is available at https://the-kamisato-sii.github.io/ChronoFlow-Policy-project-page/.
Bokai Lin, Yifu Xu, Xinyu Zhan +6
Shanghai Jiao Tong University, China · Shanghai Innovation Institute, China · Noematrix, China