Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences · State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences · Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences · Computer Science and Engineering, the Faculty of Innovation Engineering, Macau University of Science and Technology
Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.
Figures & tables
Figure 1: Viewer-centered versus object-centric representations. (a) Existing visuomotor policies rely on entangled viewer-centered representations and historical motion priors, leading to failures under spatial and visual perturbations. (b) Object-centric representations decouple semantic (“what”) features and image-plane spatial (“where”) cues for more robust visuomotor reasoning.
Figure 2: SlotFlow Architecture. The proposed CFM decouples observations into semantic (“what”) features and lightweight image-plane spatial (“where”) cues.
Figure 3: CFM Overview. (a) Object-centric slot extraction with foveated feature decoding. (b) Coarse-to-fine refinement of the spatial anchor Pglobal and semantic representation SlotFfine through localized high-resolution crops.
Method
Steps
Level 0
Level 1
Level 2
Level 3
Pos Pert.
DDPM-DiT [ 15 ]
100
86%
0%
0%
0%
56%
FM-DiT [ 24 ]
10
98%
2%
2%
2%
60%
Score-Unet [ 32 ]
100
84%
2%
0%
0%
58%
VITA [ 11 ]
6
100%
4%
2%
4%
74%
A2A [ 17 ]
6
100%
12%
18%
10%
70%
SlotFlow (Ours)
6
100%
50%
52%
52%
80%
Table 1: Performance on Close Box. Success rates under visual, kinematic, and spatial perturbations. All models are trained for 100 epochs on 99 demonstrations.
Figure 4: Qualitative Results. Top: Simulation trajectories under Level 3 perturbations and spatial displacements. Bottom: Real-world UR3 deployment under tablecloth replacement and object position perturbations.
Method
ID
Tablecloth Shift
Pos Shift
A2A [ 17 ]
100%
55%
40%
SlotFlow (Ours)
100%
85%
75%
Table 2: Real-World Close Drawer Performance. Success rates under visual and spatial perturbations on a physical UR3 manipulation setup. All results are averaged over 20 real-world evaluation trials.
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Steps
Latency (ms)
VITA
6
4.40
A2A
6
5.23
SlotFlow (Ours)
6
11.03
FM-DiT
10
21.95
DDPM-DiT
100
108.09
Score-Unet
100
116.85
Appendix
Table A1: Inference latency comparison. Average policy inference time measured on the Close Box task.
Method
Level 1
Level 2
Level 3
Pos. Pert.
A2A
12%
18%
10%
70%
SlotFlow + Gaussian init.
36%
32%
24%
76%
Mask-pooled DINO
4%
2%
4%
54%
Spatial-only
22%
22%
18%
64%
No-crop SlotFlow
26%
24%
30%
74%
Full SlotFlow
50%
52%
52%
80%
Appendix
Table A2: Controlled ablations on Close Box. Levels 1–3 correspond to the visual and kinematic perturbations defined in the main paper.
Method
5 cm
10 cm
15 cm
A2A
94%
80%
44%
SlotFlow
96%
84%
56%
Improvement
+2%
+4%
+12%
Appendix
Table A3: Recovery after mid-rollout object displacement. The object is moved after the action-history buffer has been accumulated, leaving historical actions unchanged.
Method
Steps
Level 0
Level 1
Level 2
Level 3
Pos Pert.
DDPM-DiT [ 15 ]
100
68%
2%
0%
0%
32%
FM-DiT [ 24 ]
10
74%
0%
0%
0%
32%
Score-Unet [ 32 ]
100
68%
0%
0%
0%
28%
VITA [ 11 ]
6
76%
0%
0%
0%
24%
A2A [ 17 ]
6
64%
12%
16%
16%
26%
SlotFlow (Ours)
6
76%
20%
26%
24%
38%
Appendix
Table A4: Performance on Pick Cube. Success rates across visual, kinematic, and spatial perturbation settings. All policies are trained for 100 epochs under identical training configurations.
Method
Steps
Level 0
Level 1
Level 2
Level 3
Pos Pert.
DDPM-DiT [ 15 ]
100
90%
24%
26%
22%
82%
FM-DiT [ 24 ]
10
92%
0%
4%
4%
76%
Score-Unet [ 32 ]
100
96%
24%
16%
24%
80%
VITA [ 11 ]
6
92%
18%
10%
6%
80%
A2A [ 17 ]
6
92%
72%
72%
74%
80%
SlotFlow (Ours)
6
94%
72%
74%
74%
86%
Appendix
Table A5: Performance on Push Cube. Success rates across visual, kinematic, and spatial perturbation settings. All policies are trained for 100 epochs under identical training configurations.