cs.ROSep 21, 2026

Object-Centric Conditioning for Visuomotor Flow Matching

Authors: Jijie Li, Xu Yang, Junhong Zou, Chunhai Zhao, Chaoyang Zhao, Zhen Lei, Xiangyu Zhu

Organizations: School of Artificial Intelligence, University of Chinese Academy of Sciences · State Key Laboratory of Multimodal Artificial Intelligence Systems, Institute of Automation, Chinese Academy of Sciences · Foundation Model Research Center, Institute of Automation, Chinese Academy of Sciences · Centre for Artificial Intelligence and Robotics, Hong Kong Institute of Science & Innovation, Chinese Academy of Sciences · Computer Science and Engineering, the Faculty of Innovation Engineering, Macau University of Science and Technology

Abstract

Robot visuomotor policies are commonly formulated as autoregressive, diffusion-based, or more recently, flow matching models. Among them, Action-to-Action (A2A) flow matching improves inference efficiency by initializing generation from historical action priors rather than stochastic noise. However, stale historical motion patterns and entangled global visual representations can jointly reduce robustness under spatial out-of-distribution (OOD) shifts and visual distractors. In this work, we propose SlotFlow, an object-centric flow matching policy for robust visuomotor manipulation. SlotFlow decouples scene observations into semantic ("what") features and lightweight image-plane spatial ("where") cues to provide object-aware policy conditioning and current-state grounding. The semantic representation suppresses irrelevant background correlations, while the spatial cue improves adaptation to shifted object configurations. Extensive simulation and real-world experiments demonstrate improved robustness under visual distractors and severe spatial perturbations while preserving the low-step inference efficiency of A2A. Controlled initialization and perception ablations further identify object-centric grounding as a major source of the gains and show that it complements, rather than replaces, useful historical motion priors.

Figures & tables

Appendix figures & tables5 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SeFA-Policy: Fast and Accurate Visuomotor Policy Learning with Selective Flow Alignment

    Nov 11, 2025Rong Xue, Jiageng Mao, Mingtong Zhang +1Visuomotor PolicyFlow Policies

  2. Trajectory-Consistent Flow Matching for Robust Visuomotor Policy Learning

    May 8, 2026Riad Ahmed, Sujosh Nag, Moniruzzaman Akash +2Visuomotor PolicyFlow-Matching Vision-Language-Action

  3. ChronoFlow-Policy: Unifying Past-Current-Future Interaction Flow in Visuomotor Policy Learning

    Jun 30, 2026Bokai Lin, Yifu Xu, Xinyu Zhan +6Visuomotor PolicyInteraction Data