cs.ROSep 30, 2026

Cue the Flow: Steering Flow-Matching Policies for Open-World Delivery Manipulation

Authors: Haoxuan Wang, Griffin Galimi, Junhua Huang, Selina Song, Wayne Wu, Yan Yan, Bolei Zhou

Organizations: University of Illinois Chicago · University of California, Los Angeles

Abstract

Open-world goods delivery requires mobile manipulators to follow free-form user instructions and manipulate potentially novel objects. Existing dual-system approaches use high-level grounding models to convert language into grounded visual prompts, but their low-level controllers can remain brittle under noisy perception, dynamic scenes, and contact-rich interactions. We instead use a pretrained flow-matching vision-language-action model as the low-level control interface, leveraging its reactivity and robustness to environmental changes while treating the grounding output as a spatial cue for policy steering. Our key insight is that the pretrained VLA already provides a strong manipulation prior, while the spatial cue supplies the missing target information needed to guide actions under novel language--object mappings. Concretely, we introduce a lightweight cue-conditioned adapter. The adapter is first trained with contrastive objectives to produce salient and spatially discriminative cue representations, and is then supervised to predict a diagonal affine transformation over the generated action chunk, aligning policy steering with the cued target. Across tabletop and mobile-base settings, our method improves instruction following and manipulation success on both in-domain and out-of-domain objects, achieving up to near 2×2\times improvement in average task success rate with negligible inference overhead.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Guided Action Flow: Value-Guided Sampling for Frozen Vision-Language-Action Policies

    Jul 2, 2026Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng +8Flow-Matching Vision-Language-ActionGuidance

  2. World Pilot: Steering Vision-Language-Action Models with World-Action Priors

    Jun 10, 2026Zefu Lin, Rongxu Cui, Junjia Xu +4Efficient World-Action ModelWorld Models

  3. Flow Control: Steering Vision-Language-Action Models with Simple Real-Time Inputs

    Jun 8, 2026Jonathan C. Kao, Jason Chan, Andy WangFlow-Matching Vision-Language-ActionVision-Language-Action Framework