Open-world goods delivery requires mobile manipulators to follow free-form user instructions and manipulate potentially novel objects. Existing dual-system approaches use high-level grounding models to convert language into grounded visual prompts, but their low-level controllers can remain brittle under noisy perception, dynamic scenes, and contact-rich interactions. We instead use a pretrained flow-matching vision-language-action model as the low-level control interface, leveraging its reactivity and robustness to environmental changes while treating the grounding output as a spatial cue for policy steering. Our key insight is that the pretrained VLA already provides a strong manipulation prior, while the spatial cue supplies the missing target information needed to guide actions under novel language--object mappings. Concretely, we introduce a lightweight cue-conditioned adapter. The adapter is first trained with contrastive objectives to produce salient and spatially discriminative cue representations, and is then supervised to predict a diagonal affine transformation over the generated action chunk, aligning policy steering with the cued target. Across tabletop and mobile-base settings, our method improves instruction following and manipulation success on both in-domain and out-of-domain objects, achieving up to near 2× improvement in average task success rate with negligible inference overhead.
Figures & tables
Figure 1: (a.1) The policy is trained on in-context demonstrations collected in a fixed environment. (a.2) In real-world delivery, however, the robot must execute manipulation under changing environments and novel language–object mappings. (b) We focus on improving policy generalization to these novel mappings. A grounding module converts the open-ended instruction into a spatial cue, which conditions a contrastive adapter to predict affine transformations over the frozen flow-matching policy’s action output, steering the generated action chunk toward the target object.
Figure 2: Illustration of adapter learning.
Figure 3: Qualitative rollouts under the unseen-object with novel-language setting (Setting III). The leftmost column shows the System 2 spatial cue, where the red heatmap indicates the target object. For each method, two temporal keyframes are shown. The Base policy produces untargeted behavior, while VP-VLA either oscillates or reaches the wrong object. Our method follows the spatial cue and successfully manipulates the specified target in both task examples.
Figure 4: Extension to a complete delivery pipeline. We show example rollouts that combine manual navigation with autonomous manipulation. The red heatmap indicates the correctly grounded target object. After the mobile base is manually navigated to the workspace, our policy performs autonomous cue-conditioned manipulation to complete the task.
Figure 5
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 7: Robot hardware setup.
Evaluation Setting
Prompt Templates
Setting I & II
Pick the <descriptor> bag and put it into the box Grab <descriptor> bag and place it in the box Take the <descriptor> bag and drop it into the box Move the <descriptor> bag into the box Put <descriptor> bag into the box Place the <descriptor> bag in the box
Setting III
Put bag with order number <value> into the box Grab bag with number <value> into the box Pick order number <value> bag and place it into the box
Appendix
Table 3: List of prompts used for evaluation. For seen-object evaluations, <descriptor> refers to objects observed during training, i.e., “black” or “plastic.” For unseen-object evaluations, it refers to novel descriptors such as “brown” or “red.” In the novel-language setting, <value> denotes ordinal identifiers, i.e., 1, 2, or 3, used to evaluate compositional reference grounding under previously unseen instruction patterns.
Figure 8: Representative evaluation configurations across the three experimental settings. Rows correspond to the evaluation regimes described in the main paper: Setting I uses seen objects with paraphrased language, Setting II uses unseen objects with known language, and Setting III uses unseen objects with novel language. Columns show different spatial layouts used during evaluation, varying object placement, distractor arrangement, and scene geometry. In Setting I, Configurations 0–1 correspond to the black-bag task, while Configurations 2–3 correspond to the plastic-bag task. All methods are evaluated on the same prompts and spatial configurations.
Figure 9: Example demonstrations of policy steering across the three evaluation settings. The red heatmap indicates the correctly grounded target object.
Reinforcement learning can improve vision-language-action (VLA) policies beyond supervised fine-tuning, although this typically involves further updates to the policy parameters. For flow-matching policies, iterative action generation provides an additional opportunity to incorporate task information during inference. We introduce Guided Action Flow (GAF), which learns a compact, observation-conditioned action-value critic from robot task rollouts and applies its action gradient to steer reverse-time flow sampling. The supervised-fine-tuned VLA remains frozen throughout critic learning and deployment. Physical-robot experiments show an increase in aggregate success from 60.0% to 82.5% across six nominal manipulation tasks. Under six altered-lighting and object-distractor conditions evaluated on three of these tasks, aggregate success improves from 34.2% to 49.2%. Ablations and rollout analyses support the importance of the learned guidance direction and the critic's visual and proprioceptive inputs. With approximately 2.735M trainable critic parameters alongside a 0.45B-parameter VLA, GAF enables task outcomes to inform action generation through a compact inference-time guidance module.
Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng +8
Department of Computer Science, University College London. · Department of Mechanical Engineering, University College London.
Vision-Language-Action (VLA) models inherit semantic grounding from large-scale pretraining and perform competently across in-distribution manipulation tasks. This grounding, however, is built on static image-text pairs, whereas manipulation is a continuous, contact-rich process whose dynamics this pretraining cannot capture. We present World Pilot, a VLA framework that augments the policy with priors from a World-Action Model (WAM), routed into the decision chain through two complementary pathways. Latent Steering conditions the perception layer on a scene-evolution latent, and Action Steering supplies an anticipated trajectory as a motion prior to the action generator. Together the two priors equip the VLA with an anticipated view of the scene and a trajectory-level motion hint alongside its semantic conditioning, and the scene-evolution prior remains effective even when supplied by a video-pretrained world model that has not been action-post-trained. World Pilot attains a state-of-the-art Total success rate of 84.7% on the LIBERO-Plus zero-shot OOD benchmark and the highest success rate on every real-robot setting across four manipulation tasks, with the largest margins under shifts in viewpoint, geometry, deformable state, and pose. Project Website: https://world-pilot.github.io/
Zefu Lin, Rongxu Cui, Junjia Xu +4
Institute of Automation, Chinese Academy of Sciences (CASIA) · Nanjing University · Beihang University
We introduce flow control of vision-language-action (VLA) models, a simple and effective way to steer VLA actions in real-time through generic inputs, such as a keyboard. This method can be used out-of-the-box and does not require retraining or fine-tuning VLAs. It enables relatively crude user inputs to steer a VLA to align with user intent. The VLA transforms these inputs into action samples drawn from the VLA expert action distribution learned during training, so that the generated actions are high quality (conformity to the action expert distribution) and high fidelity (reflecting the user's intent). We demonstrate that flow control has many desirable properties: (1) flow control accurately and responsively steers robot actions with user inputs, (2) it is robust to suboptimal user inputs, (3) it enables users to steer VLAs to achieve significantly higher success rates and faster task completion, and (4) fine-tuning a VLA on flow control trajectories improves the autonomous policy. Together, these results provide a simple and intuitive way for users to help steer VLA actions, increasing task performance.
Jonathan C. Kao, Jason Chan, Andy Wang
Departments of Electrical and Computer Engineering, Computer Science, and Neurobiology, University of California, Los Angeles · Department of Computer Science, University of California, Los Angeles