End-to-end robot policies trained through imitation learning remain constrained by limited data diversity, making reliable zero-shot deployment in real-world settings challenging. Shared-autonomy methods enable human correction through teleoperation, but specialized hardware and operator training hinder deployment at scale. Other approaches incorporate human guidance as additional policy inputs, often requiring architectural changes and dedicated training for steerability, which limits their applicability across policies. We present RoboPrompt, a general-purpose, lightweight robot policy steering system that enables users to guide policy behavior through intuitive, sparse inputs, including drawn traces, target points, and coarse directional instructions. RoboPrompt decouples human-intention translation from the underlying policy: a reusable module converts human guidance into action drafts, which are refined through the diffusion or flow-matching dynamics of the base policy. By controlling action generation in noise space, RoboPrompt balances human intent with the policy prior without modifying the base policy architecture or fine-tuning it for steerability. Experiments demonstrate effective steering across Diffusion Policy, π0.5, and FastWAM. We further use steered rollouts for online policy improvement through DAgger. After 2-3 rounds of iteration, average success rates increase by 15.5% for π0.5 across three tasks and by 21.3% across three policies(Diffusion Policy, π0.5, FastWAM) on the Insert Bread task, while average human intervention counts decrease by 44.0% (2.86 to 1.60) and 81.9% (2.60 to 0.47), respectively.
Figures & tables
Fig. 1 : Overview of RoboPrompt . RoboPrompt aims to enable human-in-the-loop deployment of robot policies, collecting failure-recovery rollout data that can drive continual policy improvement. It supports intuitive multimodal steering through inputs, including points, traces, and coarse directional instructions. These human-guided rollouts are collected and combined with offline data for co-training, forming a closed-loop framework for task completion and self-improvement without teleoperation.
Fig. 2 : System Architecture. RoboPrompt incorporates human intention into policy inference through two phases. In Phase I, prompt inputs are processed by a lightweight VLA to produce a temporally continuous 6D action chunk as an action draft. In Phase II, the base policy refines the action draft generated by Phase I. The fraction of noise-adding steps, γ , controls the balance between human guidance and the base policy. γ increases with successive replanning steps, gradually shifting control toward the base policy.
Fig. 3 : Prompt label generation. Over the next n frames, connected TCP projections and the final TCP projection provide ground truth for trajectory and point prompts; summed and normalized actions supervise the global action prompt.
Success Rate Improvement
Steering Quality
Policy
Task
Progress without Steering
Progress with Steering
Avg. Steering Counts
Avg. visual input
Avg. global input
Avg. Prompt Alignment (%)
π0.5
Left
59.6%
77.5%
2.50
1.45
1.90
99.44%
Right
72.4%
2.95
2.50
2.50
Pot
89.8%
2.35
1.80
1.25
DP
Left
46.5%
71.3%
4.80
3.10
3.40
97.1%
Right
64.75%
3.45
1.85
1.95
TABLE I : Steering performance of RoboPrompt across different policies and tasks. The left block compares task progress with and without steering; the right block reports steering quality (the lower the steering counts, the less human intervention is needed; the higher the prompt alignment, the more faithfully the steered behavior follows the human prompt).
Fig. 4 : Steering experiment setup and natural policy distribution. In the steering experiment, the robot picks up bread from the bowl and places it into one of three targets (pot, left toaster, right toaster). The training set contains 46 demonstrations per target, all using the same text description across the three targets. The right panel shows the natural distribution of the three base policies over the three targets before steering.
Fig. 5 : DAgger experiment setup. (a) Insert Bread : pick up bread from a bowl and insert it into the toaster’s right slot; Hang Cup : hang a cup by its handle on the second-highest peg; Push Ball : push a polyhedral ball across an uneven platform until it touches the flag. (b) Push Ball training dataset contains 5 different layouts, and is evaluated on another new layout. Steered rollouts enable task completion in the new layout and provide data improve the policy over successive DAgger rounds.
Fig. 6 : Main DAgger results with π0.5 across three real-world tasks. Round 0 denotes the original policy before DAgger fine-tuning. Across the Insert Bread, Hang Cup, and Push Ball tasks, iterative training with RoboPrompt-guided corrections improves success rate, while the required steering counts are reduced or remain comparable.
Fig. 7 : DAgger performance after replacing the low-level policy. Round 0 denotes the original base policy. With additional DAgger iterations, task progress increases while the average steering counts decrease across DP, π0.5 , and FastWAM.
Fig. 8 : Denoising dynamics ablation. Original flow-matching and diffusion-based policies exhibit non-uniform action changes over denoising time. FRS keeps the output closer to the Phase I action, while NA-FRS produces a smoother, approximately linear transition across denoising levels.
Fig. 9 : Task progress and steering counts for first-time and experienced users on bread insertion , using the same policy checkpoint and environment. In baseline, the policy executes autonomously without human steering.
Fig. 10 : Policy improvement with filtered (w/ TOT) and unfiltered (w/o TOT) rollout data. Each setting uses 20 successful trajectories for one DAgger round. Baselines show task progress and average steering counts before training.
Generalist robot policies carry broad manipulation priors from large-scale data, but specializing them to a new task remains the deployment bottleneck. This requires eliciting task-specific behavior from limited demonstrations without degrading their broad capabilities. We introduce Proxy Policy Steering (PPS), an inference-time adaptation method that resolves this challenge by training two lightweight proxy policies whose calibrated velocity-space difference steers the frozen base sampler. A reference proxy models the frozen base's behavior on target-task observations, and a task proxy, initialized from the reference, captures how this behavior changes under task supervision. Their difference forms a calibrated velocity-space residual that steers the frozen base sampler at every denoising step. We identify the conditions under which this residual isolates the change induced by task supervision, and validate them empirically. Because the base is never directly modified, its broad capabilities remain available at inference, including behaviors such as recovery from failure that the demonstrations themselves do not exercise. Adaptation requires only forward velocity predictions from the base, making PPS lightweight to train and applicable even without access to the base's parameters. On 8 real-world and 4 simulation manipulation tasks, PPS lifts the state-of-the-art pi 0.5 base policy by 53% absolute success rate on average, with zero-to-one gains on tasks the base never solves, while preserving the base's broad capabilities. PPS outperforms LoRA fine-tuning, from-scratch specialists, residual policies, and prior inference-time steering methods.
Research on robotic manipulation has developed a diverse set of policy paradigms, including vision-language-action (VLA) models, vision-action (VA) policies, and code-based compositional approaches. Concrete policies typically attain high success rates on specific task distributions, but limited generalization beyond it. Rather than proposing another monolithic policy, we propose to leverage the complementary strengths of existing approaches through intelligent policy routing. We introduce RoboRouter, a training-free framework that maintains a pool of heterogeneous policies and learns to select the best-performing policy for each task through accumulated execution experience. Given a new task, RoboRouter constructs a semantic task representation, retrieves historical records of similar tasks, predicts the optimal policy choice without requiring trial-and-error, and incorporates structured feedback to refine subsequent routing decisions. Integrating a new policy into the system requires only a lightweight evaluation and does not incur training overhead. Across simulation benchmark and real-world evaluations, RoboRouter consistently outperforms individual policies, improving the average success rate by more than 3% in simulation and 13% in real-world settings, while preserving execution efficiency. Our results demonstrate that intelligent routing across heterogeneous, off-the-shelf policies provides a practical and scalable pathway toward building more capable robotic systems.
Yiteng Chen, Zhe Cao, Hongjia Ren +9
South China University of Technology · Nanjing University · University of New South Wales +5
Generalist policies can learn a wide range of skills from diverse robot datasets. In order to solve or improve on challenging new tasks, we need a way to infer and invoke the appropriate actions from the policy's rich behavioral prior, especially when directly commanding the policy fails. We focus on flow matching generalists and propose Flow Reversal Steering (FRS): a method that takes suboptimal but ``reasonable'' actions, finds their latent noises by passing them through the flow policy in reverse, and maps them to nearby generalist action modes. We evaluate FRS across many simulated and real-world manipulation settings. First, FRS can turn coarse semantic guidance from humans or vision-language models (VLMs) into corresponding good robot actions, improving zero-shot control. These gains can be distilled with behavioral cloning by training an auxiliary policy to output noises that the generalist maps to good actions -- showing up to 95% absolute task success rate boosts in under a minute of training. Finally, FRS enables policy improvement by bootstrapping reinforcement learning with semantic knowledge, improving on several tasks that standard RL fails to improve on.