MotorMind: Scaffolding General Vision Language Models for Zero-Shot Robot Manipulation
Organizations: University of Illinois Urbana-Champaign
Abstract
Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.
Figures & tables
| Method | Action Generation | Extra Components | Execution |
| ( Black et al., 2025a ) | Learned action policy | Trained action model | Action chunks |
| MolmoAct2 ( Fang et al., 2026 ) | Learned action policy | Trained action model | Closed-loop prediction |
| OpenVLA-OFT ( Kim et al., 2025b ) | Trained action model | Fine-tuned action decoder | Action chunks |
| GR00T N1.5 ( NVIDIA, 2025 ) | Trained action model | Diffusion action model | Policy rollout |
| CaP-X ( Fu et al., 2026 ) | Generated robot code | Perception and control tools | Program execution |
| VoLoAgent ( Chen et al., 2026a ) | VLA + robot primitives | Perception models and tools | Tool orchestration |
| Model | Action | Progress | Completion | Overall | Latency |
| Qwen3.8-Flash-Next-FP8 ( Qwen Team, 2026 ) | 36.25 | 55.00 | 65.00 | 52.08 | 276 |
| HY-Embodied-0.5 MoT-2B ( Tencent Robotics X et al., 2026 ) | 20.00 | 50.00 | 47.50 | 39.17 | 402 |
| Hy-Embodied-VLM-1.0 A3B ( Wang et al., 2026 ) | 18.75 | 43.75 | 50.00 | 37.50 | 592 |
| Cosmos3-Nano ( NVIDIA, 2026 ) | 18.75 | 50.00 | 62.50 | 43.75 | 3035 |
| GLM-5.3-Flash ( Z.ai, 2026 ) | 37.50 | 52.50 | 58.75 | 49.58 | 403 |
| GPT-6 Astra ( OpenAI, 2026 ) | 60.00 | 76.25 | 82.50 | 72.92 | 8724 |
| Method | Configuration | Base success rate (%) | Perturbation success rate (%) | Base efficiency | Perturbation efficiency | |||||||||
| Goal | Spatial | Object | Avg. | Semantic | Object | Position | Task | Avg. | Time (s) | Time score | Time (s) | Time score | ||
| Fine-Tuned Policies / Agentic Methods with Fine-Tuned Policies | ||||||||||||||
| Fine-tuned | 95.0 | 100.0 | 100.0 | 98.3 | 98.3 | 95.0 | 36.7 | 23.3 | 63.3 | 5.9 | 997.1 | 6.8 | 562.1 | |
| MolmoAct2 | Fine-tuned | 100.0 | 100.0 | 100.0 | 100.0 | 95.0 | 90.0 | 36.7 | 30.0 | 62.9 | 6.6 | 914.6 | 7.7 | 492.2 |
| OpenVLA / OFT | Fine-tuned | 95.0 | 100.0 | 100.0 | 98.3 | 96.7 | 86.7 | 11.7 | 10.0 | 51.2 | 5.7 | 1033.5 | 6.7 | 459.9 |
| GR00T N1.5 | Fine-tuned | 0.0 | 95.0 | 0.0 | 31.7 | 30.0 | 30.0 | 1.7 | 16.7 | 19.6 | 15.6 | 121.8 | 16.2 | 72.4 |
| Method | Dynamic Reasoning | Scene Shift | Dynamic Manipulation | Prompt Shift |
| (10 tasks) | (10 tasks) | (5 tasks) | (5 tasks) | |
| 0% | 70% | 0% | 20% | |
| GR00T N1.5 | 0% | 10% | 0% | 0% |
| MolmoAct2 | 20% | 50% | 40% | 0% |
| CaP-X | 30% | 20% | 20% | 60% |
| MotorMind | 70% | 90% | 80% | 60% |
| Object | Direct Perception | Human Intervention | Avg. SR | ||
| Bowl | Box | Bowl | Box | ||
| Blue Cube | 100% | 100% | 100% | 100% | 100% |
| Corn | 100% | 100% | 100% | 80% | 95% |
| Battery | 90% | 100% | 80% | 90% | 90% |
| Avg. SR | 97% | 100% | 93% | 90% | 95% |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Decision being diagnosed | Visual input | Count |
| Action selection | Primary primitive to execute after the current observation | Scene and wrist views at | 80 |
| Progress verification | Whether the transition from to advances the stated subgoal | Scene and wrist views at | 80 |
| Subsubgoal completion | Whether the state at satisfies the stated subgoal and criterion | Scene and wrist views at | 80 |
| Model | Translation (48) | Rotation (19) | Gripper (13) |
| Qwen3.8-Flash-Next-FP8 | 47.92% (23/48) | 0.00% (0/19) | 46.15% (6/13) |
| HY-Embodied-0.5 MoT-2B | 20.83% (10/48) | 15.79% (3/19) | 23.08% (3/13) |
| Hy-Embodied-VLM-1.0 A3B | 31.25% (15/48) | 0.00% (0/19) | 0.00% (0/13) |
| Cosmos3-Nano (understanding tower) | 27.08% (13/48) | 0.00% (0/19) | 15.38% (2/13) |
| GLM-5.3-Flash | 41.67% (20/48) | 15.79% (3/19) | 53.85% (7/13) |
| GPT-6 Astra (medium) | 58.33% (28/48) | 63.16% (12/19) | 61.54% (8/13) |
| Condition | Configuration key | Description |
| Base | base | Original, unperturbed task setting |
| Language | lan | Instruction wording perturbations |
| Object | object | Benchmark-defined object perturbations |
| Position Swap | swap | Object placement perturbations |
| Task | task | Target or goal perturbations |
| Baseline | Control mechanism | Planning and memory |
| OpenPI | Direct VLA action prediction | No external planner or task memory |
| OpenVLA | Autoregressive action-token prediction | No external planner or task memory |
| OpenVLA-OFT | Continuous action-chunk prediction | No external planner or task memory |
| GR00T N1.5 | Native action-chunk prediction | No external planner or task memory |
| MolmoAct2 | Continuous action prediction | No external planner or task memory |
| Harness VLA | Tool-based execution with frozen | VLM planning and task-specific memory |
| Suite | Tasks | Evaluation focus |
| Interactive | 10 | Adapt to object or receiver displacement with an unchanged instruction. |
| Pure Dynamic | 5 | Identify, intercept, and place a named moving target. |
| Dynamic Reasoning | 10 | Resolve semantic, relational, and temporal references to moving targets. |
| Prompt Change | 5 | Respond to revised instructions without resetting the physical scene. |
| Task | Instruction | Intervention |
| 1 | put the ketchup in the wooden tray | Receiver during transport; 1 displacement |
| 2 | put the bowl on the plate | Receiver during transport; 1 displacement |
| 3 | put the wine bottle on the rack | Target object before grasping; 1 displacement |
| 4 | put the cream cheese in the bowl | Receiver during transport; 1 displacement |
| 5 | put the bbq sauce in the wooden tray | Target object before grasping; 2 displacements |
| 6 | put the milk in the basket | Receiver during transport; 2 displacements |
| Task | Instruction |
| 1 | pick up the red mug from the conveyor belt and place it in the basket |
| 2 | pick up the alphabet soup from the conveyor belt and place it in the basket |
| 3 | pick up the ketchup from the conveyor belt and place it in the red basket |
| 4 | pick up the blue-and-white cream cheese box from the conveyor belt and place it in the wooden tray |
| 5 | pick up the green ketchup from the conveyor belt and place it in the basket |
| Task | Reasoning requirement | Instruction |
| 1 | Contents: liquid | pick up the packaged product containing liquid from the conveyor belt and place it in the basket |
| 2 | Category: edible item | pick up the edible item from the conveyor belt and place it in the basket |
| 3 | Purpose: beverage | pick up the beverage meant for drinking from the conveyor belt and place it in the basket |
| 4 | Category exclusion: not a bowl | pick up the drinking vessel that is not a bowl from the conveyor belt and place it in the basket |
| 5 | Attribute exclusion: not red | pick up the cup that is not red from the conveyor belt and place it in the basket |
| 6 | Initial spatial relation | pick up the object that is between the two bowls in the initial arrangement on the conveyor belt and place it in the basket |
| Task | Initial instruction | Replacement instruction |
| 1 | put the ketchup in the wooden tray | Change of plan: put the ketchup on the plate instead of in the wooden tray. |
| 2 | Put the chocolate pudding on the plate. | Change of plan: put the chocolate pudding in the wooden tray instead of on the plate. |
| 3 | Put the chocolate pudding on the plate. | First put the chocolate pudding down on the table, turn on the stove and keep it on, and then continue putting the chocolate pudding on the plate. |
| 4 | put the chocolate pudding on the plate | Avoid the obstacle and continue putting the chocolate pudding on the plate. |
| 5 | Put the BBQ sauce in the wooden tray. | Change of plan: leave the BBQ sauce on the table and put only the chocolate pudding in the wooden tray. |
| Budget | Goal | Spatial | Object | Avg. SR | Time (s) | Time score |
| 300 s | 40.0% | 70.0% | 90.0% | 66.7% | 159.8 | 25.03 |
| 450 s | 50.0% | 70.0% | 80.0% | 66.7% | 210.0 | 19.05 |
| 900 s | 40.0% | 60.0% | 50.0% | 50.0% | 320.0 | 9.37 |
| 3,600 s | 50.0% | 60.0% | 70.0% | 60.0% | 296.4 | 12.14 |