Vision-language-action (VLA) models have advanced robotic manipulation, but their zero-shot generalization in new tasks and environments remains limited, and their reliance on specialized training keeps them from benefiting directly from rapidly advancing general-purpose vision-language models (VLMs). In parallel, recent agentic robotic systems leverage VLMs for high-level reasoning or coding agents for robot control, but often depend on extensive external models and tools, introducing additional complexity and cost. This motivates us to ask: Can a general-purpose VLM itself operate a robot more like the human teleoperator by reasoning directly from observations, issuing actions, and continuously adapting to execution feedback, without relying on external models such as learned action experts, coding agents or grounding tools like SAM3? In this work, we introduce MotorMind, a robot manipulation harness that connects VLM-proposed mid-level actions to deterministic robot control and feedback, with asynchronous monitoring and background memory updates. Without task-specific policy training, coding agents, or additional grounding tools such as SAM3, MotorMind achieves 66.7% success on the base LIBERO-PRO suites and 53.8% under perturbations, compared with at most 13.3% and 19.2%, respectively, for the prior zero-shot methods we evaluate. The same interface reaches 95% average success on a real xArm6 robot across direct manipulation and human-perturbation settings. Replacing the backbone with a stronger VLM further improves performance, while the remaining failures - primarily due to visual grounding, embodied reasoning, and action knowledge - decrease as VLM capability improves. These results show that a general-purpose VLM, when equipped with an appropriate mid-level action representation and asynchronous execution harness, can perform effective zero-shot robotic manipulation.
Figures & tables
Figure 1: MotorMind : Equipping general-purpose vision-language models into robot controller. VLM proposes mid-level actions (e.g. "move forward by 12.20mm"), then pass to the control layer that connects those decisions to physical execution and feedback. Asynchronous monitoring enables interruption of pending actions, and outcome verification guides recovery and replanning. This design supports zero-shot manipulation across simulation and real-world settings without task-specific training.
Method
Action Generation
Extra Components
Execution
π0.5 ( Black et al., 2025a )
Learned action policy
Trained action model
Action chunks
MolmoAct2 ( Fang et al., 2026 )
Learned action policy
Trained action model
Closed-loop prediction
OpenVLA-OFT ( Kim et al., 2025b )
Trained action model
Fine-tuned action decoder
Action chunks
GR00T N1.5 ( NVIDIA, 2025 )
Trained action model
Diffusion action model
Policy rollout
CaP-X ( Fu et al., 2026 )
Generated robot code
Perception and control tools
Program execution
VoLoAgent ( Chen et al., 2026a )
VLA + robot primitives
Perception models and tools
Tool orchestration
Table 1: Comparison of robotic manipulation methods. Most existing works rely on a learned action policy, generated robot code, or extra perception and control modules. MotorMind uses the VLM directly for semantic action generation without a coding agent, learned action model, SAM-based perception, or external inverse kinematics (IK) solver. See Appendix A for a detailed discussion of related work.
Figure 2: Diagnosing VLM capabilities for robotic manipulation. We evaluate action selection, progress assessment, and subgoal completion.
Model
Action
Progress
Completion
Overall
Latency ↓
Qwen3.8-Flash-Next-FP8 ( Qwen Team, 2026 )
36.25
55.00
65.00
52.08
276
HY-Embodied-0.5 MoT-2B ( Tencent Robotics X et al., 2026 )
20.00
50.00
47.50
39.17
402
Hy-Embodied-VLM-1.0 A3B ( Wang et al., 2026 )
18.75
43.75
50.00
37.50
592
Cosmos3-Nano ( NVIDIA, 2026 )
18.75
50.00
62.50
43.75
3035
GLM-5.3-Flash ( Z.ai, 2026 )
37.50
52.50
58.75
49.58
403
GPT-6 Astra ( OpenAI, 2026 )
60.00
76.25
82.50
72.92
8724
Table 2: Diagnostic Results. Accuracy (%) on action selection, progress verification, subgoal completion, and all 240 questions overall. Latency denotes mean inference time per query in milliseconds.
Figure 3: Overview of MotorMind . The Planner specifies subgoals and success criteria. The Executor proposes short action batches, which the Controller converts into robot motion. During execution, the Monitor evaluates updated observations and can request interruption at an action boundary, cancelling pending commands. Outcome assessment then determines whether to advance, retry, or replan. Memory summarizes execution evidence and assessed outcomes in the background for subsequent decisions.
Method
Configuration
Base success rate (%)
Perturbation success rate (%)
Base efficiency
Perturbation efficiency
Goal
Spatial
Object
Avg.
Semantic
Object
Position
Task
Avg.
Time (s) ↓
Time score ↑
Time (s) ↓
Time score ↑
Fine-Tuned Policies / Agentic Methods with Fine-Tuned Policies
π0.5
Fine-tuned
95.0
100.0
100.0
98.3
98.3
95.0
36.7
23.3
63.3
5.9
997.1
6.8
562.1
MolmoAct2
Fine-tuned
100.0
100.0
100.0
100.0
95.0
90.0
36.7
30.0
62.9
6.6
914.6
7.7
492.2
OpenVLA / OFT
Fine-tuned
95.0
100.0
100.0
98.3
96.7
86.7
11.7
10.0
51.2
5.7
1033.5
6.7
459.9
GR00T N1.5
Fine-tuned
0.0
95.0
0.0
31.7
30.0
30.0
1.7
16.7
19.6
15.6
121.8
16.2
72.4
Table 3: Main Evaluation Results. Success rates and averages are reported separately for unperturbed base tasks and perturbed tasks; walltime and time-normalized success score use the corresponding evaluation subset. Methods are grouped by whether their underlying policy uses task-specific fine-tuning. Time score is 60s/Tˉ in pp/min, where s is success on a 0–100 scale and Tˉ is wall time in seconds; it does not measure monetary or compute cost.
Method
Dynamic Reasoning
Scene Shift
Dynamic Manipulation
Prompt Shift
(10 tasks)
(10 tasks)
(5 tasks)
(5 tasks)
π0.5
0%
70%
0%
20%
GR00T N1.5
0%
10%
0%
0%
MolmoAct2
20%
50%
40%
0%
CaP-X
30%
20%
20%
60%
MotorMind
70%
90%
80%
60%
Table 5: Adaptive reasoning evaluation. Left: Success rates across four task groups. Right: Example in which the robot must reason and act while the scene evolves. Parentheses indicate the number of tasks; bold and underlined entries denote the best and second-best results, respectively.
Object
Direct Perception
Human Intervention
Avg. SR
Bowl
Box
Bowl
Box
Blue Cube
100%
100%
100%
100%
100%
Corn
100%
100%
100%
80%
95%
Battery
90%
100%
80%
90%
90%
Avg. SR
97%
100%
93%
90%
95%
Table 6: Real-robot zero-shot deployment. Left: Success rates in Direct Perception and Human Perturbation tasks. Right: Success rates on Semantic Understanding tasks. Instructions include semantic categories, visual attributes, or spatial relations.
Figure 4: Backbone sensitivity and component ablations. Left: success rate and wall time across three LIBERO-PRO base task suites on single seed with Qwen3.8-Flash-Next and GPT-6 Sol (Medium Reasoning) Right: performance after removing replanning, verification, or planning. Δ Success is measured in percentage points relative to the full system’s 66.7% success rate. Time score uses the definition in Section 4 (pp/min).
Figure 5: Failure analysis of VLM decision errors across LIBERO-PRO settings.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Decision being diagnosed
Visual input
Count
Action selection
Primary primitive to execute after the current observation t1
Scene and wrist views at t0,t1
80
Progress verification
Whether the transition from t0 to t1 advances the stated subgoal
Scene and wrist views at t0,t1
80
Subsubgoal completion
Whether the state at t1 satisfies the stated subgoal and criterion
Scene and wrist views at t1
80
Appendix
Table 8: Composition of the manipulation decision diagnostic.
Model
Translation (48)
Rotation (19)
Gripper (13)
Qwen3.8-Flash-Next-FP8
47.92% (23/48)
0.00% (0/19)
46.15% (6/13)
HY-Embodied-0.5 MoT-2B
20.83% (10/48)
15.79% (3/19)
23.08% (3/13)
Hy-Embodied-VLM-1.0 A3B
31.25% (15/48)
0.00% (0/19)
0.00% (0/13)
Cosmos3-Nano (understanding tower)
27.08% (13/48)
0.00% (0/19)
15.38% (2/13)
GLM-5.3-Flash
41.67% (20/48)
15.79% (3/19)
53.85% (7/13)
GPT-6 Astra (medium)
58.33% (28/48)
63.16% (12/19)
61.54% (8/13)
Appendix
Table 9: Action-selection results by primitive type. Accuracy is reported separately for translation, rotation, and gripper questions.
Condition
Configuration key
Description
Base
base
Original, unperturbed task setting
Language
lan
Instruction wording perturbations
Object
object
Benchmark-defined object perturbations
Position Swap
swap
Object placement perturbations
Task
task
Target or goal perturbations
Appendix
Table 10: LIBERO-Pro evaluation conditions. Each condition is applied to every evaluated task family.
Baseline
Control mechanism
Planning and memory
OpenPI π0.5
Direct VLA action prediction
No external planner or task memory
OpenVLA
Autoregressive action-token prediction
No external planner or task memory
OpenVLA-OFT
Continuous action-chunk prediction
No external planner or task memory
GR00T N1.5
Native action-chunk prediction
No external planner or task memory
MolmoAct2
Continuous action prediction
No external planner or task memory
Harness VLA
Tool-based execution with frozen π0.5
VLM planning and task-specific memory
Appendix
Table 11: Baseline configurations. Direct policies predict robot actions, whereas agent frameworks add planning, tools, or programmatic execution around their control policy.
Suite
Tasks
Evaluation focus
Interactive
10
Adapt to object or receiver displacement with an unchanged instruction.
Pure Dynamic
5
Identify, intercept, and place a named moving target.
Dynamic Reasoning
10
Resolve semantic, relational, and temporal references to moving targets.
Prompt Change
5
Respond to revised instructions without resetting the physical scene.
Appendix
Table 12: Adaptive reasoning task suites. The four suites contain 30 online manipulation tasks in total.
Task
Instruction
Intervention
1
put the ketchup in the wooden tray
Receiver during transport; 1 displacement
2
put the bowl on the plate
Receiver during transport; 1 displacement
3
put the wine bottle on the rack
Target object before grasping; 1 displacement
4
put the cream cheese in the bowl
Receiver during transport; 1 displacement
5
put the bbq sauce in the wooden tray
Target object before grasping; 2 displacements
6
put the milk in the basket
Receiver during transport; 2 displacements
Appendix
Table 13: Scene shift tasks. The instruction remains unchanged while the target object or receiving container is displaced.
Task
Instruction
1
pick up the red mug from the conveyor belt and place it in the basket
2
pick up the alphabet soup from the conveyor belt and place it in the basket
3
pick up the ketchup from the conveyor belt and place it in the red basket
4
pick up the blue-and-white cream cheese box from the conveyor belt and place it in the wooden tray
5
pick up the green ketchup from the conveyor belt and place it in the basket
Appendix
Table 14: Dynamic manipulation tasks. Each instruction identifies a named target moving on a single-pass conveyor.
Task
Reasoning requirement
Instruction
1
Contents: liquid
pick up the packaged product containing liquid from the conveyor belt and place it in the basket
2
Category: edible item
pick up the edible item from the conveyor belt and place it in the basket
3
Purpose: beverage
pick up the beverage meant for drinking from the conveyor belt and place it in the basket
4
Category exclusion: not a bowl
pick up the drinking vessel that is not a bowl from the conveyor belt and place it in the basket
5
Attribute exclusion: not red
pick up the cup that is not red from the conveyor belt and place it in the basket
6
Initial spatial relation
pick up the object that is between the two bowls in the initial arrangement on the conveyor belt and place it in the basket
Appendix
Table 15: Dynamic reasoning tasks. The target is specified through semantic, exclusion-based, spatial, or temporal constraints.
Task
Initial instruction
Replacement instruction
1
put the ketchup in the wooden tray
Change of plan: put the ketchup on the plate instead of in the wooden tray.
2
Put the chocolate pudding on the plate.
Change of plan: put the chocolate pudding in the wooden tray instead of on the plate.
3
Put the chocolate pudding on the plate.
First put the chocolate pudding down on the table, turn on the stove and keep it on, and then continue putting the chocolate pudding on the plate.
4
put the chocolate pudding on the plate
Avoid the obstacle and continue putting the chocolate pudding on the plate.
5
Put the BBQ sauce in the wooden tray.
Change of plan: leave the BBQ sauce on the table and put only the chocolate pudding in the wooden tray.
Appendix
Table 16: Prompt shift tasks. The instruction changes after a physical trigger without resetting the scene or robot.
Budget
Goal
Spatial
Object
Avg. SR
Time (s)
Time score
300 s
40.0%
70.0%
90.0%
66.7%
159.8
25.03
450 s
50.0%
70.0%
80.0%
66.7%
210.0
19.05
900 s
40.0%
60.0%
50.0%
50.0%
320.0
9.37
3,600 s
50.0%
60.0%
70.0%
60.0%
296.4
12.14
Appendix
Table 17: Execution-budget ablation. Average success does not increase monotonically with the maximum execution budget. Time score uses the definition in Section 4 (pp/min). The 300 s setting matches the success of the 450 s setting while requiring less wall time.
Figure 6: Physical failure analysis tracing MotorMind episodes through failures, recovery, and outcomes on unperturbed (left) and perturbed (right) LIBERO-PRO tasks. Each failed episode is assigned its last unrecovered failure; Recovered episodes had failures and recovered from all of them. False done : the model declared the task complete while its goal predicate was false. Band thickness is proportional to the share of episodes within each panel.
Current Vision-Language-Action (VLA) models predominantly rely on end-to-end fine-tuning. While effective, this paradigm compromises the inherent generalization capabilities of Vision-Language Models (VLMs) and incurs catastrophic forgetting. To address these limitations, we propose M2-VLA, which demonstrates that a generalized VLM is able to serve as a powerful backbone for robotic manipulation directly. However, it remains a key challenge to bridge the gap between the high-level semantic understanding of VLMs and the precise requirements of robotic control. To overcome this, we introduce the Mixture of Layers (MoL) strategy that selectively extracts task-critical information from dense semantic features. Furthermore, to facilitate efficient trajectory learning under constrained model capacity, we propose a Meta Skill Module (MSM) that integrates strong inductive biases. Extensive experiments in both simulated and real-world environments demonstrate the effectiveness of our approach. Furthermore, generalization and ablation studies validate the architecture's zero-shot capabilities and confirm the contribution of each key component. Our code and pre-trained models will be made publicly available.
Siyao Xiao, Yuhong Zhang, Zhifang Liu +9
Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China. · PengCheng Laboratory, Shenzhen, China. · College of Intelligent Robotics and Advanced Manufacturing, Fudan University, Shanghai, China. +1
General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.
The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7% vs.\ 53.3% for π0.5 and 24.4% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4% vs.\ 26.3% for π0.5 and 26.1% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5%), a language-following Pick-and-Place benchmark, LIBERO (98.9%), RoboTwin 2.0 (93.3%), and SimplerEnv-Bridge (87.3%).