World Action Agent: Harnessing VLMs for Robot Manipulation via World Action Rehearsal
Organizations: HKUST(GZ) · Knowin AI · CUHK
Abstract
General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirectly, to predict constraints or write programs, or give them a view of the scene rather than a world in which to act. We present World Action Agent (WAA), a multi-agent harness through which VLMs pilot robots with basic tools, making every decision within a visual action workspace. The workspace has three properties. Contact views, selected automatically from the scene geometry, present the scene around the current interaction. Action rehearsal turns each action into an editable proposal that the agent, alone or through an Imagination Agent, previews and revises against planning feedback before execution. In-view correction closes the loop between observation, rehearsal, and low-level execution, letting the agent remove residual offsets in the view where it observes them. Through the same workspace, WAA acquires embodied procedural knowledge in two ways: it evolves multimodal skills from expert videos and human teaching under evidence-based review and consults them through a Skill Agent, and its interaction traces train smaller VLMs to pilot the same harness. On LIBERO-Pro, WAA with skills evolved only from LIBERO-90 reaches a state-of-the-art 75.6% average success, outperforming end-to-end VLAs, code-as-policy agents, and a visual-harness baseline with the same backbone; the same skills remain effective on robosuite without further learning. Fine-tuning Qwen3.5-9B on harness traces raises its out-of-domain success from 1.7% to 43.3%.
Figures & tables
| Object | Goal | Spatial | |||||
| Method | Pos. | Task | Pos. | Task | Pos. | Task | Avg. |
| OpenVLA | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| 17.0 | 1.0 | 38.0 | 0.0 | 20.0 | 1.0 | 12.8 | |
| CaP-Agent0 (CaP-X) | 22.0 | 18.0 | 26.0 | 17.0 | 12.0 | 14.0 | 18.2 |
| CaP-Agent0 (Playful) | 27.0 | 31.0 | 29.0 | 16.0 | 13.0 | 23.0 | 23.2 |
| RATs (Playful) | 61.0 | 63.0 | 43.0 | 36.0 | 29.0 | 31.0 | 43.8 |
| Method | API cost (\downarrow$ | Model calls | Time (s) | Input token (K) | Output token (K) |
|---|---|---|---|---|---|
| Show-Harness | 0.5021 | 119.80 | 874.12 | 347.10 | 62.90 |
| WAA | 0.1996 | 31.05 | 150.47 | 193.62 | 5.45 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Function | Arguments | Operation and feedback |
|---|---|---|
| detect_region | query : string within_region_id ∗ : string | Detect and segment a region in the latest observation. Return a region reference and, when available, the centroid of its visible point cloud in the base frame. No robot motion. |
| propose_grasps | region_id : string | Generate grasp candidates with seed_id references. A candidate must be converted to an action with preview_grasp before execution. |
| locate_point | query : string within_region_id ∗ : string force_refresh ∗ : bool | Estimate a coarse 3D anchor and display its point reference on the canvas. Optionally restrict grounding to a region or refresh the cached estimate. No robot motion. |
| preview_pose | point_id : string offset_xyz_m : number[3] quaternion_xyzw ∗ : number[4] | Construct a pending pose and motion plan. Target TCP position equals the referenced point plus the base-frame offset. Omitted orientation preserves the current orientation. |
| preview_grasp | seed_id : string | Convert a grasp candidate into a pending action and motion plan for visual inspection. Return an action reference without executing or closing the gripper. |
| imagine_action | instruction : string action_id ∗ : string | Request local geometric inspection and refinement by the Imagination Agent. Start from the referenced action, or the current TCP when omitted. Return a virtual preview without physical execution. |
| Function | Arguments | Operation and feedback |
|---|---|---|
| shift_preview | delta_xyz_m : number[3] frame : base or tool | Translate the virtual target in the selected frame and update the preview for inspection, subject to the metric bounds above. |
| rotate_preview | angle_deg : number | Rotate the virtual target about its tool-local axis, with a right-handed angle in degrees. |
| finish_imagination | status : ready or failed | Return the current preview on ready , or abandon the current edits on failed . Readiness does not indicate execution or task success. |
| Function | Arguments | Operation and feedback |
|---|---|---|
| select_skill_evidence | state_ids : string[] reference_ids : string[] reason : string | Select at most two state entries and a configuration-bounded number of reference images. The selection reason is limited to 240 characters. |
| report_skill_guidance | applicability : enum reference_differences : string[] uncertainty : string[] | Report applicable , not_applicable , or uncertain . Each list contains at most two statements, each limited to 120 characters. |
| Setting | Value |
|---|---|
| Training data | 112 successful LIBERO-Pro episodes (Gemini 3.7 Flash, evolved skills) |
| Samples | 1,774 main-agent tool calls |
| Framework | LLaMA-Factory 0.9.5 |
| Fine-tuning method | LoRA on all linear layers (rank 8, , dropout 0) |
| Trainable parameters | 21.6M |
| Frozen modules | Vision encoder, multimodal projector |