RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
Abstract
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.
Figures & tables
| Evaluation | Control interfaces | Beyond manipulation | ||||
|---|---|---|---|---|---|---|
| Direct actions | Code control | Mixed calls | Navigation | Balance | Driving / flight | |
| CaP-X ( Fu et al., 2026 ) | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ |
| EmbodiedBench ( Yang et al., 2025 ) | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| VLABench ( Zhang et al., 2024 ) | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ |
| EmbodiedEval ( Cheng et al., 2025 ) | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Embodied Agent Interface ( Li et al., 2024b ) | ✓ | ✗ | ✗ | ✓ | ✗ | ✗ |
| Embodiment | Goal requirements | Dynamic control | |||
|---|---|---|---|---|---|
| Spatial goals | Constrained contact | Multiple subgoals | Balance / tracking | Timed interaction | |
| Fixed-base arms | ✓ | ✓ | ✓ | ✓ | ✓ |
| Dexterous / bimanual | ✓ | ✓ | ✓ | ✓ | ✗ |
| Mobile manipulators | ✓ | ✓ | ✓ | ✗ | ✗ |
| Humanoids / bipeds | ✗ | ✓ | ✓ | ✓ | ✓ |
| Quadrupeds | ✗ | ✗ | ✗ | ✓ | ✗ |
| Model | Manip. | Mobile manip. | Locomotion | Driving | Aerial | Overall |
|---|---|---|---|---|---|---|
| Astra | 9/38 | 4/20 | 0/11 | 2/11 | 1/4 | 16/84 (19.0%) |
| Opus 5.5 | 7/38 | 0/20 | 1/11 | 2/11 | 3/4 | 13/84 (15.5%) |
| Kimi K3 | 0/38 | 1/20 | 0/11 | 1/11 | 0/4 | 2/84 (2.4%) |
| DeepSeek V4.1 Flash | 0/38 | 0/20 | 0/11 | 0/11 | 1/4 | 1/84 (1.2%) |
| Gemini 3.8 Flash | 1/38 | 0/20 | 0/11 | 0/11 | 0/4 | 1/84 (1.2%) |
| Case / model | Agent judgement | Subsequent behaviour | Recorded result |
|---|---|---|---|
| Block sweeping Kimi K3 | Declares completion at step 861 | 69 empty-arm height commands; 139 more steps | Unfinished at the 1,000-step horizon |
| Stove navigation DeepSeek | No further motion needed at step 345 | 105 more steps with zero base-velocity commands | Unfinished at the 450-step horizon |
| Hanging mugs Opus 5.5 | 35 steps deemed insufficient to grasp and hang a mug | Returns the arm home over 35 steps | Unfinished at the 800-step horizon |
| White mug Opus 5.5 | Requests gripper release at step 340 | One of 15 requested steps executes | Native success at step 341 |
Appendix figures & tables19 assets
Supplementary material from the paper’s appendix.
Appendix
| Task source | Tasks | Reference |
|---|---|---|
| HumanoidSoccer | 1 | ( Kong et al., 2026 ) |
| ReflexBench | 1 | ( Chen et al., 2026c ) |
| Digit | 1 | ( Mittal et al., 2025 ) |
| Flamingo | 1 | ( jaykorea, n.d. ) ; software |
| Go2 Push | 1 | ( BrandoUlissi, n.d. ) ; software |
| AI-CPS | 3 | ( Zhou et al., 2024 ) |
| Domain | Primary task type | Tasks |
|---|---|---|
| Manipulation | Fitting & insertion | 7 |
| Manipulation | Placement & organisation | 13 |
| Manipulation | Pouring & processing | 3 |
| Manipulation | Multi-stage workflows | 8 |
| Manipulation | Articulation & device use | 3 |
| Manipulation | Dynamic manipulation | 4 |
| Source group | Tasks | Registered robot tools |
|---|---|---|
| RoboDojo | 15 | move_eef |
| RoboLab | 10 | move_eef , set_gripper , move_robot |
| RoboCasa | 10 | move_eef , move_base , move_torso , set_gripper , move_robot |
| BEHAVIOR-1K | 10 | move_arms , move_base , set_grippers , move_torso , move_robot |
| WheeledLab | 11 | observe , drive |
| AI-CPS | 3 | observe , move_joints |
| Adapter | Tool | Inputs and meaning |
|---|---|---|
| RoboDojo | move_eef | targets : selected left/right XYZ (world metres), wrist pitch/roll/yaw (degrees from the downward reference), and gripper opening (0 closed, 1 open). Omitted fields hold measured call-start values. The planner moves the arm before applying a combined gripper change. |
| RoboLab | move_eef | targets.position[3] and/or quaternion_wxyz[4] : absolute flange pose in the robot-root frame. Omitted pose components hold measured values. steps=1..30 at 15 Hz. |
| RoboLab | set_gripper | targets.gripper_close : 0 opens, 1 closes; the target persists. The arm holds while the segment executes. |
| RoboLab | move_robot | Combines the same arm-pose and gripper fields in one native action. Closure starts with arm motion, not after arrival. |
| RoboCasa | move_eef | targets.eef_delta[6] : normalised translation and rotation-vector increments in the base frame. Scales are 0.05 m and 0.5 rad per step. steps=1..30 ; repetition repeats the increment. |
| RoboCasa | move_base | targets.base_motion[3] : normalised XY/yaw controller inputs, not displacement. Uses base-following mode; gripper target persists. |
| Adapter | Tool | Inputs and meaning |
|---|---|---|
| BEHAVIOR-1K | move_arms | Left/right XYZ in robot-root metres and left/right_quat_xyzw . Absolute EEF IK; omitted arm holds its measured pose. All tools in this group use steps=1..30 at 30 Hz. |
| BEHAVIOR-1K | move_base | base_vx , base_vy in local-body m/s (limits +/-0.3), and base_wz in rad/s (+/-0.5). Arms hold root-relative targets; omitted base velocities are zero. |
| BEHAVIOR-1K | set_grippers | left/right_gripper : continuous opening in [0,1], with 0 closed and 1 open. Targets persist; EEFs hold measured poses and base motion stops. |
| BEHAVIOR-1K | move_torso | trunk_qpos[4] : absolute trunk joint positions in radians, using the recorded joint order and limits. This differs from the normalised slide increment in RoboCasa. |
| BEHAVIOR-1K | move_robot | Combines arm poses, gripper openings, base velocities and trunk targets simultaneously. One control step counts once, regardless of how many components are commanded. |
| WheeledLab | observe | Returns the current allowed observation without stepping physics. Authored onboard courses expose front RGB, encoders and IMU; the call does not reveal map, world pose or checkpoint progress. |
| Source group | Dim. | Meaning of the action vector |
|---|---|---|
| ReflexBench | 8 | Absolute root-frame EEF XYZ + WXYZ quaternion + gripper sign. Positive gripper opens; non-positive closes. |
| Digit | 26 | Joint-position targets with the recorded offsets and 0.5 rad input scale, in runtime joint order. |
| Flamingo | 8 | Six actuator-position channels and two wheel-velocity channels (40 rad/s scale). Leg channels use motor-space angles with gear ratio -1.5. |
| Go2 Push | 12 | Joint-position residuals: radians. |
| OmniIsaacGymEnvs | 12 | ANYmal joint-position offsets: radians, tracked by native PD. |
| OmniDrones | 4 | Four rotor inputs. -1 requests zero thrust; +1 requests maximum. Zero is half maximum steady-state thrust, not hover. |
| ID | Source / task identifier | Limit | Score | A | O | K | D | G |
|---|---|---|---|---|---|---|---|---|
| RoboDojo | ||||||||
| 01 | hang_mugs | 800 | N | ✗ | ✗ | ✗ | ✗ | ✗ |
| 02 | sweep_blocks | 1000 | N | ✗ | ✓ | ✗ | ✗ | ✗ |
| 03 | pour_liquid_into_cup | 400 | N | ✗ | ✗ | ✗ | ✗ | ✗ |
| 04 | make_toast | 1400 | N | ✗ | ✗ | ✗ | ✗ | ✗ |
| 05 | store_laptop_and_headphones | 800 | N | ✗ | ✗ | ✗ | ✗ | ✗ |
| Model | Elapsed time (min) | Robot requests | All tool calls | |||
|---|---|---|---|---|---|---|
| Median | IQR | Median | IQR | Median | IQR | |
| Astra | 30.9 | [15.2, 61.2] | 58 | [25.5, 94] | 70 | [30, 138.8] |
| Opus 5.5 | 54 | [24.5, 137.3] | 37 | [16, 87.2] | 57.5 | [31, 148] |
| Kimi K3 | 67.6 | [21.3, 181.7] | 42.5 | [11.8, 70] | 61.5 | [18.5, 95] |
| DeepSeek | 16 | [11.1, 40.5] | 22 | [5, 59.2] | 33 | [18, 70] |
| Gemini | 12.8 | [10, 25.8] | 11 | [0, 36.8] | 31 | [15, 52.2] |
| Model | Control steps | Shell calls | Image views | |||
|---|---|---|---|---|---|---|
| Median | IQR | Median | IQR | Median | IQR | |
| Astra | 627 | [298, 1165] | 4.5 | [0, 21.2] | 0 | [0, 0] |
| Opus 5.5 | 458 | [178.5, 913.8] | 15 | [0, 28] | 0 | [0, 10] |
| Kimi K3 | 449.5 | [158, 918.5] | 1 | [0, 14.2] | 0 | [0, 1.2] |
| DeepSeek | 347 | [62, 742] | 8.5 | [0, 16] | 2.5 | [0, 8] |
| Gemini | 315 | [78.5, 625] | 15 | [9.8, 23] | 0 | [0, 2.2] |
| Model | Tokens | Estimated USD | Elapsed hours |
|---|---|---|---|
| Astra | 941M | 9,912.80 | 54.6 |
| Opus 5.5 | 1.63B | 2,916.43 | 140.3 |
| Kimi K3 | 2.00B | 624.09 | 163.8 |
| DeepSeek V4.1 Flash | 907M | 34.49 | 35.7 |
| Gemini 3.8 Flash | 303M | 66.68 | 28.7 |
| Pattern | Trace evidence | Interpretation boundary |
|---|---|---|
| Sequential decomposition | Astra stacking names the placement order and reuses a stack centre after release/retraction. | An explicit plan is evidence of intent; the final checker confirms the completed arrangement. |
| Feedback-driven correction | Gemini revises a rejected 0.95 m lift to 0.85 m; the next receipt advances from step 501 to 508. | The consecutive request/receipt pair supports a local correction, not a causal effect of a general recovery policy. |
| Frequent dynamic feedback | Opus 5.5 juggling uses 274 robot requests over 800 control steps with repeated strike/descent phases. | This successful case does not establish that more calls universally improve success. |
| Calibration revision | DeepSeek changes its hover computation after reporting excessive rise. | A visible self-correction is distinct from independently verified correctness of its entire controller. |
| Contact/grasp reassessment | Kimi pouring repeatedly reports missed closure and changes approach or wrist orientation. | The agent notices difficulty; final failure should not be described as an unobserved or hallucinated success. |
| Pre-action analysis saturation | DeepSeek computer-use episode reaches a non-action boundary at step zero. | No robot motion is available for diagnosing physical control ability in that episode. |
| Task | Robot requests | All calls | Result steps |
|---|---|---|---|
| Conveyor matching | 45 / 35 | 45 / 35 | 380 / 347 |
| Plastic clutter sorting | 41 / 104 | 46 / 107 | 784 / 1497 |
| Three fruits on a plate | 21 / 31 | 22 / 58 | 284 / 419 |
| Cube left of bowl | 17 / 18 | 24 / 31 | 286 / 239 |
| White mug centring | 18 / 22 | 20 / 42 | 280 / 341 |
| Volleyball 1v1 | 30 / 152 | 30 / 157 | 135 / 498 |
| Trace | Model / task | Original | Scored step | Observed step |
|---|---|---|---|---|
| task-1 | Kimi K3 / CountertopCleanup | failure | 489 | 600 |
| task-104 | Opus 5.5 / visual | failure | 19 | 143 |
| task-13 | Kimi K3 / OrganizeMugsByHandle | unavailable | 335 | 335 |
| task-17 | Kimi K3 / MicrowaveCorrectMeal | unavailable | 372 | 372 |
| task-19 | Kimi K3 / ResetCabinetDoors | unavailable | 2420 | 2420 |
| task-2 | Opus 5.5 / SortingCleanup | unavailable | 778 | 778 |