General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.
Figures & tables
Figure 1 : RobotWorld leaderboard. Left: overall success across 84 tasks. Right: success in four displayed groups (Manipulation and Mobile manipulation are pooled), with counts shown beside each bar. All panels use the same 0–100% success scale. Each model has one retained outcome per task under the evaluation budget rules.
Figure 2 : Overview of RobotWorld. RobotWorld evaluates whether general-purpose multimodal models can turn digital capabilities into reliable robot use. Agents interpret instructions and observations, perform auxiliary computation, and act through a feedback loop across diverse robots and tasks. Trajectory analysis examines emerging capabilities—including image segmentation, geometric calibration, dynamics modelling and control program generation—alongside failure mechanisms and differences between models on shared tasks. These findings inform training and agent design. The scene illustrates task families evaluated in separate simulation environments.
Evaluation
Control interfaces
Beyond manipulation
Direct actions
Code control
Mixed calls
Navigation
Balance
Driving / flight
CaP-X ( Fu et al., 2026 )
✗
✓
✗
✓
✗
✗
EmbodiedBench ( Yang et al., 2025 )
✓
✗
✗
✓
✗
✗
VLABench ( Zhang et al., 2024 )
✓
✓
✗
✗
✗
✗
EmbodiedEval ( Cheng et al., 2025 )
✓
✗
✗
✓
✗
✗
Embodied Agent Interface ( Li et al., 2024b )
✓
✗
✗
✓
✗
✗
Table 1: Control interfaces and task coverage ( ✓ : included; ✗ : not reported in the cited protocols). Direct actions are model-emitted commands or action values; code control executes generated programs. Mixed calls alternate these interfaces within one episode. Balance denotes tasks with an explicit robot-stability objective. Optional RobotWorld code control is included.
Figure 3 : The RobotWorld framework. Under a task-specific observation, action, and budget contract, a multimodal agent alternates between analysis and robot execution, using returned observations and execution feedback to revise its actions. Direct actions execute bounded command segments; optional code control, disabled by default, runs generated programs with feedback at each control step. The analysis workspace supports perception, computation, and memory while physics is paused. Evaluation records task outcomes, execution validity, resource use, and behaviour across manipulation, locomotion, driving, and flight.
Figure 4 : Environment interaction loop. Observations and task context inform the agent’s next robot call, optionally supported by workspace computation. Native execution validates the request and advances physics within its step limit; returned observations, completed steps, and errors inform the next decision. Physics pauses during reasoning and offline computation. The evaluator checks task state and events independently. Camera views illustrate observations from a stacking episode, rather than the endpoints of a single call; the execution strip is schematic.
Figure 5 : Control is chosen per call. Direct actions and code segments can alternate within one episode. The agent replans between calls; within a code segment, the program updates actions from fresh state at each step, up to its limit of M steps.
Figure 6 : Task composition by domain and primary task type. The inner ring shows five domains and the outer ring their primary task types. Each task is counted once. Percentages use all 84 tasks, not action frequencies.
Figure 7 : What establishes task completion. Six families of success evidence, illustrated schematically. (a) Kitchen navigation checks position and heading. (b) Drawer closure checks normalised joint opening. (c) Mug placement requires support and release in addition to its spatial goal. (d) Payload hovering requires all stability bounds to hold together for at least the final two seconds; a broken continuous hold resets its duration. (e) Repeated piano notes require release before re-pressing. (f) Reverse parking combines entry history, final pose, and boundary constraints. Thresholds are specific to the illustrated tasks; these are checker components, not universal rules for all 84 tasks.
Embodiment
Goal requirements
Dynamic control
Spatial goals
Constrained contact
Multiple subgoals
Balance / tracking
Timed interaction
Fixed-base arms
✓
✓
✓
✓
✓
Dexterous / bimanual
✓
✓
✓
✓
✗
Mobile manipulators
✓
✓
✓
✗
✗
Humanoids / bipeds
✗
✓
✓
✓
✓
Quadrupeds
✗
✗
✗
✓
✗
Table 2: Representative task coverage by embodiment ( ✓ : represented; ✗ : no selected example). Labels overlap and describe task requirements, not model success.
Figure 8 : Success coverage across the 84 tasks. Each aligned column represents one task, grouped by its primary domain. Coloured cells indicate success and grey cells indicate failure. Tasks are ordered identically for all models, with successful tasks grouped within each domain. The bottom row marks tasks solved by at least one model: 21 tasks are covered, leaving 63 unsolved. Each model has one evaluated episode per task; counts use the outcomes at the applicable evaluation boundary.
Model
Manip.
Mobile manip.
Locomotion
Driving
Aerial
Overall
Astra
9/38
4/20
0/11
2/11
1/4
16/84 (19.0%)
Opus 5.5
7/38
0/20
1/11
2/11
3/4
13/84 (15.5%)
Kimi K3
0/38
1/20
0/11
1/11
0/4
2/84 (2.4%)
DeepSeek V4.1 Flash
0/38
0/20
0/11
0/11
1/4
1/84 (1.2%)
Gemini 3.8 Flash
1/38
0/20
0/11
0/11
0/4
1/84 (1.2%)
Table 3: Main results across five task domains. Entries give successful tasks / evaluated tasks; overall success rates are in parentheses. One episode is retained per model–task pair. Figure 1 pools the first two domains for compact presentation.
Figure 9 : Representative successful trajectories of Astra (left) and Opus 5.5 (right). Each strip contains five chronological video frames. The first two rows compare shared successes; the remaining rows show tasks solved only by the displayed model in the illustrated trajectory comparison. Sampling intervals vary, and final frames may precede recorded success.
Figure 10 : Recorded resource consumption. Points and labels show medians; horizontal segments span the 25th–75th percentiles across 84 retained runs per model, including failed runs. Models use the same colours and row order in every panel. Elapsed time includes setup, waiting and execution. Total tool calls include robot requests, shell calls and explicit image-view calls. Full logs may extend beyond retrospective adjudication cutoffs; the intervals describe variation across tasks, rather than uncertainty over repeated trials.
Figure 11 : Success versus aggregate resource consumption. Each point represents one model evaluated on 84 tasks; all panels use the same success axis. (a) Estimated list-price cost. (b) Total tokens, including cached input and output. (c) Cumulative elapsed time, including unsuccessful runs. All values follow the project website: valid per-task mean multiplied by 84. Token and cost estimates use 57 valid records for Kimi and 84 for each other model; all elapsed-time totals use 84 records. Missing usage is estimated by this scaling, rather than counted as zero. Resource accounting covers full recorded runs and may extend beyond retrospective scoring cutoffs; these plots do not measure cost or time to first success.
Figure 12 : Feedback-dependent correction and its limits. (a) Kimi K3 shortens successive drawer pushes; withdrawal executes two of twelve requested steps before success. (b) Opus 5.5 sustains payload hover: shading marks all stability conditions jointly satisfied for 2.15 seconds, exceeding the two-second requirement. (c) Astra changes the idle arm’s position after a downward camera rotation is rejected at step 196: it lowers the arm by step 207 and successfully retries the rotation by step 230. The episode subsequently completes charger insertion at step 323. Frames show the corresponding states from the retained recording; the rejected call itself executes no steps.
Figure 13 : Execution without task completion. (a) Kimi K3 topples the bottle during approach, fails to secure it with the left arm, and switches to a right-arm attempt late in the 400-step episode. Images are sampled from the current run at the labelled control steps. (b) Astra crosses the tracking bound (20 cm) before recovering below the terminal bound (10 cm); shading marks all terminal conditions, including speed and tilt. The final joint hold lasts only 0.565 seconds. (c) Chopping never advances to the next subgoal: 69 positive-step calls cover the entire 2,000-step episode. Sweeping ends with 69 calls that vary only the empty left arm’s height over the final 139 steps, while the model asserts completion and the native task remains unsuccessful. Each bar spans its own episode horizon. White dividers mark chopping calls; the purple sweeping segment marks the final 139 steps. The chopping count covers the whole episode; the sweeping count covers only this final segment.
Figure 14 : Sustaining contact under the same initial observation. (a) Matched frames from Astra and Opus 5.5 near steps 143 and 188. Both have made three valid contacts near step 143; Astra terminates at 188 while Opus 5.5 continues. (b) Ball height over the first 200 steps; the dashed line marks the 3.5 m qualification threshold. (c) Valid and height-qualified contact counts, ending at each trajectory’s termination. The first contact is not height-qualified. Astra ends with 3 valid / 1 qualified contacts; Opus 5.5 reaches 15 / 14 at step 800. Both use a median of three executed steps per action. Frames are nearest samples from 12-fps recordings.
Figure 15 : Different routes from observation to grasping. Milestones from the two stacking episodes on a shared control-step axis. Filled circles mark observations or physical milestones; open squares mark explicit computation, during which physics is paused. Astra closes the gripper at step 157 and lifts at 179, before its first camera fit at 204. Opus 5.5 writes colour-segmentation code before motion, samples the arm and fits camera geometry at step 45, then lifts at 188. Both reach first placement and retraction at similar physical steps (270 versus 268), but their information-gathering sequences differ. Astra completes the stack at step 635; Opus 5.5 reaches the applicable auxiliary-call budget at step 313. These observations describe strategies in the illustrated episodes and do not identify the models’ training data.
Case / model
Agent judgement
Subsequent behaviour
Recorded result
Block sweeping Kimi K3
Declares completion at step 861
69 empty-arm height commands; 139 more steps
Unfinished at the 1,000-step horizon
Stove navigation DeepSeek
No further motion needed at step 345
105 more steps with zero base-velocity commands
Unfinished at the 450-step horizon
Hanging mugs Opus 5.5
35 steps deemed insufficient to grasp and hang a mug
Returns the arm home over 35 steps
Unfinished at the 800-step horizon
White mug Opus 5.5
Requests gripper release at step 340
One of 15 requested steps executes
Native success at step 341
Table 4: Completion judgements and subsequent execution. Holding and return-to-home commands advance simulation even when they do not advance the unfinished task. In white-mug placement, the tool closes without a final observation or success score; the release request may itself trigger success and does not establish post-success over-action.
Figure 16 : Tool calls and physical execution measure different costs. Stacking ends at the shared first blue-on-red placement and retraction, before Opus 5.5’s step-313 budget cutoff. Cube and fruit comparisons cover full successful episodes. Separate panels count robot requests, shell executions, explicit image-view calls and control steps. Total tool calls (Astra / Opus 5.5) are 23 / 46 for stacking, 24 / 31 for cube and 22 / 58 for fruit. Images embedded in robot feedback are not additional image-view calls.
Figure 17 : Interaction distributions across 84 tasks. Points represent the full retained episode for each task, including failures; boxes show medians and interquartile ranges, with whiskers extending to the most extreme points within 1.5 interquartile ranges. Robot requests include rejected and zero-step requests; shell counts exclude polling. Explicit image views exclude images embedded in feedback. Astra has zero explicit views on 66/84 tasks and Opus 5.5 on 43/84, so both medians are zero; their upper quartiles are 0 and 10 calls, respectively. Symmetric-log axes retain zeros. These distributions describe tool use, not efficiency or success-conditioned cost.
Figure 18 : Resource use on the eight tasks solved by both models. Absolute tool-call and control-step counts cover each full retained successful episode. Tool calls include robot requests, shell executions and explicit image views; control steps use the final episode count. Astra uses fewer tool calls on seven tasks and fewer control steps on six. These paired costs condition on shared success; the all-task distributions in Figure 17 include unsuccessful runs.
Figure 19 : Exploratory task-demand profiles. Bars show success rates with exact counts. Labels overlap and both models are evaluated on the same tasks within each group. All 84 tasks are scored for both models. Group membership is assigned from task requirements, not inferred from the outcome; these descriptive differences do not isolate component abilities.
targets : selected left/right XYZ (world metres), wrist pitch/roll/yaw (degrees from the downward reference), and gripper opening (0 closed, 1 open). Omitted fields hold measured call-start values. The planner moves the arm before applying a combined gripper change.
RoboLab
move_eef
targets.position[3] and/or quaternion_wxyz[4] : absolute flange pose in the robot-root frame. Omitted pose components hold measured values. steps=1..30 at 15 Hz.
RoboLab
set_gripper
targets.gripper_close : 0 opens, 1 closes; the target persists. The arm holds while the segment executes.
RoboLab
move_robot
Combines the same arm-pose and gripper fields in one native action. Closure starts with arm motion, not after arrival.
RoboCasa
move_eef
targets.eef_delta[6] : normalised translation and rotation-vector increments in the base frame. Scales are 0.05 m and 0.5 rad per step. steps=1..30 ; repetition repeats the increment.
Left/right XYZ in robot-root metres and left/right_quat_xyzw . Absolute EEF IK; omitted arm holds its measured pose. All tools in this group use steps=1..30 at 30 Hz.
BEHAVIOR-1K
move_base
base_vx , base_vy in local-body m/s (limits +/-0.3), and base_wz in rad/s (+/-0.5). Arms hold root-relative targets; omitted base velocities are zero.
BEHAVIOR-1K
set_grippers
left/right_gripper : continuous opening in [0,1], with 0 closed and 1 open. Targets persist; EEFs hold measured poses and base motion stops.
BEHAVIOR-1K
move_torso
trunk_qpos[4] : absolute trunk joint positions in radians, using the recorded joint order and limits. This differs from the normalised slide increment in RoboCasa.
BEHAVIOR-1K
move_robot
Combines arm poses, gripper openings, base velocities and trunk targets simultaneously. One control step counts once, regardless of how many components are commanded.
WheeledLab
observe
Returns the current allowed observation without stepping physics. Authored onboard courses expose front RGB, encoders and IMU; the call does not reveal map, world pose or checkpoint progress.
ANYmal joint-position offsets: qtarget=qdefault+0.5u radians, tracked by native PD.
OmniDrones
4
Four rotor inputs. -1 requests zero thrust; +1 requests maximum. Zero is half maximum steady-state thrust, not hover.
Appendix
Table 29
ID
Source / task identifier
Limit
Score
A
O
K
D
G
RoboDojo
01
hang_mugs
800
N
✗
✗
✗
✗
✗
02
sweep_blocks
1000
N
✗
✓
✗
✗
✗
03
pour_liquid_into_cup
400
N
✗
✗
✗
✗
✗
04
make_toast
1400
N
✗
✗
✗
✗
✗
05
store_laptop_and_headphones
800
N
✗
✗
✗
✗
✗
Appendix
Table 11 : Complete 84-task inventory and five-model outcomes.
Model
Elapsed time (min)
Robot requests
All tool calls
Median
IQR
Median
IQR
Median
IQR
Astra
30.9
[15.2, 61.2]
58
[25.5, 94]
70
[30, 138.8]
Opus 5.5
54
[24.5, 137.3]
37
[16, 87.2]
57.5
[31, 148]
Kimi K3
67.6
[21.3, 181.7]
42.5
[11.8, 70]
61.5
[18.5, 95]
DeepSeek
16
[11.1, 40.5]
22
[5, 59.2]
33
[18, 70]
Gemini
12.8
[10, 25.8]
11
[0, 36.8]
31
[15, 52.2]
Appendix
Table 12: Time and interaction use. Medians and interquartile ranges across tasks.
Model
Control steps
Shell calls
Image views
Median
IQR
Median
IQR
Median
IQR
Astra
627
[298, 1165]
4.5
[0, 21.2]
0
[0, 0]
Opus 5.5
458
[178.5, 913.8]
15
[0, 28]
0
[0, 10]
Kimi K3
449.5
[158, 918.5]
1
[0, 14.2]
0
[0, 1.2]
DeepSeek
347
[62, 742]
8.5
[0, 16]
2.5
[0, 8]
Gemini
315
[78.5, 625]
15
[9.8, 23]
0
[0, 2.2]
Appendix
Table 13: Physical execution and auxiliary-tool use. Medians and interquartile ranges across tasks.
Model
Tokens
Estimated USD
Elapsed hours
Astra
941M
9,912.80
54.6
Opus 5.5
1.63B
2,916.43
140.3
Kimi K3
2.00B
624.09
163.8
DeepSeek V4.1 Flash
907M
34.49
35.7
Gemini 3.8 Flash
303M
66.68
28.7
Appendix
Table 14: Aggregate consumption for the five-model evaluation. Monetary values are reported list-price estimates. Kimi token and cost totals extrapolate from 57 valid records to 84 tasks.
Figure 20 : Astra stacking. Recorded external-camera frames from astra-62 . The final shown frame still precedes the terminal release request. The complete run, rather than this frame alone, supplies the success label.
Figure 21 : Opus 5.5 juggling. Reviewer-camera frames from task-158 . The policy receives a native actor vector, not these rendered views. Hit counts and final-window validity come from the retained checker.
Figure 22 : Gemini conveyor matching. Recorded head-camera frames from task-272 . The decisive recovery is resolved by consecutive tool receipts at events 550–559, rather than inferred solely from sampled frames.
Figure 23 : DeepSeek volleyball 1v1. Reviewer video from task-350 . The result is the native terminal win for the controlled actor; surviving or touching the ball alone would not suffice.
Pattern
Trace evidence
Interpretation boundary
Sequential decomposition
Astra stacking names the placement order and reuses a stack centre after release/retraction.
An explicit plan is evidence of intent; the final checker confirms the completed arrangement.
Feedback-driven correction
Gemini revises a rejected 0.95 m lift to 0.85 m; the next receipt advances from step 501 to 508.
The consecutive request/receipt pair supports a local correction, not a causal effect of a general recovery policy.
Frequent dynamic feedback
Opus 5.5 juggling uses 274 robot requests over 800 control steps with repeated strike/descent phases.
This successful case does not establish that more calls universally improve success.
Calibration revision
DeepSeek changes its hover computation after reporting excessive rise.
A visible self-correction is distinct from independently verified correctness of its entire controller.
Contact/grasp reassessment
Kimi pouring repeatedly reports missed closure and changes approach or wrist orientation.
The agent notices difficulty; final failure should not be described as an unobserved or hallucinated success.
Pre-action analysis saturation
DeepSeek computer-use episode reaches a non-action boundary at step zero.
No robot motion is available for diagnosing physical control ability in that episode.
Appendix
Table 38
Task
Robot requests
All calls
Result steps
Conveyor matching
45 / 35
45 / 35
380 / 347
Plastic clutter sorting
41 / 104
46 / 107
784 / 1497
Three fruits on a plate
21 / 31
22 / 58
284 / 419
Cube left of bowl
17 / 18
24 / 31
286 / 239
White mug centring
18 / 22
20 / 42
280 / 341
Volleyball 1v1
30 / 152
30 / 157
135 / 498
Appendix
Table 16 : Matched successful tasks. Every numeric pair is Astra / Opus 5.5.
Figure 24 : Kimi pouring failure. Head-camera frames from task-25 . The bottle is upright initially and later lies on the table. Agent notes acknowledge failed acquisition; the retained episode ends unsuccessfully at step 400.
Figure 25 : Opus 5.5 stacking: historical continuation. Frames from the full retained video of task-82 . The later two views lie beyond the scored step-313 budget boundary; they illustrate continued work, not valid success within that budget.
Recent advances in multimodal foundation models have made them capable generalist physical agents for a range of manipulation tasks. However, successful operation in an unfamiliar environment may require an agent to seek task-relevant information through interaction when it is absent from the observations: it may need to determine where a relevant object is, inspect an unobserved property, or discover the effect of an unfamiliar tool. We thus introduce RoboQuest, a benchmark for goal-directed embodied exploration, where agents must actively acquire task-relevant information through physical interaction, use the resulting evidence to adapt subsequent actions, and autonomously decide when to commit to task completion. RoboQuest comprises ten mobile manipulation tasks centered on three forms of uncertainty: search, manipulation-based inspection, and interactive testing. We evaluate five frontier multimodal agents through a common visuomotor interface, as well as a π0.5 policy fine-tuned on the full-episode demonstrations we release. The best agent succeeds in only 23% of the episodes, and the fine-tuned policy almost never succeeds. Isolated tests of the execution skills the tasks are built from, with the hidden information supplied, show that the agents can carry out most of the required actions, and our failure analysis attributes only a minority of the failures to execution. Our failure analysis further finds that the agents often stop exploring too early as they make decisions before observing the required evidence for task completion. We also find that agents rarely prevent or repair the disturbances caused by their exploration. Moreover, learning by trial and error remains difficult for most models.
Spatial reasoning is a foundational capability for multimodal large language models (MLLMs) to perceive and operate within the physical world. However, existing benchmarks predominantly rely on passive evaluation (e.g., static VQA) or simulator-specific pipelines, failing to assess general interactive spatial understanding. We introduce SpatialWorld, a unified benchmark designed specifically for evaluating the interactive spatial understanding of multimodal agents in complex real-world tasks. Integrating eight heterogeneous simulation backends under a shared, simulator-agnostic protocol, SpatialWorld features 760 human-annotated tasks across diverse domains (e.g., household routines, travel, social collaboration). Agents must solve tasks under vision-only partial observability, actively gathering egocentric visual evidence and expressing decisions via a unified, text-based action interface native to MLLMs. For reliable evaluation, each task includes a human-validated initial state, a reference trajectory, and a terminal-state verifier. Evaluating 15 advanced agents reveals that robust spatial task solving remains challenging: the strongest model, GPT-5, achieves an average task success rate (TSR) of only 17.4%, while the leading open-source model, Qwen-3.5, reaches 14.1%. Further analysis exposes a clear mismatch between task success and execution efficiency, alongside substantial domain-specific performance variations. These bottlenecks in active exploration and long-horizon planning position SpatialWorld as a rigorous testbed for future spatial agents.
Hongcheng Gao, Hailong Qu, Jingyi Tang +18
Tsinghua University · Chongqing University · Peking University +6
General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot manipulation tasks. Rather than asking agents to submit task-level Python control programs or operate through high-level robot skills, LIBERO-Agent provides an interactive robotic environment where agents can select which observations to inspect, process them with their own tools, and issue native action commands. LIBERO-Agent integrates 200 tasks into a common interaction framework and provides a 30-task primary suite that separates perception, short-horizon execution, and long-horizon composition. Results reveal a pronounced reliability gap: while agents perform well on perception and easy short-horizon tasks, their performance degrades substantially on hard short-horizon and long-horizon tasks. Richer observations improve short-horizon manipulation, while demonstration benefits depend on the agent and format. Among these agents, GPT-6 Astra achieves the strongest overall performance. Further analysis shows its major advantage lies in mechanism interaction, especially when sustained physical contact is needed, while its remaining failures stem from cross-stage interference and geometric errors.
Zijie Diao, Yitong Chen, Sicheng Xie +7
Institute of Trustworthy Embodied AI, Fudan University · Shanghai Innovation Institute