General-purpose agents can plan, use tools, and revise their behavior from feedback, but it remains unclear whether these capabilities transfer from digital environments to embodied manipulation. To investigate this question, we introduce LIBERO-Agent, an agent-native benchmark for evaluating these agents in robot manipulation tasks. Rather than asking agents to submit task-level Python control programs or operate through high-level robot skills, LIBERO-Agent provides an interactive robotic environment where agents can select which observations to inspect, process them with their own tools, and issue native action commands. LIBERO-Agent integrates 200 tasks into a common interaction framework and provides a 30-task primary suite that separates perception, short-horizon execution, and long-horizon composition. Results reveal a pronounced reliability gap: while agents perform well on perception and easy short-horizon tasks, their performance degrades substantially on hard short-horizon and long-horizon tasks. Richer observations improve short-horizon manipulation, while demonstration benefits depend on the agent and format. Among these agents, GPT-6 Astra achieves the strongest overall performance. Further analysis shows its major advantage lies in mechanism interaction, especially when sustained physical contact is needed, while its remaining failures stem from cross-stage interference and geometric errors.
Figures & tables
Figure 1 : Overview of LIBERO-Agent: agent-native interaction, benchmark task suite, and key findings.
Agent
Perception
Short-Horizon
Long-Horizon
Overall
Resource Usage
Judg. Acc.
Task SR
Easy SR
Hard SR
Easy SC
Hard SC
Score
Avg. Time
Avg. Tokens
GPT-6 Astra
100.0%
100.0%
100.0%
40.0%
50.0%
22.0%
45.0
11.5 min
1.38M
GPT-5.6 Sol
90.0%
10.0%
40.0%
0.0%
16.7%
0.0%
10.8
22.0 min
3.44M
Claude Opus 5
100.0%
80.0%
60.0%
0.0%
16.7%
0.0%
19.3
22.0 min
5.49M
Claude Fable 5.1
90.0%
60.0%
40.0%
0.0%
16.7%
0.0%
15.8
22.4 min
3.01M
DeepSeek-V4.1-Flash
90.0%
0.0%
0.0%
0.0%
0.0%
0.0%
4.5
21.2 min
0.39M
Table 1 : Stable performance on the primary 30-task suite. Metrics follow the stable criterion in Section 4.1 . Judg. Acc.: target-identification accuracy; SR: success rate; SC: stage completion; M: million tokens. Score denotes the 100-point Performance Score. Bold marks the best task-performance result in each column.
Table 3
Figure 2 : Stage and failure taxonomy. (a–b) Stage types: Object Manipulation and Mechanism Interaction; (c–d) failure types: Local Execution and Cross-stage Interference.
Stage Progress (%)
Execution SR (%)
Failure Mode (%)
Agent
Reached
Success
Object Manipulation
Mechanism Interaction
Local Execution
Cross-stage Interference
GPT-6 Astra
65.0
36.7
66.7
94.1
23.5
52.9
GPT-5.6 Sol
43.3
13.3
66.7
22.2
61.1
27.8
Claude Opus 5
41.7
5.0
28.6
7.1
81.8
0.0
Claude Fable 5.1
26.7
6.7
20.0
37.5
83.3
0.0
Table 4 : Stage-level diagnosis on the hard long-horizon tasks. Reached and Success are computed over all stage opportunities. Object Manipulation and Mechanism Interaction report execution success rates after excluding failures attributed to cross-stage interference. Failure modes are percentages of reached-but-failed stages and may overlap.
Figure 3 : Sustained contact during mechanism interaction. (a) Mean sustained-contact ratio (SCR) across agents for entered opening operations. (b) Local opening success across SCR intervals. Error bars in (b) show 95% Wilson confidence intervals. SCR intervals are used only for visualization. Pearson correlation is computed over the original 64 opening operations.
Figure 4 : Residual failure modes of Astra. (a) A placement blocks access required by a later stage. (b) A subsequent action disturbs previously established state. (c) The intended operation is executed with an invalid object orientation.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Agent harness
Version
GPT-6 Astra
Codex
0.154.0
GPT-5.6 Sol
Codex
0.154.0
Claude Opus 5
Claude Code
2.1.139
Claude Fable 5.1
Claude Code
2.1.139
DeepSeek V4.1 Flash
DeepSeek Harness
0.1.5-rc.2
Kimi K3
Kimi Code
0.42.0
Appendix
Table 5 : Agent harnesses and versions used in our experiments.
ID
Primary information
Required physical outcome
T01
Object pose
Place the only upright marker in the collection bin.
T02
Spatial relation
Place the object between the other two objects in the basket.
T03
Color
Place the red object in the collection bin.
T04
Texture
Place the striped object in the collection bin.
T05
Instance marking
Place the object with exactly one square mark in the collection bin.
T06
Shape
Press the button paired with the small ball, among ball, cylinder, and cube candidates.
Appendix
Table 6 : Perception-focused tasks. The target-selection requirement is distinct from the complete physical outcome. The identifiers T01–T30 are shared across the task catalogs and execution examples.
ID
Split
Task
Success requirement
T11
Easy
Pick and place
Alphabet soup is in the basket.
T12
Hard
Push plate
The plate reaches the specified region in front of the stove.
T13
Easy
Press button
The red button is physically actuated.
T14
Easy
Close drawer
The cabinet’s top drawer reaches its closed state.
T15
Easy
Open drawer
The cabinet’s top drawer reaches its open state.
T16
Hard
Open microwave
The microwave door reaches its open state.
Appendix
Table 7 : Short-horizon manipulation tasks. Easy and hard refer to the evaluation split. A short-horizon task can contain several physical events, but its reported task success requires the complete operation.
ID
Split
K
Goals or stages
T21
Hard
5
Prepare the cooking area from text: moka pot on its stove, that stove on, frying pan on its stove, that stove on, and microwave open.
T22
Hard
5
Match the cooking goal image, using the same five physical goals as T21.
T23
Easy
8
Open and close the top, middle, and bottom drawers in order; reopen the occupied drawer; place butter inside it. Final closure is optional.
T24
Easy
4
Lift the tomato-sauce bottle, complete two distinct pours over the frying pan, and place the bottle in the bowl drainer.
T25
Hard
4
Move the left bowl to the empty plate, the middle bowl to the left plate, the right bowl to the middle plate, and the buffered bowl to the right plate. No bowl may touch the tabletop.
T26
Easy
3
Transfer tomato sauce, milk, and orange juice from the first cabinet to the second cabinet.
Appendix
Table 8 : Long-horizon task requirements and scoring units. K denotes the number of scored stages or goals. Cooking goals are evaluated at the terminal state without an imposed order.
Reference content
None
Video
Video + Traj.
Head/wrist RGB video
—
Yes
Yes
RGB contact sheets
—
Yes
Yes
Ordered native OSC action vectors
—
—
Yes
Reference depth, calibration, or robot-state stream
—
—
—
Private object states or checker progress
—
—
—
Appendix
Table 9 : Contents of the demonstration bundle. These columns describe historical reference data; the live observation profile is held fixed across the context comparison.
Sequence
Control batches
Control steps
Initial error
Final error
1
6
60
121.88∘
0.76∘
2
4
42
89.68∘
2.26∘
3
5
61
87.99∘
0.14∘
4
4
46
90.42∘
0.61∘
Appendix
Table 10 : Repeated orientation corrections in the successful T23 episode. Each sequence uses fresh state before successive control batches and retains a common target orientation. Angular error is measured from the recorded end-effector quaternion to the agent-selected target.
Probe
Sensor Fz (N)
EEF height (m)
Jaw width (m)
Left mug
−8.635
0.207
0.0613
Middle mug
−12.195
0.180
0.0622
Right mug
−6.861
0.204
0.0616
Appendix
Figure 8 : Active sensing in a successful weighing episode. The same lift-and-hold command pattern is issued for each mug, followed by a state and proprioception read. Values are rounded from the returned observations. Mug locations refer to the initial head-camera view; height is in the robot-base frame and force is in the sensor frame.
Agent
T01 Upright marker
T02 Between objects
T03 Red object
T04 Striped object
T05 Marked instance
T06 Ball shape
T07 Cola drink
T08 Heaviest mug
T09 Highest reflectance
T10 Largest cylinder
GPT-6 Astra
3/3
3/3
3/3
3/3
3/3
3/3
3/3
3/3
3/3
3/3
GPT-5.6 Sol
1/3
0/3
2/3
1/3
2/3
3/3
1/3
1/3
1/3
2/3
Claude Opus 5
3/3
2/3
3/3
3/3
3/3
3/3
3/3
2/3
3/3
3/3
Claude Fable 5.1
2/3
2/3
1/3
3/3
3/3
3/3
3/3
2/3
3/3
3/3
DeepSeek V4.1 Flash
1/3
1/3
1/3
1/3
0/3
1/3
0/3
1/3
1/3
2/3
Kimi K3
1/3
0/3
0/3
2/3
0/3
3/3
1/3
1/3
1/3
2/3
Appendix
Table 11 : Perception: full-task successes out of three rollouts.
Agent
T11 Pick soup
T12 Push plate
T13 Press button
T14 Close drawer
T15 Open drawer
T16 Open microwave
T17 Turn on stove
T18 Insert tube
T19 Wipe spill
T20 Pour wine
GPT-6 Astra
3/3
3/3
3/3
3/3
3/3
3/3
2/3
0/3
2/3
3/3
GPT-5.6 Sol
3/3
0/3
2/3
3/3
0/3
0/3
2/3
0/3
0/3
2/3
Claude Opus 5
3/3
1/3
3/3
2/3
3/3
0/3
1/3
0/3
1/3
2/3
Claude Fable 5.1
3/3
0/3
3/3
2/3
1/3
0/3
1/3
0/3
1/3
2/3
DeepSeek V4.1 Flash
1/3
0/3
1/3
0/3
0/3
0/3
1/3
0/3
1/3
1/3
Kimi K3
2/3
0/3
1/3
2/3
1/3
0/3
1/3
0/3
1/3
1/3
Appendix
Table 12 : Short-horizon manipulation: full-task successes out of three rollouts.
Figure 9 : Perception rollouts on T04. The instruction is to place the striped object in the collection bin. The successful execution transfers the striped object to the bin. In the unsuccessful execution, the robot also transports the striped object toward the bin, but its final placement does not satisfy the task’s success check. The wrist view exposes the texture cues and the objects near the gripper, while the head view shows their positions relative to the bin.
Figure 10 : Short-horizon manipulation rollouts. Top: a successful T16 execution opens the microwave door. The selected frames show approach, gripper reorientation, and door opening. Bottom: an unsuccessful T19 execution attempts to grasp the sponge but leaves the spill unwiped when the time budget expires. The brown spill remains visible throughout the selected frames. The paired views show the gripper configuration and its relation to the handle or sponge.
Figure 11 : Long-horizon rollouts. Top: a successful T23 execution opens and closes the drawers to inspect their contents, reopens the occupied drawer, and places butter inside. The selected frames show the top, middle, and bottom drawer openings, the occupied drawer reopened, butter placement, and the final configuration. All eight required stages are completed. Bottom: an unsuccessful T27 execution opens the middle drawer and places cookies inside, but does not complete the subsequent chocolate placement before the time budget expires. It completes two of four required stages; the final frame shows the chocolate outside the drawer.