How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, external specialist perception or grasp modules, privileged scene state, or predefined task policies, agents should select visual evidence, form task-relevant 3D estimates, construct manipulation targets, and iteratively execute and revise their policies. A shared robot API provides RGB observations, calibrated geometric operations, robot feedback, and bounded motion, leaving task-dependent decisions to the evaluated agent. Fixed task instances, resource budgets, and a hidden physical-outcome verifier support controlled comparisons across models and harness configurations. Extensive evaluations across nine configurations and 675 attempts achieve success rates ranging from 2.7% to 73.3%. The strongest configuration, GPT-6 Astra with Codex CLI, solves 22 of 25 tasks at least once in three attempts, demonstrating the best performance while still leaving substantial room for improvement. Trajectory analyses reveal difficulties in spatial alignment, object retention, and completion judgment, including task failures despite successfully completed motions. CodeActionBench provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process.
Figures & tables
Figure 1: Robot-control paradigms and agentic Code-as-Policy. (a) VLA learns action policies from large-scale robot datasets; Code-as-Policy generates programs that compose robot APIs to perform tasks. (b) CodeActionBench evaluates agents that autonomously construct task-relevant 3D estimates, manipulation targets, and policies, without depth sensing, external perception/grasp specialists, or privileged scene state.
Figure 2: Benchmark system and information boundaries. Each evaluated agent comprises a model and its harness. Task instances, the robot API, resource budgets, and verification rules are fixed across configurations. A hidden verifier scores physical outcomes after termination without returning its verdict to the agent, separating task outcome from the agent’s completion judgment.
Figure 3: Task performance. (A) Success rate over 75 attempts. (B) Percentage of the 25 fixed instances solved at least once in three independent attempts.
Figure 4: Success rate against inference cost and wall time. Each point represents one configuration’s 75 attempts. Both panels show success rate vertically; the horizontal axes are (A) total API-equivalent inference cost and (B) median wall time per attempt.
Figure 5: Subgoal and spatial checkpoint coverage across configurations. Panel A reports final coverage in descending order of subgoal checkpoint coverage. Panel B shows spatial checkpoint coverage over charged calls within each attempt. Each measure is averaged over three attempts per task, and all 25 tasks receive equal weight.
Figure 6: Task outcomes and failure stages. The horizontal axis counts attempts, from 0 to 75 per configuration. Configurations are ordered by task success rate. Each attempt contributes once to success or the earliest supported failure stage; unresolved cases remain separate.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Task
Seed
Scoring
Btool
Bsim (s)
Solution API calls
Solution sim (s)
Tool use and contact processes (4 tasks)
beat_block_hammer
0
latch
42
60
40
28
There is a hammer and a block on the table, use the arm to grab the hammer and beat the block.
click_bell
0
latch
28
60
20
23
Click the bell’s top center on the table. Use the arm on the bell’s side and keep its gripper closed.
press_stapler
0
latch
28
60
20
23
Appendix
Table 3: Tasks and evaluation budgets. Each of the 25 fixed task instances is listed with a representative instruction, scene seed, scoring rule, and reference-solution resource use. Btool and Bsim bound charged calls and simulated seconds. Scoring uses required events during execution (latch) or physical conditions at termination (final).
Tool
Agent-supplied inputs
Returned information
get_world_frame
None
World-axis directions and coordinate conventions.
get_embodiment
None
Robot dimensions, gripper geometry, and TCP/camera-mount conventions.
get_camera_info
Camera
Camera intrinsics, extrinsics, and image dimensions.
get_arm_pose
Arm
Measured end-effector and TCP poses, including orientation axes.
get_gripper_state
Arm
Gripper drive readback, physical finger gap, and measurement availability.
get_robot_state
Optional arm selection
Combined arm poses, gripper state, contact feedback, and joint state.
Appendix
Table 4: Robot-state measurements and RGB observations. State queries and camera tools provide feedback for spatial estimation and action assessment. Contact measurements identify the contacting fingers but not the contacted objects.
Figure 7: Visual feedback for spatial estimation and target construction. The panels show measured gripper geometry, selected image points, a proposed gripper pose, and a comparison of pose candidates. The agent supplies the image points and proposed poses.
Tool
Agent-supplied inputs
Returned information
ray
Observation and pixel
World-frame origin and direction of the viewing ray.
project
Observation and 3D point
Pixel coordinates of the point in the selected view.
plane_intersect
Observation, pixel, plane point/normal, and uncertainty
Estimated 3D intersection and propagated uncertainty.
capture_motion_pair
Arm and displacement
Wrist RGB views before/after motion, measured camera displacement, and execution feedback.
triangulate_correspondence
Motion pair and corresponding pixels
Estimated 3D point, uncertainty, and triangulation diagnostics.
scale_from_object_size
Observation, bounding box, size prior, and extent axis
Coarse depth estimate from the selected image extent and size prior.
Appendix
Table 5: Spatial estimation and target-pose construction. Tools convert agent-selected pixels, geometric assumptions, and gripper axes into 3D estimates and candidate poses. Projection and reachability checks support inspection before execution.
Tool
Agent-supplied inputs
Returned information
reach_tcp
Arm, target position, and optional orientation
Measured TCP pose and planning/execution outcome.
reach_both_tcp
Target positions and optional orientations for both arms
Measured poses of both arms and coordinated execution outcome.
move_delta
Arm, world-frame displacement, and optional path mode
Measured TCP pose and execution outcome for the relative motion.
move_both_delta
World-frame displacements for both arms
Measured poses of both arms and coordinated execution outcome.
probe_contact_along
Arm, direction, travel bound, and step size
Contact change, measured TCP and finger-link poses, and stopping reason.
set_gripper
Arm and normalized opening
Drive readback, physical finger gap, contact feedback, and execution outcome.
Appendix
Table 7: TCP motion, contact probing, and gripper control. Actions execute agent-specified targets and return measured robot state and execution outcomes, including partial motion before a stop. TCP denotes the tool center point.
Configuration
Model identifier
Reasoning
Output tokens
GPT-6 Astra (Codex CLI)
gpt-6-astra
High
Harness-managed
Claude Opus 5 (Reference)
claude-opus-5
High
128,000
Claude Opus 5 (Claude Code)
claude-opus-5
High
Harness-managed
Gemini 3.6 Flash (Reference)
gemini-3.6-flash
High
65,536
GPT-5.6 Sol (Reference)
gpt-5.6
High
128,000
Qwen 3.8 Max (Reference)
qwen3.8-max
Medium
131,072
Appendix
Table 8: Detailed agent configurations. Each row contains 75 attempts. Reasoning gives the requested effort; output is the configured per-request token limit. For Claude Code and Codex CLI, the vendor harness manages the output limit. The benchmark sets no additional token cap.
Configuration
Program share (%)
Program errors count (%)
GPT-6 Astra (Codex CLI)
92.6
10/2,167 ( 0.46% )
Claude Opus 5 (Reference)
85.6
65/2,316 (2.81%)
Claude Opus 5 (Claude Code)
85.7
76/2,387 (3.18%)
Gemini 3.6 Flash (Reference)
87.4
159/2,893 (5.50%)
GPT-5.6 Sol (Reference)
18.9
30/784 (3.83%)
Qwen 3.8 Max (Reference)
35.2
91/1,020 (8.92%)
Appendix
Table 9: Program use and execution errors over 75 attempts per configuration. Program share is the proportion of charged calls submitted through run_code . Errors include Python exceptions and sandbox rejections; each entry gives the count over returned programs and its rate. Bold marks the highest program share and lowest error rate.
Configuration
Success (/75) ↓
Coverage (/25)
0/3
1/3
2/3
3/3
GPT-6 Astra (Codex CLI)
55
22
3
4
3
15
Claude Opus 5 (Reference)
37
19
6
8
4
7
Claude Opus 5 (Claude Code)
34
14
11
3
2
9
Gemini 3.6 Flash (Reference)
15
8
17
4
1
3
GPT-5.6 Sol (Reference)
12
9
16
7
1
1
Qwen 3.8 Max (Reference)
11
6
19
3
1
2
Appendix
Table 10: Attempt success and task coverage across three attempts. Each configuration makes three attempts on each of 25 fixed task instances. Success counts successful attempts, and Coverage counts tasks solved at least once. Columns 0/3–3/3 count tasks by their number of successful attempts.
Figure 8: Task-level success across nine agent configurations. Each cell reports successful attempts out of three for one task and configuration. All configurations use the same fixed scene for each of the 25 tasks. Configurations without a harness label use the reference harness.
Calls
Median time (s)
Inference cost (USD)
Configuration
Mean
Sim.
Wall
Wait
Total
Median
IQR
GPT-6 Astra (Codex CLI)
31.2
32.7
477.2
–
252.87
3.09
2.06–4.35
Claude Opus 5 (Reference)
36.1
45.3
1,193.1
910
297.84
3.71
2.48–5.29
Claude Opus 5 (Claude Code)
37.1
57.7
1,359.4
–
380.09
4.88
3.34–6.35
Gemini 3.6 Flash (Reference)
44.1
45.0
781.6
501
52.00
0.61
0.36–0.88
GPT-5.6 Sol (Reference)
55.2
28.2
636.4
447
259.85
3.04
2.10–4.54
Appendix
Table 11: Resource use over 75 attempts per configuration. Calls are averaged per attempt; simulated, wall, and provider-wait times are medians. Costs are API-equivalent USD, with median and interquartile range (IQR) computed per attempt. Waiting time is available only for the reference harness. Bold marks the lowest total cost and median wall time.
Figure 9: Task outcomes and confirmed checkpoint progress. Each configuration has 75 attempts. The bars separate task success from failed attempts with at least one subgoal attained, spatial matches only, or no confirmed checkpoint. Configurations are ordered by the number of attempts with task success or at least one confirmed checkpoint.
Task
SG
SP
Selected subgoals
beat_block_hammer
1
2
(1) Hammer head aligned with and contacting the block.
blocks_ranking_size
3
5
(1) Large/middle blocks aligned. (2) Middle/small blocks aligned. (3) Both alignments and size order hold together.
click_bell
1
1
(1) Required bell contact with the selected gripper closed, or its retained success event.
dump_bin_bigbin
3
2
(1) Small bin raised to at least 1.0 m. (2) All five balls in the verifier’s height band. (3) The lift and all ball-height conditions hold together.
grab_roller_dual_contact
4
4
(1) Left gripper contacts the roller. (2) Right gripper contacts the roller. (3) Roller above 0.80 m. (4) Both closed grippers contact the raised roller.
handover_block
1
3
(1) Moved block’s bottom point aligned with and seated on the support’s top point.
Appendix
Table 12: Subgoal checkpoints and annotation counts for all 25 tasks. Numbered descriptions list all 55 selected subgoals. The count columns give subgoal (SG) and spatial (SP) checkpoints; the spatial total is 72.
Checkpoint
Arm
Spatial condition (cm)
Prerequisite
beat_block_hammer
1. Approach hammer
Either
∣Δx∣≤2 , ∣Δy∣≤6 , Δz∈[−3,4] .
–
2. Strike
Either
d3≤5 .
1
blocks_ranking_size
1. Approach block 1
Either
dxy≤2 , Δz∈[−3,4] .
–
2. Approach block 2
Either
dxy≤2 , Δz∈[−3,4] .
–
Appendix
Table 13: Spatial rules for all 72 checkpoints. Distances are in centimetres. Static regions use expert TCP references; moving-reference rules are explained above. L/R denotes arm selection. Prerequisites refer to earlier actions in the same attempt. For relative lifts, zo,0 is initial object height and zapproach is the same arm’s matched approach height.
Figure 10: Observed failure stages and action patterns. Each of the 491 failed attempts appears once. The inner ring shows the five failure stages. The outer ring divides each stage by the observed action patterns.
Configuration
Non-completion count (rate)
Planner count
Stall count
Deviation count
Allowance count
GPT-6 Astra (Codex CLI)
305/2,055 ( 14.8% )
187
110
3
5
Claude Opus 5 (Reference)
410/2,323 (17.6%)
171
209
6
24
Claude Opus 5 (Claude Code)
444/2,450 (18.1%)
191
221
4
28
Gemini 3.6 Flash (Reference)
665/3,029 (22.0%)
400
257
3
5
GPT-5.6 Sol (Reference)
310/1,611 (19.2%)
124
166
1
19
Qwen 3.8 Max (Reference)
703/2,385 (29.5%)
393
265
3
42
Appendix
Table 14: Non-completed robot actions. Each row covers 75 attempts and includes actions within programs. Non-completion reports the number of incomplete actions over all actions, followed by the rate. The remaining columns separate planner refusals, stalls, trajectory deviations, and motion-allowance stops. Bold marks the lowest rate.
Configuration
Stopping condition
Completion claim
Done
Tool
Phys.
Other
Correct
Over
Under
Absent
GPT-6 Astra (Codex CLI)
72
0
3
0
68
3
1
3
Claude Opus 5 (Reference)
63
0
12
0
50
13
0
12
Claude Opus 5 (Claude Code)
54
2
19
0
42
11
1
21
Gemini 3.6 Flash (Reference)
37
17
21
0
18
19
0
38
GPT-5.6 Sol (Reference)
21
52
2
0
12
9
0
54
Appendix
Table 15: Attempt termination and completion claims. Each configuration has 75 attempts. Correct denotes agreement with the verifier. Over and Under denote incorrect success and failure claims, respectively. Absent indicates no Boolean claim. Bold marks the most frequent stopping condition and completion-claim category within each row.
"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an open-access framework for systematically studying Code-as-Policy agents in robot manipulation. At its core is CaP-Gym, an interactive environment in which agents control robots by synthesizing and executing programs that compose perception and control primitives. Building on this foundation, CaP-Bench evaluates frontier language and vision-language models across varying levels of abstraction, interaction, and perceptual grounding. Across 12 models, CaP-Bench reveals a consistent trend: performance improves with human-crafted abstractions but degrades as these priors are removed, exposing a dependence on designer scaffolding. At the same time, we observe that this gap can be mitigated through scaling agentic test-time computation--through multi-turn interaction, structured execution feedback, visual differencing, automatic skill synthesis, and ensembled reasoning--substantially improves robustness even when agents operate over low-level primitives. These findings allow us to derive CaP-Agent0, a training-free framework that recovers human-level reliability on several manipulation tasks in simulation and on real embodiments. We further introduce CaP-RL, showing reinforcement learning with verifiable rewards improves success rates and transfers from sim2real with minimal gap. Together, CaP-X provides a principled, open-access platform for advancing embodied coding agents.
Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
We demonstrate that a general-purpose agent can directly drive a physical robot throughout task execution without any task-specific or environment-specific training. We introduce Agent as Policy (AGP), which places task planning and execution under the agent's control. Given a task and a robot interface, the agent interprets visual evidence, writes executable programs, issues motion commands, and revises its actions in response to physical outcomes. This brings the agent's reasoning and programming capabilities into continuous interaction with the physical world. We study AGP across multiple real-world manipulation tasks spanning precision manipulation, dynamic motions, and deformable objects. These include assembly from human videos, block construction from goal images, dice flipping, targeted throwing, and bimanual towel folding. Across assembly, block construction, and dice flipping, AGP succeeds in at least eight of ten trials for each evaluated task configuration. We further study efficiency through task experience accumulation and find that reusing saved procedures and programs shortens execution time across repeated trials. These findings support a path for general-purpose agents to act as robot policies, extending their autonomy to physical manipulation through runtime reasoning, programming, and interaction.