CodeActionBench: Evaluating Agentic Code-as-Policy for Embodied Manipulation
Organizations: Nanyang Technological University, Singapore · Independent Researcher
Abstract
How well can general-purpose multimodal models turn visual understanding and reasoning into embodied manipulation via executable code? We introduce CodeActionBench, a benchmark of 25 manipulation tasks that evaluates this capability through agentic Code-as-Policy. Without task-specific fine-tuning, demonstrations, external specialist perception or grasp modules, privileged scene state, or predefined task policies, agents should select visual evidence, form task-relevant 3D estimates, construct manipulation targets, and iteratively execute and revise their policies. A shared robot API provides RGB observations, calibrated geometric operations, robot feedback, and bounded motion, leaving task-dependent decisions to the evaluated agent. Fixed task instances, resource budgets, and a hidden physical-outcome verifier support controlled comparisons across models and harness configurations. Extensive evaluations across nine configurations and 675 attempts achieve success rates ranging from 2.7% to 73.3%. The strongest configuration, GPT-6 Astra with Codex CLI, solves 22 of 25 tasks at least once in three attempts, demonstrating the best performance while still leaving substantial room for improvement. Trajectory analyses reveal difficulties in spatial alignment, object retention, and completion judgment, including task failures despite successfully completed motions. CodeActionBench provides a controlled testbed for measuring how general-purpose models translate their capabilities into manipulation behavior and for examining typical failure scenarios in that process.
Figures & tables
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
| Task | Seed | Scoring | (s) | Solution API calls | Solution sim (s) | |
|---|---|---|---|---|---|---|
| Tool use and contact processes (4 tasks) | ||||||
| beat_block_hammer | 0 | latch | 42 | 60 | 40 | 28 |
| There is a hammer and a block on the table, use the arm to grab the hammer and beat the block. | ||||||
| click_bell | 0 | latch | 28 | 60 | 20 | 23 |
| Click the bell’s top center on the table. Use the arm on the bell’s side and keep its gripper closed. | ||||||
| press_stapler | 0 | latch | 28 | 60 | 20 | 23 |
| Tool | Agent-supplied inputs | Returned information |
|---|---|---|
| get_world_frame | None | World-axis directions and coordinate conventions. |
| get_embodiment | None | Robot dimensions, gripper geometry, and TCP/camera-mount conventions. |
| get_camera_info | Camera | Camera intrinsics, extrinsics, and image dimensions. |
| get_arm_pose | Arm | Measured end-effector and TCP poses, including orientation axes. |
| get_gripper_state | Arm | Gripper drive readback, physical finger gap, and measurement availability. |
| get_robot_state | Optional arm selection | Combined arm poses, gripper state, contact feedback, and joint state. |
| Tool | Agent-supplied inputs | Returned information |
|---|---|---|
| ray | Observation and pixel | World-frame origin and direction of the viewing ray. |
| project | Observation and 3D point | Pixel coordinates of the point in the selected view. |
| plane_intersect | Observation, pixel, plane point/normal, and uncertainty | Estimated 3D intersection and propagated uncertainty. |
| capture_motion_pair | Arm and displacement | Wrist RGB views before/after motion, measured camera displacement, and execution feedback. |
| triangulate_correspondence | Motion pair and corresponding pixels | Estimated 3D point, uncertainty, and triangulation diagnostics. |
| scale_from_object_size | Observation, bounding box, size prior, and extent axis | Coarse depth estimate from the selected image extent and size prior. |
| Tool | Agent-supplied inputs | Returned information |
|---|---|---|
| reach_tcp | Arm, target position, and optional orientation | Measured TCP pose and planning/execution outcome. |
| reach_both_tcp | Target positions and optional orientations for both arms | Measured poses of both arms and coordinated execution outcome. |
| move_delta | Arm, world-frame displacement, and optional path mode | Measured TCP pose and execution outcome for the relative motion. |
| move_both_delta | World-frame displacements for both arms | Measured poses of both arms and coordinated execution outcome. |
| probe_contact_along | Arm, direction, travel bound, and step size | Contact change, measured TCP and finger-link poses, and stopping reason. |
| set_gripper | Arm and normalized opening | Drive readback, physical finger gap, contact feedback, and execution outcome. |
| Configuration | Model identifier | Reasoning | Output tokens |
|---|---|---|---|
| GPT-6 Astra (Codex CLI) | gpt-6-astra | High | Harness-managed |
| Claude Opus 5 (Reference) | claude-opus-5 | High | 128,000 |
| Claude Opus 5 (Claude Code) | claude-opus-5 | High | Harness-managed |
| Gemini 3.6 Flash (Reference) | gemini-3.6-flash | High | 65,536 |
| GPT-5.6 Sol (Reference) | gpt-5.6 | High | 128,000 |
| Qwen 3.8 Max (Reference) | qwen3.8-max | Medium | 131,072 |
| Configuration | Program share (%) | Program errors count (%) |
|---|---|---|
| GPT-6 Astra (Codex CLI) | 92.6 | 10/2,167 ( 0.46% ) |
| Claude Opus 5 (Reference) | 85.6 | 65/2,316 (2.81%) |
| Claude Opus 5 (Claude Code) | 85.7 | 76/2,387 (3.18%) |
| Gemini 3.6 Flash (Reference) | 87.4 | 159/2,893 (5.50%) |
| GPT-5.6 Sol (Reference) | 18.9 | 30/784 (3.83%) |
| Qwen 3.8 Max (Reference) | 35.2 | 91/1,020 (8.92%) |
| Configuration | Success (/75) | Coverage (/25) | 0/3 | 1/3 | 2/3 | 3/3 |
|---|---|---|---|---|---|---|
| GPT-6 Astra (Codex CLI) | 55 | 22 | 3 | 4 | 3 | 15 |
| Claude Opus 5 (Reference) | 37 | 19 | 6 | 8 | 4 | 7 |
| Claude Opus 5 (Claude Code) | 34 | 14 | 11 | 3 | 2 | 9 |
| Gemini 3.6 Flash (Reference) | 15 | 8 | 17 | 4 | 1 | 3 |
| GPT-5.6 Sol (Reference) | 12 | 9 | 16 | 7 | 1 | 1 |
| Qwen 3.8 Max (Reference) | 11 | 6 | 19 | 3 | 1 | 2 |
| Calls | Median time (s) | Inference cost (USD) | |||||
|---|---|---|---|---|---|---|---|
| Configuration | Mean | Sim. | Wall | Wait | Total | Median | IQR |
| GPT-6 Astra (Codex CLI) | 31.2 | 32.7 | 477.2 | – | 252.87 | 3.09 | 2.06–4.35 |
| Claude Opus 5 (Reference) | 36.1 | 45.3 | 1,193.1 | 910 | 297.84 | 3.71 | 2.48–5.29 |
| Claude Opus 5 (Claude Code) | 37.1 | 57.7 | 1,359.4 | – | 380.09 | 4.88 | 3.34–6.35 |
| Gemini 3.6 Flash (Reference) | 44.1 | 45.0 | 781.6 | 501 | 52.00 | 0.61 | 0.36–0.88 |
| GPT-5.6 Sol (Reference) | 55.2 | 28.2 | 636.4 | 447 | 259.85 | 3.04 | 2.10–4.54 |
| Task | SG | SP | Selected subgoals |
|---|---|---|---|
| beat_block_hammer | 1 | 2 | (1) Hammer head aligned with and contacting the block. |
| blocks_ranking_size | 3 | 5 | (1) Large/middle blocks aligned. (2) Middle/small blocks aligned. (3) Both alignments and size order hold together. |
| click_bell | 1 | 1 | (1) Required bell contact with the selected gripper closed, or its retained success event. |
| dump_bin_bigbin | 3 | 2 | (1) Small bin raised to at least 1.0 m. (2) All five balls in the verifier’s height band. (3) The lift and all ball-height conditions hold together. |
| grab_roller_dual_contact | 4 | 4 | (1) Left gripper contacts the roller. (2) Right gripper contacts the roller. (3) Roller above 0.80 m. (4) Both closed grippers contact the raised roller. |
| handover_block | 1 | 3 | (1) Moved block’s bottom point aligned with and seated on the support’s top point. |
| Checkpoint | Arm | Spatial condition (cm) | Prerequisite |
|---|---|---|---|
| beat_block_hammer | |||
| 1. Approach hammer | Either | , , . | – |
| 2. Strike | Either | . | 1 |
| blocks_ranking_size | |||
| 1. Approach block 1 | Either | , . | – |
| 2. Approach block 2 | Either | , . | – |
| Configuration | Non-completion count (rate) | Planner count | Stall count | Deviation count | Allowance count |
|---|---|---|---|---|---|
| GPT-6 Astra (Codex CLI) | 305/2,055 ( 14.8% ) | 187 | 110 | 3 | 5 |
| Claude Opus 5 (Reference) | 410/2,323 (17.6%) | 171 | 209 | 6 | 24 |
| Claude Opus 5 (Claude Code) | 444/2,450 (18.1%) | 191 | 221 | 4 | 28 |
| Gemini 3.6 Flash (Reference) | 665/3,029 (22.0%) | 400 | 257 | 3 | 5 |
| GPT-5.6 Sol (Reference) | 310/1,611 (19.2%) | 124 | 166 | 1 | 19 |
| Qwen 3.8 Max (Reference) | 703/2,385 (29.5%) | 393 | 265 | 3 | 42 |
| Configuration | Stopping condition | Completion claim | ||||||
|---|---|---|---|---|---|---|---|---|
| Done | Tool | Phys. | Other | Correct | Over | Under | Absent | |
| GPT-6 Astra (Codex CLI) | 72 | 0 | 3 | 0 | 68 | 3 | 1 | 3 |
| Claude Opus 5 (Reference) | 63 | 0 | 12 | 0 | 50 | 13 | 0 | 12 |
| Claude Opus 5 (Claude Code) | 54 | 2 | 19 | 0 | 42 | 11 | 1 | 21 |
| Gemini 3.6 Flash (Reference) | 37 | 17 | 21 | 0 | 18 | 19 | 0 | 38 |
| GPT-5.6 Sol (Reference) | 21 | 52 | 2 | 0 | 12 | 9 | 0 | 54 |