Systematically Exploring the Capabilities of GPT-6 Astra as Embodied Policies
Organizations: Galbot Team
Abstract
GPT-6 Astra exhibits a remarkable ability to generate numerical robot actions, extending its role beyond high-level planning. To assess Astra's capabilities as general-purpose embodied policies, we conduct comprehensive evaluations across six domains, examining direct control, cooperation with learned policies, and feedback-driven adaptation. In gripper manipulation, Astra can correct task targets and prepare contact conditions for subsequent policy execution; hybrid control with π0.5 achieves 48% success on the evaluated RoboDojo subset. In dexterous manipulation, hybrid control achieves 50% success in ten experience-guided DexJoCo trials, while direct in-hand control struggles to coordinate finger contacts. In mobile manipulation, hybrid control reaches 38.7% success on the evaluated RoboCasa365. In navigation, Astra leads our local comparisons, reaching 92% success on RxR instruction following and 82% on HM3D object search, although search incurs substantial detours. In locomotion, dense motion-reference generation remains unreliable: none of five sequential attempts on a single obstacle course reaches the goal, despite improvements in stability and forward progress. In humanoid loco-manipulation, Astra exceeds baseline methods on 13 of 30 HumanoidBench tasks with pretrained whole-body controllers. These findings reveal a gap between useful task decisions and reliable physical control. Inference latency further constrains practical control: across 50 RoboDojo instances per condition, policy-assisted and direct control consume 624.8 million and 1.132 billion tokens. A 30-second locomotion run requires 250 model calls averaging 39.86 seconds each, with physics paused during inference.
Figures & tables
| Domain | Tasks and sample sizes | What is compared |
| Gripper manipulation | RoboDojo and RoboLab each cover 10 tasks with 5 results per task and method. | RoboDojo compares paired Direct and Hybrid runs with published policy scores. RoboLab compares Direct and Hybrid with three zero-shot policies. Control settings and evaluation procedures appear in Appendix B . |
| Dexterous manipulation | 10 manipulation tasks with 5 cases each; 4 in-hand tasks with 5 initial states each. | Manipulation compares Direct, Hybrid, and a standalone policy on the same cases. In-hand control compares Astra and task-specific RL from identical physical states, with different observations and action rates. |
| Mobile manipulation | 15 RoboCasa tasks with 5 episodes per task and method. | Direct, Hybrid, and a locally evaluated standalone policy share initial states and 20-step action segments. The Astra conditions differ in policy access, context management, and feedback guards. |
| Navigation | 4 dataset subsets with 50 episodes per system in each. | Astra and released navigation policies use the same episode lists, scoring rules, and action budgets. Camera views and the conversion of outputs to navigation actions differ between systems. |
| Locomotion | 5 sequential Astra attempts on one course. Two reference interfaces, each tested on 6 courses in 2 physics backends: 12 rollouts per interface. | Astra and PASSAGE [ 26 ] share the initial state, goal, and frozen tracker. Separate analytic tests compare five-point and whole-body reference interfaces without Astra. |
| Humanoid loco-manipulation | HumanoidBench covers 30 tasks. SIMPLE covers 6 L2 tasks with 10 scenes each. | HumanoidBench compares Astra returns with published DreamerV3, TD-MPC2, and SAC results. SIMPLE uses the official codebase and protocol for comparison. |
| Reweighted public references | Our evaluation | ||||||
| Task | DM0.5 | Galaxea | Xiaomi R1 | OpenWAM | Hybrid | Direct | |
| Organize the table | 44.00 | 46.33 | 57.67 | 62.50 | 23.33 | 60.00 | 30.00 |
| Classify by language | 0.47 | 1.07 | 2.00 | 1.33 | 0.60 | 38.00 | 60.00 |
| Imitate a sorting sequence | 1.80 | 1.67 | 2.50 | 2.90 | 1.60 | 53.00 | 0.00 |
| Arrange the largest number | 7.85 | 4.11 | 8.56 | 4.36 | 2.29 | 50.00 | 57.00 |
| Pack objects into a box | 14.72 | 17.12 | 18.69 | 20.83 | 18.36 | 50.00 | 50.00 |
| Task | Direct | Hybrid | |
| Pot lift and hold | 40.0 | 44.0 | 62.0 |
| Headphones in box | 42.0 | 22.0 | 42.0 |
| Toy retrieval | 100.0 | 22.0 | 100.0 |
| Mug hanging | 62.0 | 4.0 | 82.0 |
| Mahjong tile storage | 48.0 | 18.0 | 88.0 |
| Bottles/cans sorting | 72.0 | 14.0 | 100.0 |
| Task | Controller | Steps | Full | Budget | Drop | At-goal | Error (rad) | Speed MAE |
| Cylinder, 20 s | Astra Direct | 1,749 | 3/5 | 1/5 | 1/5 | 0.51% | 1.626 | 0.960 |
| RL | 2,000 | 5/5 | – | 0/5 | 76.90% | 0.173 | 0.276 | |
| Cuboid, 10 s | Astra Direct | 1,000 | 5/5 | 0/5 | 0/5 | 4.40% | 0.742 | 0.143 |
| RL | 1,000 | 5/5 | – | 0/5 | 63.50% | 0.096 | 0.063 |
| Task | Controller | Endpoint | Position (mm) | Rotation (deg) | Success |
| Translation | Astra Direct | 15 s | 59.05 | – | 1/5 |
| RL | 15 s | 17.29 | – | 4/5 | |
| Translation + rotation | Astra Direct | 120 decisions | 47.35 | 32.72 | 0/5 |
| RL | matched steps | 16.22 | 8.31 | 4/5 | |
| RL | 15 s | 16.08 | 3.49 | 5/5 |
| Task group | alone | Astra Direct | Astra Hybrid |
| Atomic seen | 13/25 (52.0%) | 7/25 (28.0%) | 13/25 (52.0%) |
| Composite seen | 3/25 (12.0%) | 4/25 (16.0%) | 7/25 (28.0%) |
| Composite unseen | 1/25 (4.0%) | 14/25 (56.0%) | 9/25 (36.0%) |
| All tasks | 17/75 (22.7%) | 25/75 (33.3%) | 29/75 (38.7%) |
| Task | Dataset | Split | Successes | SR | SPL | nDTW | sDTW |
| VLN-CE | R2R | Val-Unseen | 39/50 | 78 | 65.27 | 72.20 | 59.35 |
| VLN-CE | RxR (English) | Val-Unseen | 46/50 | 92 | 77.25 | 84.73 | 80.42 |
| ObjectNav | MP3D | Val | 28/50 | 56 | 22.43 | – | – |
| ObjectNav | HM3D | Val | 41/50 | 82 | 43.69 | – | – |
| VLN-CE | ObjectNav | ||||||||
| R2R | RxR | MP3D | HM3D | ||||||
| System | Views | SR | SPL | SR | SPL | SR | SPL | SR | SPL |
| Astra | 1 | 78.00 | 65.27 | 92.00 | 77.25 | 56.00 | 22.43 | 82.00 | 43.69 |
| LightNav-0 | 1 | 58.00 | 54.14 | 72.00 | 64.08 | 20.00 | 8.39 | 66.00 | 37.75 |
| Uni-NaVid 7B | 1 | 30.00 | 27.76 | 34.00 | 25.67 | 8.00 | 5.46 | 30.00 | 20.48 |
| Multi-view methods (separate observation regime) | |||||||||
| Planner / Attempt | Outcome | Time (s) | Progress (m) | Final error (m) | Fall | Reached |
| PASSAGE | Goal reached | 13.18 | 8.622 | 0.495 | No | Yes |
| Astra Attempt 1 | Fall | 9.80 | 1.561 | 7.556 | Yes | No |
| Astra Attempt 2 | No progress | 20.28 | -0.024 | 9.141 | No | No |
| Astra Attempt 3 | No progress | 8.00 | 0.012 | 9.104 | No | No |
| Astra Attempt 4 | No progress | 8.48 | 0.159 | 8.958 | No | No |
| Astra Attempt 5 | Time horizon | 30.00 | 2.598 | 6.519 | No | No |
| Interface | Strict success | Goal + stop | Falls | Final error (m) | Travel (m) | Peak contact (N) | Non-foot contact (N) | Five-point RMSE (m) | |
| Five-point | 12 | 0 | 4 (33.3%) | 1 (8.3%) | 3.441 | 5.745 | 1853 | 644 | 0.161 |
| WholeBody-14 | 12 | 0 | 3 (25.0%) | 4 (33.3%) | 5.334 | 3.866 | 2293 | 1562 | 0.197 |
| Task | DreamerV3 | TD-MPC2 | SAC | Astra | Threshold |
| Maze | 272.3 | 244.3 | 144.8 | 1358.8 | 1200 |
| Reach | 7580.9 | 7316.1 | 4565.1 | 11430.2 | 12000 |
| Walk | 800.2 | 782.0 | 31.7 | 848.7 | 700 |
| Run | 633.8 | 93.3 | 5.0 | 642.1 | 700 |
| Crawl | 878.8 | 957.4 | 330.0 | 971.9 | 700 |
| Stair | 131.1 | 70.4 | 14.1 | 272.5 | 700 |
| Task | Successes | Rate |
| Handover | 8/10 | 80% |
| Mobile pick/place | 7/10 | 70% |
| Tabletop | 9/10 | 90% |
| XMove bend pick | 9/10 | 90% |
| XMove pick | 9/10 | 90% |
| Bend | 8/10 | 80% |
| Setting | Recorded resource use | Scope and interpretation |
| RoboDojo, 50 task instances per condition | Hybrid / Direct: 624.8M / 1,132.3M total tokens; 607.6M / 1,107.3M cached input; 15.90M / 22.91M uncached input; 1.31M / 2.09M output | Earlier attempts excluded. Cached input dominates the totals. |
| RoboCasa, 75 episodes per condition | Direct / Hybrid: 7,910 / 8,941 recorded model requests; median 85 / 108 per episode | Counts cover saved request logs, including recovery history, rather than all inference calls. |
| Dense locomotion, final attempt | 250 synchronous calls for 30 s of simulated robot motion; mean model latency 39.86 s per call | Physics pauses during inference; PASSAGE planning takes approximately 0.08 s per call. |
Appendix figures & tables32 assets
Supplementary material from the paper’s appendix.
Appendix
| Native Score | Success rate (%) | |||
| Task | Direct | Hybrid | Direct | Hybrid |
| Organize Table | 30.0 | 60.0 | 0 | 0 |
| Classify Objects By Language | 60.0 | 38.0 | 40 | 20 |
| Imitate Sorting Sequence | 0.0 | 53.0 | 0 | 40 |
| Arrange Largest Number | 57.0 | 50.0 | 40 | 40 |
| Pack Objects Into Box | 50.0 | 50.0 | 20 | 20 |
| Task | Direct | Hybrid | Cosmos | DreamZero | |
| Blocks into bin | 5/5 | 5/5 | 0/5 | 2/5 | 0/5 |
| Pumpkins in clutter | 5/5 | 4/5 | 0/5 | 0/5 | 0/5 |
| Butter on raisin box | 5/5 | 5/5 | 0/5 | 1/5 | 2/5 |
| Stack blocks in order | 5/5 | 4/5 | 0/5 | 0/5 | 0/5 |
| Reorient red mug | 5/5 | 4/5 | 2/5 | 1/5 | 1/5 |
| Larger raisin box into bin | 4/5 | 4/5 | 3/5 | 0/5 | 0/5 |
| Configuration | Seed role | Seed | Native return | Controls | Physical outcome |
| Focused Push | Development | 0 | 872.58 | 396 | 4.66 cm; success |
| Focused Push | Development | 3 | 866.75 | 423 | 4.78 cm; success |
| Focused Push | Development | 5 | 848.01 | 457 | 4.85 cm; success |
| Focused Push | Transfer | 7 | 900.37 | 323 | 4.81 cm; success |
| Walk, 1.8 Hz | Development | 1 | 743.54 | 1000 | No fall |
| Walk, 1.8 Hz | Development | 2 | 748.92 | 1000 | No fall |
| Metric | Hybrid | Direct |
| Executed control steps | 42,750 | 38,221 |
| Executed action segments | 3,776 | 7,729 |
| Mean simulated duration per slot (s) | 34.20 | 30.58 |
| Total tokens, including cached input | 624,762,828 | 1,132,343,772 |
| Cached input tokens | 607,555,840 | 1,107,349,760 |
| Uncached input tokens | 15,901,963 | 22,907,448 |
| Task | Threshold | Astra |
| Stand | 800 | 954.2 0.1 |
| Walk | 700 | 848.7 5.9 |
| Run | 700 | 642.1 2.8 |
| Kitchen | 4 | 0.0 0.0 |
| Maze | 1200 | 1358.8 7.4 |
| Hurdle | 700 | 183.0 0.3 |
| Stage | Condition | Valid trials | Success + hold |
| Grasp stabilization | Grasp entry, enclosure, and support transfer | 5 | 2/5 |
| Depth-assisted grasp | Relative depth and side-entry grasp guidance | 5 | 3/5 |
| Fallen-carton recovery | Table-edge support, regrasping, and handover | 1 | 0/1 |
| Depth-assisted recovery | Relative depth and recovery notes | 2 | 0/2 |
| Trajectory-informed guidance | Scene 1; instructions derived from contrasting trajectories | 1 | 1/1 |
| Post-trial reflection | Scene 1; guidance informed by the preceding trial | 1 | 0/1 |
| Task | Task instruction | (s) | Verification (s) | |
| Pot lift and hold | Lift the pot at least 10 centimeters above the table and hold it steadily without dropping it. | 1 | 30 | 3 |
| Headphones in box | Pick the headset and place it into the box. | 1 | 20 | 2 |
| Toy retrieval | Take the toy out of the bin and place it on the table. | 2 | 20 | 2 |
| Mug hanging | Hang the mug on the mug rack without requiring handle-specific alignment. | 1 | 25 | 2 |
| Two mahjong tiles | Place both mahjongs into the basket. | 2 | 30 | 2 |
| Bottles/cans sorting | Place bottles into the front-left basket and cans into the front-right basket. | 4 | 60 | 2 |
| Configuration | Specification |
| Demonstration budget | 100 trajectories per task; ten tasks; 1,000 trajectories in total. |
| S1 adaptation | One multi-task policy, shared by the standalone baseline and Hybrid. |
| Direct action responsibility | Astra specifies wrist/arm targets and finger-joint targets. |
| Hybrid action responsibility | supplies manipulation actions; Astra can edit wrist and finger targets. |
| S1 initialization and training | base initialization; full-parameter finetuning; 50,000 optimization steps; global batch size 256. |
| S1 optimizer | AdamW with , , weight decay , and gradient-norm clipping at 1.0. |
| Task | Subgoal predicates and task-specific verification conditions |
| Pot lift and hold | The pot root is at least 10 cm above its initial height. The 0.8 tier requires this lift condition at termination. Full-score verification requires continued hand contact, no pot–table contact, and stability for 3 s. |
| Headphones in box | The root projects inside the box-bottom polygon; at least 50% of the vertically constrained minimum bounding-box volume lies in the vertical prism extending upward from that polygon; and hand contact is absent. The prism has no upper height limit, so a released headset resting across the rim can qualify. |
| Toy retrieval | Two independent subgoals: (i) the root projects outside the box-bottom polygon; (ii) the toy contacts the table and its root height is at most the table height plus its reference resting height plus 3 cm. Reference resting height is the initial root height above the box-bottom reference plane. |
| Mug hanging | Conjunctive subgoal: lift from the initial root height cm, root height above the table cm, and root-to-rack horizontal distance cm. Handle alignment is unnecessary. |
| Two mahjong tiles | One subgoal per tile: the root projects inside the basket interior (half-widths 10.85 and 6.67 cm), and hand contact is absent. No constraint is imposed on the tiles' vertical positions. Full credit requires both released tiles to remain stable for 2 s. |
| Bottles/cans sorting | One subgoal per object: the root projects inside its assigned basket interior and its height relative to the basket reference plane is in cm. Bottles belong in the front-left basket and cans in the front-right basket; placement in the other basket receives no credit. |
| Task | Platform | Goal | Endpoint | Reported measures |
| Cylinder rotation | Sharpa / Isaac Lab | Rotate about world at | 20 s (400 steps) | Orientation error, axial-speed MAE, At-goal, completion, drop |
| Cuboid rotation | Sharpa / Isaac Lab | Rotate about world at | 10 s (200 steps) | Orientation error, axial-speed MAE, At-goal, completion, drop |
| Cylinder translation | Allegro / Isaac Gym | Reach a Cartesian target | 15 s (300 steps) | Terminal position error, success, drop |
| Translation + rotation | Allegro / Isaac Gym | Reach the Cartesian target and rotate the long axis by | Astra at 120 decisions; RL at matched steps and 15 s | Terminal position and rotation errors, joint success, drop |
| Task | Policy construction | Training setup |
| Cylinder rotation | PPO from scratch; four distributed ranks; speed curriculum | , , 1,000 updates; transitions |
| Cuboid rotation | PPO initialized from the cylinder actor and actor normalizer | , , 4,000 updates; transitions |
| Cylinder translation | Released privileged-teacher/tactile-student policy | Released checkpoint; no retraining |
| Translation + rotation | Privileged PPO teacher initialized from the released oracle | , , 20,000 updates; transitions |
| Case | Controller | End | Steps | Decisions | At-goal | Error (rad) | Speed MAE |
| 000 | Astra Direct | Budget | 223 | 100 | 0.45% | 1.756 | 0.977 |
| RL | Horizon | 400 | 400 | 73.25% | 0.193 | 0.208 | |
| 004 | Astra Direct | Drop | 326 | 84 | 0.31% | 1.532 | 0.906 |
| RL | Horizon | 400 | 400 | 77.00% | 0.182 | 0.335 | |
| 008 | Astra Direct | Horizon | 400 | 96 | 0.75% | 1.610 | 0.957 |
| RL | Horizon | 400 | 400 | 75.50% | 0.201 | 0.297 |
| Case | Controller | Split | Steps | Decisions | At-goal | Error (rad) | Speed MAE |
| 000 | Astra Direct | overlap | 200 | 50 | 0.50% | 0.798 | 0.143 |
| RL | overlap | 200 | 200 | 81.00% | 0.071 | 0.053 | |
| 001 | Astra Direct | overlap | 200 | 56 | 6.50% | 0.917 | 0.168 |
| RL | overlap | 200 | 200 | 51.00% | 0.123 | 0.067 | |
| 002 | Astra Direct | overlap | 200 | 54 | 5.50% | 0.690 | 0.139 |
| RL | overlap | 200 | 200 | 65.00% | 0.111 | 0.063 |
| Sample | Controller | Steps | Decisions | Final error (mm) | Outcome |
| 000 | Astra Direct | 300 | 102 | 74.73 | Fail |
| RL | 300 | 300 | 10.02 | Pass | |
| 001 | Astra Direct | 300 | 79 | 19.13 | Pass |
| RL | 300 | 300 | 31.12 | Fail | |
| 002 | Astra Direct | 300 | 101 | 61.67 | Fail |
| RL | 300 | 300 | 9.84 | Pass |
| Sample | Astra steps | Astra endpoint | RL matched | RL full, 15 s | Outcome (A / M / F) |
| 000 | 153 | 51.22 / 38.77 | 18.84 / 3.22 | 19.37 / 3.39 | Fail / Pass / Pass |
| 001 | 129 | 38.12 / 28.56 | 15.53 / 3.57 | 15.38 / 3.55 | Fail / Pass / Pass |
| 002 | 134 | 49.68 / 35.16 | 21.07 / 21.70 | 20.94 / 5.71 | Fail / Fail / Pass |
| 003 | 120 | 50.60 / 27.41 | 10.42 / 9.56 | 11.26 / 1.87 | Fail / Pass / Pass |
| 004 | 132 | 47.15 / 33.67 | 15.22 / 3.53 | 13.45 / 2.91 | Fail / Pass / Pass |
| Task / dataset | Successful STOP | Unsuccessful STOP | Step limit | Mean steps | Mean path (m) |
| VLN-CE / R2R | 39 | 11 | 0 | 85.2 | 12.03 |
| VLN-CE / RxR | 46 | 4 | 0 | 98.1 | 12.95 |
| ObjectNav / MP3D | 28 | 12 | 10 | 220.8 | 26.58 |
| ObjectNav / HM3D | 41 | 4 | 5 | 163.2 | 16.80 |
| Released policy | Source commit | Checkpoint commit |
| LightNav-0 | 3015508b70fb | 826dc5fbfa37 |
| Uni-NaVid 7B | 79ef5ea3fea1 | 0437222534b2 |
| OmniNav Flow | e8b485953c65 | f73a8d094f53 |
| Task / dataset | System | Successes | nDTW | sDTW | Failed STOP | Step limit |
| VLN-CE / R2R | Astra | 39/50 | 72.20 | 59.35 | 11 | 0 |
| VLN-CE / R2R | LightNav-0 | 29/50 | 72.77 | 48.89 | 21 | 0 |
| VLN-CE / R2R | Uni-NaVid 7B | 15/50 | 61.76 | 26.38 | 14 | 21 |
| VLN-CE / R2R | OmniNav Flow (3 views) | 28/50 | 71.18 | 49.23 | 22 | 0 |
| VLN-CE / RxR | Astra | 46/50 | 84.73 | 80.42 | 4 | 0 |
| VLN-CE / RxR | LightNav-0 | 36/50 | 77.44 | 64.89 | 14 | 0 |
| R2R | RxR | ||||
| Method | Observations | SR | SPL | SR | SPL |
| Published full-benchmark references (published protocols) | |||||
| Uni-NaVid [ 61 ] | 1 RGB | 47.0 | 42.7 | 48.7 | 40.9 |
| NavFoM [ 62 ] | 1 RGB | 56.2 | 51.2 | 57.4 | 49.4 |
| LightNav-0 [ 49 ] | 1 RGB | 68.5 | 62.8 | 73.6 | 64.5 |
| Qwen-RobotNav-4B [ 63 ] | 1 RGB | 66.9 | 60.5 | 71.3 | 61.5 |
| MP3D | HM3D v2 | ||||
| Method | Observations | SR | SPL | SR | SPL |
| Published full-benchmark references (published protocols) | |||||
| CogNav [ 6 ] | RGB-D + pose | 46.6 | 16.1 | – | – |
| LightNav-0 [ 49 ] | 1 RGB | 53.3 | 21.2 | 79.5 | 43.7 |
| Qwen-RobotNav-4B [ 63 ] | RGB | 52.2 | 16.0 | 75.6 | 30.6 |
| Qwen-RobotNav-8B [ 63 ] | RGB | 48.8 | 17.7 | 71.2 | 33.0 |
| Group | Task | Horizon | alone | Direct | Hybrid |
| Atomic seen | CoffeeSetupMug | 600 | 3/5 | 0/5 | 0/5 |
| OpenDrawer | 750 | 3/5 | 2/5 | 3/5 | |
| OpenStandMixerHead | 450 | 1/5 | 0/5 | 2/5 | |
| PickPlaceDrawerToCounter | 750 | 2/5 | 0/5 | 3/5 | |
| PickPlaceSinkToCounter | 900 | 4/5 | 5/5 | 5/5 | |
| Composite seen | DeliverStraw | 2550 | 0/5 | 1/5 | 1/5 |
| Comparison ( vs. ) | Both succeed | only | only | Both fail | Exact |
| Hybrid vs. Direct | 16 | 13 | 9 | 37 | 0.5235 |
| Hybrid vs. | 11 | 18 | 6 | 40 | 0.0227 |
| Direct vs. | 7 | 18 | 10 | 40 | 0.1849 |
| Harness configuration | Direct | Hybrid |
| Standard, xhigh , 20 steps | 2 | 19 |
| Conservative optimization | 0 | 8 |
| Conservative optimization with guards | 73 | 48 |
| Quantity | alone | Direct | Hybrid |
| Episodes | 75 | 75 | 75 |
| Executed simulator steps | 140,958 | 130,665 | 130,549 |
| Policy proposals | 7,072 | 0 | 6,185 |
| Astra-authored action segments | 0 | 6,560 | 2,941 |
| Accepted policy segments | – | 0 | 3,621 |
| Recorded model requests, retained chains | 0 | 7,910 | 8,941 |