An Unexpected Robot Policy: Early Evaluations of GPT-6 Astra on RoboDojo and Beyond
Authors: Wenbo Zhang, Kaixuan Wang, Yutao Ouyang, Xiaoyu Huang, Liyang Li, Kailun Su, Weiyang Jin, Wenhao Chai, +4 more
Organizations: RoboProbe · RoboDojo · The University of Hong Kong · Tsinghua University · University of California, Berkeley · Princeton University · Massachusetts Institute of Technology · Peking University
Embodied AI systems are often organized into System 1 and System 2. System 1 is typically a pretrained policy that generates actions at high frequency, whereas System 2 is often instantiated as a vision-enabled language model for high-level planning. We ask whether a large language model (LLM) can act as the policy for robot manipulation without task-specific finetuning. We call this setting LLM as policy. We evaluate three LLMs on all 42 RoboDojo tasks and compare their scores with 40 public policies. Astra and GPT-5.5 use the official 50-episode-per-task protocol; DeepSeek-Flash uses 10 episodes per task. GPT-6 Astra achieves 22.48% average success rate and 28.97 Score over 2,100 trials, ranking above every public entry. Yet GPT-5.5 and DeepSeek-Flash reach only 0.88% and 1.92% average success rate with the same post-processing. We find that Astra exhibits a sharply polarized capability profile. It generalizes well to tasks that require semantic understanding but not high-precision control. In contrast, it performs poorly on tasks that require precision, dynamic control, or complex bimanual coordination. In-context experiments show no aggregate benefit from one-shot demonstrations, while selected interaction traces show within-episode corrections under perturbations. Overall, the evaluated LLMs vary substantially in manipulation performance. Astra stands out and provides initial evidence for the potential of a general-purpose manipulation model, although reliable precision and dynamic control remain limitations in the evaluated setting.
Figures & tables
Figure 1: Two ways to turn an observation into an executable action. Top: a dual-process controller separates a language model for semantic planning from a pretrained motor policy for execution. VLA denotes a vision language action model; WAM denotes a world action model. Bottom: in the LLM-as-policy setting, the language model selects motion targets directly. Non-learned post-processing converts these targets into executable commands without a pretrained motor policy.
Figure 2: The model-facing context and execution loop. A frozen language model chooses world-frame end-effector (EEF) targets at the grasp point through move_eef . The action conversion module is non-learned. It validates the request, converts grasp-point targets to flange poses, plans joint paths, and constructs executable joint commands. Requested gripper changes follow the arm-motion phase. The two return routes are the two branches of a request. A rejected target returns an error as the tool result and causes no motion, so the model retries against the same observation; this is the dashed route. An accepted target returns its planned duration, executes, and only then produces the next observation with the achieved state and arrival error; this is the solid route. give_up requests termination, but only the environment determines success. The context retains the latest two RGB frames from each camera.
Rank
Model
Average
Gen.
Prec.
Long-Hor.
Memory
Open
1
GPT-6 Astra
28.97/22.48
33.36/30.50
12.65/4.00
21.45/8.25
43.04/38.67
34.36/31.00
2
DM0.5
24.90/19.34
15.77/10.95
24.82/16.75
33.70/19.50
47.74/47.44
2.43/2.08
3
GalaxeaVLA (G0.5)
20.23/14.88
18.46/12.83
28.25/20.42
44.12/32.25
8.61/7.33
1.73/1.58
4
Xiaomi-Robotics-1
20.07/13.93
23.54/17.00
26.69/18.83
38.39/23.67
7.81/6.56
3.94/3.58
5
OpenWAM- α
17.18/11.92
20.71/14.83
18.45/9.25
34.93/25.33
10.41/9.11
1.41/1.08
6
Meituan-Robotics-0
14.95/9.53
13.75/8.17
16.77/7.75
29.61/18.58
10.06/8.89
4.54/4.25
Table 1: RoboDojo-Sim board, Score/SR% per cell: the top ten of 43 ranked entries plus the two remaining LLM controllers, which rank 28 and 33. Ranks are over all 43; the public submission contains only the 40 policy rows, where the same order gives DM0.5 rank 1. † DeepSeek-Flash is 1 seed × 10 episodes per task, not 50. Policy rows are a 2026-09-10 leaderboard snapshot. The full 43-row board is Table 5 .
Astra ≥ 20%, policy < 5%
Policy ≥ 20%, Astra < 5%
Task
Astra
Best policy
Task
Astra
Best policy
push_T
60.0
0.7 Xiaomi
make_kong
0.0
90.0 G0.5
arrange_largest_number
60.0
4.7 Xiaomi
build_tower
2.0
78.7 G0.5
align_blocks
50.0
0.0 DM0.5
insert_tubes
0.0
59.3 DM0.5
solve_equation
40.0
0.0 DM0.5
play_tic_tac_toe
0.0
58.7 OpenWAM
stack_blocks_by_language
40.0
2.7 DM0.5
pour_balls_into_vase
4.0
46.0 Xiaomi
Table 2: Complementary task-level outcomes. Left: Astra reaches at least 20% SR while every public policy remains below 5%. Right: the strongest public policy reaches at least 20% while Astra remains below 5%. Xiaomi denotes Xiaomi-Robotics-1; G0.5, GalaxeaVLA (G0.5).
Figure 3: Capabilities and limits of Astra used as a policy. Each panel shows one episode in three camera views; the task and observed evidence appear below. Panels (a)–(c) show capabilities and panels (d)–(f) show limits. These examples establish occurrence, not frequency.
Condition
Episodes
SR
vs zero-shot
Zero-shot (baseline)
340
78/340 = 22.9%
—
Image + end-effector demonstration
340
61/340 = 17.9%
− 5.0
Text demonstration
340
44/340 = 12.9%
− 10.0
Table 3: Demonstration ICL on 340 matched task–layout pairs: the same 34 tasks and ten layouts per task under each condition. The two demonstration conditions are new runs of 340 episodes each; the zero-shot baseline is the official 50-episode Astra campaign restricted to those exact pairs.
Perturbed component
Condition
Successes
LLM as policy (GPT-6 Astra)
None
No perturbation
8/8
Visual observation
Top–bottom image flip
8/8
Visual observation
Left–right image mirror
6/8
Visual observation
No head-camera view
6/8
Visual observation
Right-wrist view only
3/8
Table 4: Performance under perturbations on eight general_pickup layouts. Visual perturbations transform or remove camera views. Per-move pose jitter offsets each executed pose by a random displacement with mean magnitude 10 cm. Negated Cartesian axes change the coordinate convention and the advertised target bounds. The layouts are the subset Astra solves without perturbation, and each cell contains one episode per layout.
Early turn
Later turn
Left wrist
Head
Right wrist
Left wrist
Head
Right wrist
wrong arm
correct arm
(a) Left–right image mirror — general_pickup layout 0, pick up the scissors
Reasoning written by the model between the two turns: “The camera views show that the right arm is the one beside the scissors. I am restoring the idle left arm and preparing the right hand for a downward grasp.”
masked
masked
searching
masked
masked
found left-hand grasp
(b) Right-wrist view only — general_pickup layout 9
Figure 4: The observations the model receives under three perturbations, defined in Appendix D.2 . Each row is one condition, named under its frames; the left group of three views is an early turn and the right group a later turn of the same episode. The line under row (a) is the model’s own reasoning from the episode, and the lines under rows (b) and (c) describe an observed behaviour instead. Row (a) shows real frames from the flip_vision_lr episode on layout 0. The mirror reverses apparent left–right positions without changing camera identities. Under (b) only the right wrist camera is available, and the masked cells mark the views the model does not receive. Row (b) comes from the supplementary rerun with 170 calls and 400 environment steps; the standard-budget run failed. Under (c) every executed move is offset by a random displacement of mean magnitude 10 cm. The model receives both images and motion feedback; their contributions are not isolated (Appendix D.4 ).
Appendix figures & tables10 assets
Supplementary material from the paper’s appendix.
Appendix
Rank
Model
Average
Gen.
Prec.
Long-Hor.
Memory
Open
1
GPT-6 Astra
28.97/22.48
33.36/30.50
12.65/4.00
21.45/8.25
43.04/38.67
34.36/31.00
2
DM0.5
24.90/19.34
15.77/10.95
24.82/16.75
33.70/19.50
47.74/47.44
2.43/2.08
3
GalaxeaVLA (G0.5)
20.23/14.88
18.46/12.83
28.25/20.42
44.12/32.25
8.61/7.33
1.73/1.58
4
Xiaomi-Robotics-1
20.07/13.93
23.54/17.00
26.69/18.83
38.39/23.67
7.81/6.56
3.94/3.58
5
OpenWAM- α
17.18/11.92
20.71/14.83
18.45/9.25
34.93/25.33
10.41/9.11
1.41/1.08
6
Meituan-Robotics-0
14.95/9.53
13.75/8.17
16.77/7.75
29.61/18.58
10.06/8.89
4.54/4.25
Appendix
Table 5: RoboDojo-Sim board in the official Score/SR% cell format: the 40 public policies plus GPT-6 Astra, DeepSeek-Flash, and GPT-5.5, sorted by Average Score and ranked 1–43 together. Ranking the three LLM controllers alongside the board makes the comparison legible; the public submission itself contains only the 40 policy rows, where the same order gives DM0.5 rank 1. † DeepSeek-Flash is 1 seed × 10 episodes per task, not 50. Snapshot of the policy rows taken 2026-09-10 from the public leaderboard.
Task
GPT-6 Astra
GPT-5.5
DS-Flash †
DM0.5
Galaxea G0.5
Xiaomi-1
OpenWAM- α
Generalization
stack_bowls
66.5
0.3
3.0
30.5
42.2
56.4
49.9
push_T
60.0
0.0
0.0
0.0
0.0
0.7
0.0
pack_objects_into_box
11.6
0.2
2.0
15.1
18.2
19.9
22.5
fold_clothes
76.4
1.2
2.0
32.5
36.3
47.3
57.5
hang_mugs
0.9
0.0
0.0
9.0
8.6
15.7
16.3
Appendix
Table 6: Per-task Score on RoboDojo-Sim (process-reward mean × 100). Public columns copied from the official Per-Task board (2026-09-10); only the strongest four public policies are shown here. Astra and GPT-5.5: 1 seed, 50 episodes per task. † DeepSeek-Flash: 1 seed, 10 episodes per task. Generalization cells are the mean of Gen-Std and Gen-Rand. Best value per row in bold.
Task subset
Pairs
Zero-shot
Image demo
Text demo
Zero-shot 0/10 (15 tasks)
150
0/150 = 0.0%
2/150 = 1.3%
0/150 = 0.0%
At least one success (19 tasks)
190
78/190 = 41.1%
59/190 = 31.1%
44/190 = 23.2%
All tasks (34 tasks)
340
78/340 = 22.9%
61/340 = 17.9%
44/340 = 12.9%
Appendix
Table 7: Where the demonstration conditions lose episodes. The 34 tasks split into those the zero-shot controller never solves and those it solves at least once. The first row is the subset count reported in Section 4.5 ; the second row is the total of Table 3 minus the first row. No cell is a new measurement.
Task
Zero-shot
With demonstration
Condition
press_by_number
5/10
7/10
image
press_by_number
5/10
8/10
text
Appendix
Table 8: An illustrative task from the demonstration campaign. Each condition contains ten matched task–layout pairs. Higher success on this task does not identify the cause of the difference.
Condition
Layout
Outcome
Calls
Top–bottom image flip
0
success
14
Left–right image mirror
0
success
15
Negated Cartesian axes
0
success
19
Negated Cartesian axes
4
failure
22 of 100
No head-camera view
0
success
16
Per-move pose jitter
0
success
14
Appendix
Table 9: Selected episodes from the perturbation sweep. Calls counts the model calls spent; a failed episode reports the calls used before the budget or the episode ended. All rows are general_pickup at a 100-call budget except the last, which repeats layout 9 at 170 calls and 400 environment steps.
Call
Model’s own note, and the targets it requested
4
“The scissors are ahead and slightly to the right in the wrist view. I will move toward their hinge while staying above the tabletop.” right_x −0.40 , right_y 0.38 , right_z −1.07
5
“The wrist view now shows the near table edge, so the forward/back estimate was reversed. I will move back across the table toward the scissors.” right_y 0.02 , right_z −1.18
12
“The closer view places the hinge slightly right and ahead of the jaw center. I will make a small lateral correction over the metal hinge.” right_x −0.389 , right_y 0.025
Intervening calls omitted.
17
“The gripper has closed against the scissors rather than fully shutting. I will lift a few millimetres first to confirm they are secure.” right_z −0.78
19
“The scissors are securely held and clear of the table. I will lift them to more than 10 cm above their starting height.” right_z −0.89
Appendix
Table 10: Selected verbatim notes from a successful negate_xyz episode on general_pickup layout 0, 19 calls, success. The note is the text the model attached to its own move_eef call, followed by the dimensions that call requested. Call 5 refers to the wrist view and revises the forward/back estimate; its requested motion is rejected as unreachable. Targets use the transformed bounds described in Appendix D.2 .
Type
Entry
n
Raw trial scores
Score
SR
ARX X5
cover_blocks
2
0 / 0
0
0%
ARX X5
insert_tubes
3
0 / 0 / 0
0
0%
ARX X5
make_bread
1
0
0
0%
ARX X5
make_food
2
0 / 0
0
0%
ARX X5
store_in_safe
1
0.4
40
0%
Piper
fill_pen_holder
3
0 / 0 / 0
0
0%
Appendix
Table 11: RoboDojo-Real diagnostic sample. Diagnostic rows report Astra’s retained 12-task, 33-trial material; Score is the mean process score × 100 and SR is the full-success rate. The official 18-task rows are complete-protocol references and are not directly comparable.
Body
Task
Observation
Media
Franka
Pour water
initial miss; later flow reaches cup, then stops
Report
Franka
Cup on shelf
placement and mid-air drop observed
Report
Humanoid
Pick and place
rigid and soft objects manipulated
Not available
Humanoid
Make breakfast
toast-into-slot fails
Not available
Humanoid
Pour water
high pour, some spill
Not available
Humanoid
Fold a shirt
corners first; neatness stalls
Not available
Appendix
Table 12: Qualitative notes from separate joint-position deployments, not an official Score/SR. Franka clips are shown in the online report’s hardware section; the humanoid rows have no accompanying media. The report is available at https://robodojo-benchmark.com/report/gpt-6-astra-eval .
Score
Notes
F1
P / R
Tries
Wall
Twinkle, one hand
14
0.907
1.00 / 0.91
22
31 min
Twinkle, two hands
34
0.902
0.99 / 0.89
33
34 min
Chopin nocturne excerpt
117
0.599
0.85 / 0.58
150
40 min
Appendix
Table 13: Piano-performance F1 from the final verification episode of each run. Tries counts practice episodes plus that verification. Two-hand Twinkle reaches 0.902; the RL reference replay scores 0.886 in this setup.
Score
Change
Notes
F1
P / R
Correct hand
Twinkle, two hands
as delivered
34
0.902
0.99 / 0.89
35/35
Twinkle, two hands
time × 0.85
34
0.883
0.98 / 0.87
32/32
Twinkle, two hands
− 7 semitones, time × 1.25
34
0.837
0.94 / 0.82
33/38
Twinkle, two hands
+5 semitones, time × 1.1
34
0.805
0.93 / 0.79
30/33
Twinkle, two hands
+2 semitones
34
0.683
0.89 / 0.65
32/36
Twinkle, two hands
− 3 semitones
34
0.655
0.85 / 0.63
31/37
Appendix
Table 14: The delivered two-hand Twinkle program re-run without modification. Notes is the score length; Correct hand counts sounded notes assigned to the staff the score specifies. Time factors multiply note timestamps, not tempo. P/R and F1 are reported by the piano evaluator. One episode per row.