PhysEvo: Astra Can Act, Let It
Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Independent Researcher · Tsinghua University
Abstract
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.
Figures & tables
| Method | General. | Precision | Long-horizon | Memory | Open | Average | ||||||
| Score | SR (%) | Score | SR (%) | Score | SR (%) | Score | SR (%) | Score | SR (%) | Score | SR (%) | |
| Robot policies | ||||||||||||
| Simate-beta | 35.09 | 27.95 | 34.35 | 26.92 | 57.84 | 43.42 | 33.33 | 33.00 | 9.12 | 8.50 | 33.95 | 27.96 |
| Liber-0 Preview | 24.99 | 18.61 | 38.28 | 33.17 | 45.98 | 33.17 | 37.77 | 37.33 | 6.68 | 5.33 | 30.74 | 25.52 |
| Liber-0 Lite | 25.36 | 19.11 | 36.60 | 31.08 | 45.50 | 32.92 | 35.68 | 35.44 | 3.06 | 2.58 | 29.24 | 24.23 |
| DM0.5 | 15.78 | 10.95 | 24.82 | 16.75 | 33.70 | 19.50 | 47.74 | 47.44 | 2.43 | 2.08 | 24.90 | 19.34 |
| Task | G0.5 | Xiaomi | O-WAM | Hybrid | Direct | PhysEvo | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Score | SR | Score | SR | Score | SR | Score | SR | Score | SR | Score | SR | |
| Organize table | 46.33 | 11.33 | 57.67 | 14.67 | 62.50 | 13.33 | 60.00 | 0.00 | 30.00 | 0.00 | 60.00 | 0.00 |
| Classify by language | 1.07 | 0.00 | 2.00 | 0.67 | 1.33 | 0.00 | 38.00 | 20.00 | 60.00 | 40.00 | 100.00 | 100.00 |
| Imitate sorting | 1.67 | 0.00 | 2.50 | 0.67 | 2.90 | 0.67 | 53.00 | 40.00 | 0.00 | 0.00 | 100.00 | 100.00 |
| Largest number | 4.11 | 0.53 | 8.56 | 3.73 | 4.36 | 0.53 | 50.00 | 40.00 | 57.00 | 40.00 | 100.00 | 100.00 |
| Pack objects | 17.12 | 2.93 | 18.69 | 2.40 | 20.83 | 3.73 | 50.00 | 20.00 | 50.00 | 20.00 | 70.00 | 40.00 |
| Task | Official Astra | G0.5 | Simate-beta | Liber-0 Preview | RoboDawn | PhysEvo | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Score | SR | Score | SR | Score | SR | Score | SR | Score | SR | Score | SR | |
| Make kong | 0.00 | 0.00 | 90.00 | 90.00 | 76.00 | 76.00 | 18.00 | 18.00 | 60.00 | 60.00 | 60.00 | 60.00 |
| Build tower | 16.40 | 2.00 | 82.93 | 78.67 | 87.40 | 83.33 | 84.33 | 81.00 | 54.00 | 40.00 | 40.00 | 20.00 |
| Insert tubes | 6.00 | 0.00 | 58.53 | 42.67 | 41.20 | 22.00 | 82.53 | 73.00 | 88.00 | 80.00 | 88.00 | 80.00 |
| Tic-tac-toe | 10.70 | 0.00 | 65.23 | 40.00 | 86.27 | 60.67 | 95.67 | 89.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| Pour balls into vase | 4.00 | 4.00 | 28.00 | 28.00 | 45.33 | 45.33 | 30.67 | 31.00 | 0.00 | 0.00 | 80.00 | 80.00 |
| Task | Score | SR (%) |
|---|---|---|
| Stack blocks | 100.00 | 100.00 |
| Stack bowls | 100.00 | 100.00 |
| Place pens | 100.00 | 100.00 |
| Pour water | 60.00 | 60.00 |
| Write PhysEvo | 93.00 | 60.00 |
| Overall | 90.60 | 84.00 |
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
| Inputs: frozen Astra , initial task and meta harnesses , and an interaction budget. | |
|---|---|
| 1 | Execute a task using the current task harness. |
| 2 | Give the meta-agent the task outcome, relevant images and robot states, action feedback, and current code. |
| 3 | Let the meta-agent inspect the failure, retrieve evidence, and run diagnostic experiments. |
| 4 | Propose a candidate revision to tools, skills, execution logic, or meta resources. |
| 5 | Run the edited code and check its outputs. For task-behavior changes, compare the candidate with the current harness on the same task and initial scene. |
| 6 | Adopt a candidate with a score gain, or progress beyond the failure without a score decrease. Reject a candidate that lowers the score. Test diagnostic changes against the error or measurement they address. |
| Method | Mean score | SR (%) |
|---|---|---|
| PhysEvo (ours) | 62.88 | 55.00 |
| Simate-beta | 61.97 | 50.83 |
| Liber-0 Preview | 54.22 | 46.44 |
| RoboDawn (1-shot) | 52.50 | 45.00 |
| GalaxeaVLA (G0.5) | 53.70 | 43.09 |
| Xiaomi-Robotics-1 | 44.73 | 31.83 |
| Tool | Tool-specific arguments and semantics |
|---|---|
| control_arms | left , right , steps , settle . Each non-null arm has control , joints , and gripper fields, as shown above. |
| review_frames | frames: [{step, camera}] . Retrieves 1–6 same-episode RGB frames and their public poses without motion. |
| project_pixels | points: [{pixel: [u,v], z}] . Projects 1–32 head-view pixels onto model-specified height planes without motion. |
| triangulate_point | step , feature , views: [{camera, pixel}] . Matches one physical feature in 2–3 distinct current views; returns geometry estimates and diagnostics without motion. |
| triangulate_temporal_point | step , camera , feature , stationary_target , camera_hand_empty , stationarity_evidence , and samples: [{step, pixel}] . Uses 2–3 reviewed views from one wrist camera. Both Boolean fields must be true ; sampling steps increase, end at the current step, and span at most 24 native steps. |
| verify_completion | No additional arguments. Advances six hold steps and checks native termination. |