Organizations: University of Chinese Academy of Sciences · Institute of Automation, Chinese Academy of Sciences · Independent Researcher · Tsinghua University
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improvement. This process develops joint-level control, evidence-seeking observation, and reusable manipulation skills without model-weight updates or a separately trained action policy. Across 42 RoboDojo tasks, held-out-layout evaluation of retained task-specific deployment versions yields a five-dimension average score of 68.14/100 and 62.00% success, compared with 47.17% for RoboDawn's one-shot Astra agent, the strongest published reference in our comparison. On eight manipulation tasks challenging direct Astra, PhysEvo achieves 55.00% success, compared with 1.25% for the direct-Astra reference. Deploying the simulation-evolved harness on AgileX PiPER and continuing skill revision yields 90.60/100 average score and 84.00% success across 25 trials on five real-world tasks. PhysEvo turns the consequences of action into persistent, testable changes to how a frozen model acts and improves.
Figures & tables
Figure 1: Physical RSI in PhysEvo: execute, diagnose, revise, test, and inherit. The lower panel compares success on eight challenging tasks for Direct Astra ( table 3 ).
Figure 2: Physical RSI in PhysEvo. A frozen Astra acts through the task harness and revises both task and meta resources from execution feedback. Accepted revisions support later episodes and improvement rounds.
Method
General.
Precision
Long-horizon
Memory
Open
Average
Score
SR (%)
Score
SR (%)
Score
SR (%)
Score
SR (%)
Score
SR (%)
Score
SR (%)
Robot policies
Simate-beta
35.09
27.95
34.35
26.92
57.84
43.42
33.33
33.00
9.12
8.50
33.95
27.96
Liber-0 Preview
24.99
18.61
38.28
33.17
45.98
33.17
37.77
37.33
6.68
5.33
30.74
25.52
Liber-0 Lite
25.36
19.11
36.60
31.08
45.50
32.92
35.68
35.44
3.06
2.58
29.24
24.23
DM0.5
15.78
10.95
24.82
16.75
33.70
19.50
47.74
47.44
2.43
2.08
24.90
19.34
Table 1: RoboDojo performance across five capability dimensions. The average weights dimensions equally.
Task
G0.5
Xiaomi
O-WAM
Hybrid
Direct
PhysEvo
Score
SR
Score
SR
Score
SR
Score
SR
Score
SR
Score
SR
Organize table
46.33
11.33
57.67
14.67
62.50
13.33
60.00
0.00
30.00
0.00
60.00
0.00
Classify by language
1.07
0.00
2.00
0.67
1.33
0.00
38.00
20.00
60.00
40.00
100.00
100.00
Imitate sorting
1.67
0.00
2.50
0.67
2.90
0.67
53.00
40.00
0.00
0.00
100.00
100.00
Largest number
4.11
0.53
8.56
3.73
4.36
0.53
50.00
40.00
57.00
40.00
100.00
100.00
Pack objects
17.12
2.93
18.69
2.40
20.83
3.73
50.00
20.00
50.00
20.00
70.00
40.00
Table 2: Comparison on the ten-task GPT-as-Policy evaluation set. Hybrid combines Astra with π0.5 ; Direct uses Astra alone.
Task
Official Astra
G0.5
Simate-beta
Liber-0 Preview
RoboDawn
PhysEvo
Score
SR
Score
SR
Score
SR
Score
SR
Score
SR
Score
SR
Make kong
0.00
0.00
90.00
90.00
76.00
76.00
18.00
18.00
60.00
60.00
60.00
60.00
Build tower
16.40
2.00
82.93
78.67
87.40
83.33
84.33
81.00
54.00
40.00
40.00
20.00
Insert tubes
6.00
0.00
58.53
42.67
41.20
22.00
82.53
73.00
88.00
80.00
88.00
80.00
Tic-tac-toe
10.70
0.00
65.23
40.00
86.27
60.67
95.67
89.00
100.00
100.00
100.00
100.00
Pour balls into vase
4.00
4.00
28.00
28.00
45.33
45.33
30.67
31.00
0.00
0.00
80.00
80.00
Table 3: Eight manipulation tasks challenging direct Astra. PhysEvo results average five layouts per task. The task set is defined by the RoboDojo report, with VLA SR ≥20% and Astra SR <5% .
Figure 3: Joint reconfiguration enables recovery from an obstructive arm posture.
Figure 4: Skill revision maintains a useful view throughout pouring and recovery.
Figure 5: An inherited review tool supports a later packing-skill revision.
Task
Score
SR (%)
Stack blocks
100.00
100.00
Stack bowls
100.00
100.00
Place pens
100.00
100.00
Pour water
60.00
60.00
Write PhysEvo
93.00
60.00
Overall
90.60
84.00
Table 4: Real-world performance on AgileX PiPER. Each task has five trials.
Figure 6: Five AgileX PiPER tasks after trajectory-driven skill revision. Each row shows one execution in time; alternate final views are marked.
Appendix figures & tables9 assets
Supplementary material from the paper’s appendix.
Appendix
Inputs: frozen Astra Mθ , initial task and meta harnesses H0 , and an interaction budget.
1
Execute a task using the current task harness.
2
Give the meta-agent the task outcome, relevant images and robot states, action feedback, and current code.
3
Let the meta-agent inspect the failure, retrieve evidence, and run diagnostic experiments.
4
Propose a candidate revision to tools, skills, execution logic, or meta resources.
5
Run the edited code and check its outputs. For task-behavior changes, compare the candidate with the current harness on the same task and initial scene.
6
Adopt a candidate with a score gain, or progress beyond the failure without a score decrease. Reject a candidate that lowers the score. Test diagnostic changes against the error or measurement they address.
Appendix
Table 5: Physical RSI procedure in PhysEvo.
Method
Mean score
SR (%)
PhysEvo (ours)
62.88
55.00
Simate-beta
61.97
50.83
Liber-0 Preview
54.22
46.44
RoboDawn (1-shot)
52.50
45.00
GalaxeaVLA (G0.5)
53.70
43.09
Xiaomi-Robotics-1
44.73
31.83
Appendix
Table 6: Eight-task means for the original references and the extended comparison. Bold and underlining mark the best and second-best values.
Figure 7: Simulation executions, cases S01–S07 from top to bottom. Task names and instructions appear above each strip; frames progress from left to right.
Figure 8: Simulation executions, cases S08–S14 from top to bottom. Task names and instructions appear above each strip; frames progress from left to right.
Figure 9: Simulation executions, cases S15–S21 from top to bottom. Task names and instructions appear above each strip; frames progress from left to right.
Figure 10: Simulation executions, cases S22–S28 from top to bottom. Task names and instructions appear above each strip; frames progress from left to right.
Figure 11: Simulation executions, cases S29–S35 from top to bottom. Task names and instructions appear above each strip; frames progress from left to right.
Figure 12: Simulation executions, cases S36–S42 from top to bottom. Task names and instructions appear above each strip; frames progress from left to right.
Tool
Tool-specific arguments and semantics
control_arms
left , right , steps , settle . Each non-null arm has control , joints , and gripper fields, as shown above.
review_frames
frames: [{step, camera}] . Retrieves 1–6 same-episode RGB frames and their public poses without motion.
step , feature , views: [{camera, pixel}] . Matches one physical feature in 2–3 distinct current views; returns geometry estimates and diagnostics without motion.
triangulate_temporal_point
step , camera , feature , stationary_target , camera_hand_empty , stationarity_evidence , and samples: [{step, pixel}] . Uses 2–3 reviewed views from one wrist camera. Both Boolean fields must be true ; sampling steps increase, end at the current step, and span at most 24 native steps.
verify_completion
No additional arguments. Advances six hold steps and checks native termination.
School of Information, Renmin University of China · Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China · Tsinghua University +2