A generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can already repair programs from execution feedback, yet it remains a central challenge to organize this experience around the task structure that gives it meaning, so that each repair is attributed to the responsible capability, supported by execution evidence, and validated before it is reused. We introduce RoboRSI, a robot self-improvement system built on Top-Down Skill Refinement (TSR). TSR decomposes tasks into compound, atomic, and base skills with scoped responsibilities and explicit input--output contracts, attributes each execution outcome to the responsible branch, and confines revision to that branch. Building upon this structure, a Manager, Planner, Engineer, and Reviewer coordinate planning, execution, diagnosis, and the validated release of new skills, while people steer the process through objectives and corrections; stable skill sequences are further consolidated into reusable compound skills. On a mobile manipulator, RoboRSI develops multi-object household cleanup over 104 rounds. In simulation, it achieves the highest success rate on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, exceeding the strongest baseline by 2.7 to 11.0 percentage points.
Figures & tables
Figure 1: RoboRSI at a glance. The multi-agent self-improvement loop (center), real-world mobile manipulation and compound-skill gains (left), and simulation results (right).
Figure 2: RoboRSI framework. Given a task, the Manager organizes its skill hierarchy, the Planner composes an execution plan, and the Engineer runs the required skills. The Reviewer diagnoses failures from observations and tool traces, returning them for TSR-guided revision. Tested updates enter the next iteration, and stable branches become parameterized skills for subsequent tasks.
Figure 3: TSR during physical cleanup. (a) Four representative task–skill structures show the progression from an initial task proposal to a developed skill library; colored blocks mark the skill operations at each snapshot. (b) Round 25 (top): the robot holds a wrapper at the correct bin without releasing it (i), although the arm has reached its observation pose (ii); the motion call reported a tolerance miss and the fallback exhausted its budget. TSR traces the failure to this observation motion and revises only the Place Object branch, which rechecks the reached pose and executes the complete code-produced release sequence; the next placement succeeds (iii). Round 33 (bottom): the destination bucket appears only at the image edge in the chassis view (iv) but completely in the in-hand view of the same sweep (v); the consumer rejected the producer’s destination, and the run ended with the second item held (vi). The revision leaves identity to the producer and prefers complete views. Both revisions passed offline replay of the four most recent runs and the existing regression tests (Appendix D.4 ).
Figure 4: Placement progress in selected physical cleanup runs. Each point shows the number of objects placed among the first four items considered in that run; placements were confirmed on site. Marker colors indicate the audited task outcome. Points are equally spaced in selected-run order, while tick labels give the development run numbers. Dashed segments connect selected observations only. The four-item denominator is a common display reference: task manifests and object configurations vary across runs.
Figure 5: Two-scene physical mobile cleanup. Selected frames show preparation, approach, grasping, placement, and the recorded final scene. The return arrow indicates repeated object handling. Objects and receptacles were rearranged between scenes, while the agent configuration remained unchanged.
Benchmark
Method
Success
Coverage
LIBERO
CaP-X
36/150 (24.0%)
16/30 (53.3%)
Maestro
62/150 (41.3%)
23/30 (76.7%)
OpenETA
76/150 (50.7%)
23/30 (76.7%)
RoboRSI
84/150 (56.0%)
24/30 (80.0%)
LIBERO-PRO
CaP-X
111/600 (18.5%)
46/120 (38.3%)
Maestro
198/600 (33.0%)
77/120 (64.2%)
Table 1: Simulation results. Success is successful/valid episodes (rate); coverage is solved/total tasks (rate), counting each task once after its first success. LIBERO uses 30 tasks × 5 episodes, LIBERO-PRO uses 120 tasks × 5 episodes, LIBERO-Plus uses 840 perturbation instances of 30 tasks, and RoboTwin uses 50 tasks × 3 episodes for the baselines; for RoboRSI we report the first 154 episodes of its online run on the same 50 tasks. RoboRSI improves online on LIBERO, LIBERO-PRO, and RoboTwin and uses its fixed LIBERO library on LIBERO-Plus. Maestro is evaluated on LIBERO and LIBERO-PRO because of its computational cost.
Task suite
Perturbation type
Method
Spatial
Object
Goal
Language
Object
Position
Task
All
CaP-X
24.5
13.5
17.5
21.3
16.7
17.3
18.7
18.5
Maestro
39.5
38.0
21.5
34.7
30.0
32.0
35.3
33.0
OpenETA
50.0
36.5
29.0
42.7
33.3
36.0
42.0
38.5
RoboRSI
57.0 (+7.0)
54.5 (+16.5)
37.0 (+8.0)
52.0 (+9.3)
48.7 (+15.3)
45.3 (+9.3)
52.0 (+10.0)
49.5 (+11.0)
Table 2: LIBERO-PRO by task suite and perturbation type. Success rate (%); each suite contains 200 episodes and each perturbation type 150 episodes per method. Bold: best in the column; gray: RoboRSI’s gain over the strongest baseline in percentage points.
Figure 6: Failure analysis. (a) Composition of failed episodes on LIBERO-PRO and LIBERO-Plus; n is the number of failed episodes, and the premature-completion share is labeled. (b) Success on LIBERO-Plus by perturbation type, 120 instances each; numbers give RoboRSI’s difference from OpenETA in percentage points. (c) Share of episodes that still succeed after an execution failure, over the 209 LIBERO-PRO and 258 LIBERO-Plus (task, seed) pairs in which every method has a complete trace and at least one execution failure (Appendix C.1 ).
Method
Spatial
Object
Goal
All
CaP-X
12/50
15/50
9/50
36/150
Maestro
22/50
24/50
16/50
62/150
OpenETA
23/50
29/50
24/50
76/150
RoboRSI
30/50
33/50
21/50
84/150
Table 3: LIBERO by suite. Successful/valid episodes per suite.
Figure 7: One-day self-iteration. Cumulative number of tasks solved at least once after each self-iteration round; each task is counted at its first recorded success.
Figure 8: Reuse of a revised skill on new tasks. The revised grasp_object , developed on a single task, is invoked on three tasks outside its development set, each of which had failed in the earlier frozen evaluation. Columns show the first recorded frame, the state after the grasp_object call, and the final successful state.
Figure 9: Simulation cases with the same initial state across methods. Images show the first and last recorded frames of the RoboRSI episode; text summarizes the decisive tool calls of each method.
Condition
Roles
Tasks solved
Single agent
Engineer
9/50 (18.0%)
Multi-agent
Planner, Engineer, Reviewer
36/50 (72.0%)
Table 4: Single-agent and multi-agent development on RoboTwin. Tasks solved at least once over the full online run, out of 50; the multi-agent condition is the RoboTwin run of Table 1 , whose first 154 episodes are reported there.
Condition
Files
Changed lines
Atomic success
Released
TSR
3
50
2/2
yes
Flat, candidate 1
4
104
1/2
no
Flat, candidate 2
1
63
1/2
no
Table 5: TSR and flat revision on the same failed task. Files and lines count changes relative to the common starting code; Atomic success is over the two development seeds.
Figure 10: Effect of code consolidation. (a) Task success over 600 matched episodes per condition. (b) Median token, model-call, and time cost on the 118-task efficiency sample, normalized to Code off.
LIBERO-PRO success (%)
RoboTwin
LIBERO-PRO episodes
Backbone
All
Spatial
Object
Goal
success
Model calls
Budget exhausted
GPT-5.6-SOL
38.3
38.3
48.3
28.3
24/150
115
79
GPT-5.5
48.6
53.3
55.0
37.5
19/150
122
80
GPT-5.4
40.3
55.8
39.2
25.8
21/150
124
101
Table 6: Backbone comparison. The study uses a separate panel of 120 LIBERO-PRO tasks × 3 seeds run independently of Table 1 , so only the relative differences between backbones are compared. LIBERO-PRO success over 360 episodes per backbone, overall and by task suite; RoboTwin successes over 150 episodes. Model calls are medians per LIBERO-PRO episode; budget exhausted counts LIBERO-PRO episodes that ended at the interaction limit. Bold: best in the column.
Figure 11: ACT policies trained on RoboRSI executions. Success with absolute and chunk-relative joint targets under the same demonstrations, architecture, and training budget.
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 12: A local skill revision enables execution in the bell-pressing case. Left: the original target-association check rejects the near-boundary box before action. Middle: the revised check accepts the compact target. Right: the post-execution observation from the successful revised run. Boxes mark the detected targets; the simulator supplies the task outcomes.
Figure 13: Observations from the published tap skill in a successful click_bell development trial: the initial view, feature localization, and post-execution view. The outcome is scored by the simulator; images are 320 × 240 pixels.
Figure 14: Physical corrections that required human knowledge. (a) Resetting the scene after a confirmed misplacement; (b) a safe grasp speed specified by the operator; (c) a wrapper released outside the bin; (d) holding protection after termination.
Vision-language-action and world-action models have demonstrated impressive capabilities in robotics, yet generalization to unseen tasks remains challenging. More recently, general-purpose multimodal agents have shown great potential for zero-shot robotic task solving. However, they often incur high execution costs by reasoning and exploring the physical world from scratch. To reduce these costs, we introduce RoboSkill, a framework that connects skill acquisition and reuse through an Explore, Execute, Evolve loop. Within this loop, the agent explores to gather task-relevant information, executes tasks while adapting to feedback, and evolves its skill library based on execution records. It then reuses these skills to guide exploration and execution in the next cycle, closing the loop. To improve loop efficiency, we complement vision with tactile feedback to reduce uncertainty during physical interaction. We further augment textual guidance with reusable code to reduce reasoning overhead during skill reuse. On LIBERO-10, RoboSkill improves first-episode success rates by 12.5--25.0 percentage points and reduces average runtime by 7.6--72.4% across four agents. On real robots, it improves success rates by 8.3 percentage points and reduces average runtime for successful trials by at least 14.4%.
Sicheng Xie, Yitong Chen, Haidong Cao +3
Institute of Trustworthy Embodied AI, Fudan University · Shanghai Innovation Institute · NeoteAI.
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
Traditional robot programming is challenging: it requires orchestrating multimodal perception, managing physical contact dynamics, and handling diverse configurations and execution failures. We introduce ASPIRE (Agentic Skill Programming through Iterative Robot Exploration), a continual learning system that autonomously writes and refines robot control programs in a code-as-policy paradigm while compounding experience into a reusable skill library. ASPIRE discovers skills that persist across tasks, simulation and real-world settings, and embodiments. It operates in an open-ended loop with three components: (1) a closed-loop robot execution engine that exposes fine-grained multimodal traces, enabling autonomous failure diagnosis, repair synthesis, and validation; (2) a continually expanding skill library that distills validated fixes into reusable, transferable knowledge; and (3) evolutionary search that generates diverse task sequences and control programs to explore beyond single-trajectory refinement. ASPIRE surpasses prior methods by up to 77% on LIBERO-Pro manipulation under perturbation, 72% on Robosuite bimanual handover, and 32% on BEHAVIOR-1K long-horizon household tasks. Its accumulated library also enables zero-shot generalization to unseen long-horizon tasks: on LIBERO-Pro Long, ASPIRE achieves 31% success versus 4% for prior methods despite their use of test-time reasoning and retries. Finally, simulation-discovered skills provide initial evidence of sim-to-real transfer, substantially reducing real-robot programming effort across different embodiments and robot APIs.