Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
Organizations: UC Berkeley · Amazon FAR · MIT · University of Chicago
Abstract
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
Figures & tables
| Source task | Constructed practice task |
|---|---|
| Organize sunglasses | Close the case; opening and loading are encoded in the initial state. |
| Organize makeup | Close the drawer; preceding manipulation is already satisfied. |
| Roll towels | Fold once; a cloth manipulation analogue with a different final shape. |
| No corresponding source task | Transfer an object between arms using a task specification. |
| Video Analyzer | Privileged Agent | Success |
|---|---|---|
| Yes | Yes | 78.2% |
| No | Yes | 51.8% |
| Yes | No | 50.9% |
| No | No | 50.9% |
| Task | Before repair | After repair |
| Two-arm lift | 6/10 | 10/10 |
| Place dish on rack | 5/10 | 9/10 |
| Nut assembly | 5/10 | 8/10 |
| Place bottles in bin | 10/10 | 10/10 |
| Total | 26/40 | 37/40 |
| Success rate | 65.0% | 92.5% |
| CaP-Agent0 | RPG | ||
|---|---|---|---|
| Task / Stage | Gemini | GPT-6 | Gemini |
| Store ball in drawer | 3/10 | 4/10 | 10/10 |
| Place ball in drawer | 7/10 | 6/10 | 10/10 |
| Fold towel | 0/10 | 9/10 | 10/10 |
| Transfer bowl | 0/10 | 1/10 | 10/10 |
| Move bowl to center | 6/10 | 8/10 | 10/10 |
| Task | Split | R1 | R2 | R3 | R4 | R5 | R6 | R7 | R8 | R9 | R10 | R11 | R12 | R13 | R14 | R15 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Extract from fixture | D | 2/10 | 5/10 | 5/10 | 10/10 | 10/10 | 10/10 | 9/10 | 10/10 | 9/10 | 10/10 | 9/10 | 10/10 | 8/10 | 10/10 | 10/10 |
| Insert into fixture | D | 2/10 | 3/10 | 4/10 | 7/10 | 9/10 | 9/10 | 8/10 | 10/10 | 9/10 | 10/10 | 10/10 | 8/10 | 10/10 | 9/10 | 10/10 |
| Place dish on rack | D | 0/10 | 0/10 | 0/10 | 2/10 | 3/10 | 6/10 | 5/10 | 5/10 | 8/10 | 4/10 | 9/10 | 8/10 | 9/10 | 9/10 | 10/10 |
| Serve onto plate | H | 1/10 | 5/10 | 8/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 9/10 |
| Sort cubes (extended) | D | 0/10 | 0/10 | 0/10 | 3/10 | 2/10 | 8/10 | 10/10 | 10/10 | 9/10 | 10/10 | 10/10 | 9/10 | 10/10 | 8/10 | 10/10 |
| Lift cube | D | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 |
| Task | R1 | R4 | R8 | R10 | R11 | R12 | R15 |
|---|---|---|---|---|---|---|---|
| Serve onto plate | 1/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 9/10 |
| Two-arm handover | 3/10 | 9/10 | 7/10 | 9/10 | 9/10 | 7/10 | 10/10 |
| Transfer object | 4/10 | 10/10 | 10/10 | 10/10 | 10/10 | 10/10 | 9/10 |
| Close sunglasses case | 0/10 | 10/10 | 10/10 | 8/10 | 10/10 | 10/10 | 10/10 |
| Grasp plate | 2/10 | 7/10 | 9/10 | 10/10 | 10/10 | 6/10 | 6/10 |
| Failure mode | Typical failure | Learned correction | Representative tasks |
|---|---|---|---|
| Localization and geometry uncertainty | The policy acts on a stale, ambiguous, or poorly localized object, receptacle, or support surface. | Re-observe after scene changes; reconstruct geometry from multiple visible points and depth; invalidate stale targets after contact or failed motion. | Place bottles in bin, Insert into fixture, Place dish on rack |
| Grasp identity and retention | A blocked or object-sized gripper opening is treated as proof that the desired object is held, even when the grasp is empty, mixed, or on the wrong object. | Combine gripper state with wrist/top observations and a short test lift; require visible target co-motion before transport; preserve uncertain grasps rather than opening blindly. | Extract from fixture, Nut assembly, Stack blocks |
| Incomplete path feasibility | An endpoint is reachable, but an approach, lift, transit, release, or retreat waypoint along the actual manipulation route is not. | Check the complete manipulation path before committing, including approach, grasp, lift, transport, release, and retreat, using the actual arm configuration and yaw. | Place bottles in bin, Place dish on rack, Stack blocks |
| End-effector-centric geometry | The fingertip reaches the nominal target while the relevant object feature does not: e.g., the bottle misses the opening or the nut hole misses the peg. | Track the geometry of the manipulated object and its task-relevant feature; estimate object-to-gripper offsets and align the object, hole, bottom surface, or footprint rather than the fingertip. | Place bottles in bin, Nut assembly, Insert into fixture |
| Contact and support calibration | Motion is commanded relative to the fingertip without accounting for pad extension, object extent, table height, or support geometry, causing penetration or ineffective contact. | Estimate the support surface and relevant tool/object offset; use surface-relative heights and small monitored increments instead of large open-loop contact motions. | Spill wipe, Fold towel |
| Multi-effector coordination | Each arm is controlled locally without verifying that both contacts are valid or that the required relation between the two arms is maintained. | Verify both grasps, perform a paired test motion, preserve relative geometry, and use synchronized incremental motions with an explicit relation tolerance. | Two-arm lift |
| R1 | R2 | R3 | R4 | R5 | R6 | R7 | R8 | R9 | R10 | R11 | R12 | R13 | R14 | R15 | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Fold towel | 0/10 | 1/10 | 0/10 | 0/10 | 0/10 | 1/10 | 4/10 | 5/10 | 4/10 | 4/10 | 3/10 | 3/10 | 3/10 | 4/10 | 3/10 |
| Fold towel (hard) | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 | 0/10 |
| Split | # Tasks | Examples |
|---|---|---|
| Development | 17 | insertion, stacking, wiping, two-arm lift |
| Validation | 1 | restack cubes |
| Feedback-held-out | 4 | serving, handover, case closing |