Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting, where task-specific simulations are constructed from target deployment scenarios and used as low-cost environments for iterative policy improvement before transfer back to the physical world. EmbodiRSI uses policy execution feedback to guide subsequent experience acquisition and policy updates. Two complementary mechanisms close this loop: Collaborative Error Correction generates agent-assisted corrective trajectories from policy-reached states, while Adaptive Data Collection directs expert demonstration generation toward the current policy's weaknesses. The task-specific simulation serves as a reusable workspace for policy warm-up, repeatable evaluation, failure diagnosis, and targeted data generation across successive RSI rounds. Across three tabletop environments and 14 subtasks, EmbodiRSI increases scene-balanced autonomous simulation success from 50.4% to 83.5% over two RSI updates. With 400 adaptive simulated trajectories and only ten real-world refinement trajectories per subtask, EmbodiRSI achieves 83.1% scene-balanced autonomous real-world success, compared with 75.0% for adaptation using 200 real-world demonstrations per subtask. These results demonstrate that feedback-driven recursive improvement in deployment-specific simulations can enable data-efficient adaptation of embodied policies to physical environments.
Figures & tables
Figure 1: Overview of EmbodiRSI. Real-world observations and task goals guide the construction of an executable simulation environment and the generation of validated demonstrations. Within this environment, policy rollout feedback guides Collaborative Error Correction and Adaptive Data Collection. The resulting experience updates the policy for the next round, forming a recursive self-improvement loop. A small real-world human-in-the-loop (HIL) dataset then refines the policy for autonomous deployment.
Figure 2: The recursive self-improvement loop in EmbodiRSI. The rollout monitor combines multi-camera observations and simulator states to trigger recovery from the current state without resetting. Failure analysis guides targeted data collection. Successful rollouts, including those completed after correction, and newly generated demonstrations are accumulated to update the policy for the next round. External recovery is used only during training; evaluation is autonomous.
Figure 3: Paired physical and simulated experimental settings. Columns from left to right show Office Table (S1), Kitchen Table (S2), and Household Table (S3). The top row shows the physical workspaces, and the bottom row shows their corresponding simulated environments. Each column presents the same scene in the real world and simulation.
Figure 4: Autonomous policy performance across RSI rounds. Thin lines show individual subtasks; markers and error bars indicate scene means ± sample SD across subtasks. Discrete rounds R0/R1/R2 use 50/200/400 retained simulated training trajectories per subtask, respectively.
Method
Simulation
R2S2R
R2S2R + Human in the Loop
EmbodiRSI w/o RSI
59.3
46.3
67.9
EmbodiRSI
79.9
59.2
82.1
Gain (pp)
+20.6
+12.9
+14.2
Table 1: Effect of RSI under equal data budgets. Scene-balanced autonomous success (%) across S1 and S3 at 400 simulated trajectories per subtask. Complete 200/400 comparisons, scene-level dispersion and task-level exceptions are retained in Appendix E .
Training setting
S1: Office n=6 tasks
S2: Kitchen n=4 tasks
S3: Household n=4 tasks
Scene-balanced mean
R1
43.3±22.5
60.0±8.2
52.5±9.6
51.9
R1+10
63.3±13.7
82.5±5.0
75.0±17.3
73.6
R2
53.3±32.0
70.0±14.1
65.0±12.9
62.8
R2+10
81.7±13.3
85.0±12.9
82.5±9.6
83.1
Real-only 10
13.3±13.7
45.0±5.8
20.0±8.2
26.1
Real-only 100
56.7±15.1
62.5±5.0
52.5±5.0
57.2
Table 2: Data-efficient real-world adaptation. Success rates (%) are mean ± sample SD across subtasks; the final column weights scenes equally. R1/R2 use 200/400 simulated trajectories per subtask; +10 adds ten real human-in-the-loop (HIL) trajectories. Real-only labels give demonstration counts per subtask. Each subtask/setting has ten autonomous test trials.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Figure B.1: From deployment observations to task-defining scenes. Object reconstruction and support-aware collision rectification produce the initial scene. The scene agent iteratively refines object poses using multi-view feedback and geometric validation to obtain the target scene. The scene pair defines manipulation goals for trajectory generation.
Subtask
Goal
S1-1
Place glasses on the stand.
S1-2
Put the red-cap marker in the green cup.
S1-3
Put the yellow-cap marker in the basket.
S1-4
Place the mouse on the mouse pad.
S1-5
Put the red pen in the green box.
S1-6
Put the black pen in the pen holder.
Appendix
Table C.1: Manipulation subtask definitions. Initial and goal relations are kept fixed across the corresponding simulation and real-world task definitions.
Figure C.1: Qualitative examples of scene reconstruction and organization. a,b , Mobile-camera and robot-camera observations paired with reconstructed scenes; views are not pixel-aligned. c , Initial and target layouts illustrating object-level manipulation goals.
Figure C.2: Iterative scene organization through critique, editing, and validation. The critic compares rendered views and object states with target relations to guide pose adjustments. Collision and support checks validate each edit before re-rendering. The accepted layout defines object-level goals for trajectory generation.
Figure C.3: Qualitative input–reconstruction pairs. a–d , Mobile-camera, web, robot-camera, and AI-generated inputs, respectively. Each pair shows the input on the left and reconstruction on the right. Examples extend beyond the evaluation benchmark; views are not aligned in pose or scale.
Data item
Source and learning role
Budget accounting
D0
Initial validated simulation demonstrations.
50 per subtask at R0.
Dkmon
Policy-generated trajectories recorded during monitored training collection.
Part of rollout-derived data.
Dkrec
Corrections generated from eligible intermediate states reached during collection.
Part of rollout-derived data.
Dkroll
Retained data assembled from monitored and corrective sources.
50 added at R1; 100 at R2.
Dknew
Newly generated demonstrations after collection configurations are updated.
100 added at each round.
Real HIL data
Real-world training rollouts with operator corrections.
10 per subtask; pooled within scene.
Appendix
Table D.1: Experience sources and retained-data notation. The notation separates monitored training rollouts, corrective data and newly generated demonstrations. The reported quotas distinguish newly generated from rollout-derived data, without a separate count for each rollout-derived source. No numerical allocation between base and targeted collection is inferred from those quotas. Policy and correction symbols denote sources within episodes, not separately counted episode datasets.
Stage
Newly generated demonstrations
Retained rollout- derived data
Cumulative total
R0
50
0
50
R1
100
50
200
R2
100
100
400
Appendix
Table D.2: Simulation training schedule per subtask. The two data columns are additions at the indicated stage; the total is cumulative. For R1/R2, newly generated demonstrations are Dknew and retained rollout-derived data are Dkroll . R0 contains the initial generated demonstrations. Counts do not include performance-test episodes.
Dataset / evaluation
S1 (6 tasks)
S2 (4 tasks)
S3 (4 tasks)
R0 simulation training
300
200
200
R1 simulation training
1,200
800
800
R2 simulation training
2,400
1,600
1,600
HIL+10 real training
60
40
40
Real-only 10 training
60
40
40
Real-only 100 training
600
400
400
Appendix
Table D.3: Scene-level trajectory budgets. Counts equal the reported per-subtask budget multiplied by the number of subtasks covered by the scene policy. Real-only 10 and HIL+10 contain the same number of trajectories but are different training settings. The last row is evaluation, not training; each real-world subtask is tested in ten autonomous rollouts.
Setting
Real-only
Static simulation
EmbodiRSI (adaptive)
Scene policy
One per scene
One per scene
One per scene
Architecture
Same policy architecture across compared groups
Initial simulation data
Not used
Shared 50 demonstrations per subtask
Same initial dataset as Static
Subsequent reset sampling
Real collection
Original configuration distribution
Feedback-guided configuration update
Simulation trajectory budget
None
200 or 400
200 (R1) or 400 (R2)
Training recipe
Same policy-training recipe across compared groups
Appendix
Table D.4: Training and collection settings. Budgets are retained trajectories per subtask. The table distinguishes the adaptation data sources and configuration-sampling procedures. Equal real-data counts do not imply identical trajectories.
Task
R0 (50)
R1 (200)
R2 (400)
S1: Office Table
Glasses → stand
44.5
64.9
80.7
Red-cap marker → green cup
34.0
50.1
75.3
Yellow-cap marker → basket
26.1
30.9
57.0
Mouse → mouse pad
33.3
70.1
85.1
Red pen → green box
17.5
36.0
72.0
Appendix
Table E.1: Task-level simulation success (%). Scene summaries report mean ± sample SD across subtasks (S1: six; S2/S3: four each), not confidence intervals. The final row weights scenes equally and reports no SD. Numbers in column headers denote simulated training trajectories per subtask.
Task
R1
R1+10
R2
R2+10
Real 10
Real 100
Real 200
S1: Office Table
Glasses → stand
7/10
8/10
10/10
10/10
3/10
7/10
7/10
Red-cap marker → green cup
3/10
6/10
3/10
7/10
0/10
4/10
7/10
Yellow-cap marker → basket
6/10
6/10
6/10
9/10
1/10
5/10
7/10
Mouse → mouse pad
6/10
8/10
8/10
9/10
3/10
8/10
8/10
Red pen → green box
2/10
5/10
2/10
7/10
1/10
5/10
7/10
Appendix
Table E.2: Task-level real-world performance. Task entries are successes out of ten autonomous tests. Scene summaries report mean success (%) ± sample SD across subtasks; the final row weights scenes equally, without an SD. R1/R2 use 200/400 simulated training trajectories per subtask; +10 adds ten real HIL trajectories per subtask. Real 10/100/200 denote real-only training with 10/100/200 demonstrations per subtask.
Figure E.1: Task-level real-world transfer and sampling precision. a , Success rates for R2+10 and Real-only 200, with pointwise 95% Wilson intervals from ten autonomous tests per task and setting. b , Observed differences matched by subtask identity; positive values favor R2+10. R2+10 uses 400 simulated and ten real-world training trajectories per subtask, versus 200 real-world demonstrations for Real-only 200. Task IDs follow Table C.1 .
Scene
Sim. per subtask
HIL per subtask
Static collection
Adaptive collection
Δ (pp)
a Simulation evaluation
S1
200
—
40.7±14.6
49.0±15.8
+8.3±12.3
400
—
51.6±10.2
74.4±9.7
+22.9±9.3
S3
200
—
58.1±28.1
78.5±11.0
+20.5±18.0
400
—
67.0±23.2
85.4±8.7
+18.4±16.0
b Real-world evaluation
Appendix
Table E.3: Scene-level matched-count comparisons. a , Simulation evaluation without physical refinement. b , Real-world performance before and after HIL refinement. Entries are mean autonomous success (%) ± sample SD across subtasks ( n=6 for S1 and n=4 for S3). Δ is the mean ± SD of task-matched differences, in percentage points. Equal counts do not equate total acquisition cost.
Initial
200 per subtask
400 per subtask
Task
R0 (50)
Static
R1
Static
R2
S1: Office Table
Glasses → stand
44.5
32.6
64.9
41.4
80.7
Red-cap marker → green cup
34.0
47.2
50.1
59.7
75.3
Yellow-cap marker → basket
26.1
30.0
30.9
38.6
57.0
Mouse → mouse pad
33.3
67.6
70.1
59.1
85.1
Appendix
Table E.4: Task-level simulation collection comparison. R0 is the shared 50-trajectory-per-subtask initialization. Static and adaptive policies have matched cumulative training budgets per subtask. Scene rows report mean ± sample SD across subtasks. The final row weights S1 and S3 equally; it is distinct from the three-scene average in the simulation-evolution table.
Task
Static 200
Static 200+10
R1
R1+10
Static 400
Static 400+10
R2
R2+10
S1: Office Table
Glasses → stand
7/10
7/10
7/10
8/10
7/10
8/10
10/10
10/10
Red-cap marker → green cup
2/10
4/10
3/10
6/10
3/10
6/10
3/10
7/10
Yellow-cap marker → basket
4/10
4/10
6/10
6/10
6/10
8/10
6/10
9/10
Mouse → mouse pad
4/10
6/10
6/10
8/10
8/10
8/10
8/10
9/10
Red pen → green box
1/10
4/10
2/10
5/10
1/10
4/10
2/10
7/10
Appendix
Table E.5: Task-level real-world collection comparison. Task entries are successes out of ten autonomous test rollouts. All budgets are training trajectories per subtask, and one policy covers each scene. R1 and R2 correspond to simulation budgets of 200 and 400; +10 adds ten real-world HIL training trajectories per subtask for either collection method. Summary rows report mean success (%) ± sample SD across subtasks within each scene.
Figure E.2: Illustrative autonomous real-world executions. a , Real-only adaptation with 200 real demonstrations per subtask. b , EmbodiRSI with 400 simulated and ten real HIL trajectories per subtask. All tests use the learned policy alone. Quantitative results are reported in Table 2 .
Scaling data volume and diversity is critical for generalizing embodied intelligence. While synthetic data generation offers a scalable alternative to expensive physical data acquisition, transferring robotic manipulation policies from simulation to the real world (sim-to-real) remains a formidable challenge due to the domain gap. This paper presents HyperSim, a holistic framework spanning from synthetic data generation to policy training and seamless real-world deployment. To systematically bridge the sim-to-real gap, HyperSim is realized through three core pillars: high-fidelity environment synthesis, adversarial trajectory generation, and sim-and-real co-training. Collectively, these modules address domain discrepancies by enhancing visual fidelity, expanding data coverage, and enforcing domain-invariant representations. We rigorously validate HyperSim through a large-scale empirical study involving 400 real-world task executions across two representative manipulation models. Assessed across three fine-grained metrics, our complete pipeline achieves remarkable sim-to-real success rates of 80% and 95% with ACT and π_{0}, respectively. Furthermore, policies trained on our adversarial trajectories exhibit significantly enhanced robustness against dynamic uncertainties, achieving a 35% higher completion rate under physical perturbations.
Junyi Dong, Haotian Luo, Ziwei Xu +11
CloudRobo Lab, Huawei Cloud Computing Technologies Co.,Ltd. · Shanghai Jiao Tong University · The University of Hong Kong
Adapting visuomotor policies to new manipulation tasks often requires substantial manual engineering or teleoperated data collection. Simulation can provide task-specific data at scale, but constructing the scene, designing expert behavior, and configuring data generation still require significant per-task effort. We present ARSTAG, an agentic Real2Sim2Real system that turns a single RGB image and a natural-language instruction directly into robot policy-learning data. A hierarchy of language agents constructs a task-scoped simulation scene, generates robot-feasible demonstrations, and expands the training distribution through task-consistent randomization, while a coordinator agent manages cross-stage feedback and recovery. Across seven manipulation tasks spanning grasping, placement, and stacking, the ARSTAG-generated demonstrations enable sim-to-real transfer of three visuomotor policy architectures to a dual-arm robot, with pi0.5 achieving an average real-world success rate of 74.6%. Ablations show that task-consistent randomization substantially improves robustness, and policy performance increases with generated dataset size. Project webpage: https://boweili666.github.io/ARSTAG/.
Bowei Li, Yuner Zhang, Changliu Liu
Carnegie Mellon University, Pittsburgh, PA 15213, USA
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/