Real2Gym: Building Gyms from Videos, Bringing Skills to Robots
Authors: Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai, Mulin Yu, Tao Lu
Organizations: Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Zhejiang University · University of Science and Technology of China · The Chinese University of Hong Kong · The University of Hong Kong
Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.
Figures & tables
Figure 1: Real2Gym converts human or robot videos into visually aligned and physically executable Blender and MuJoCo environments, where reconstructed interactions are validated under native physics and expanded into feasible task variations. Rather than treating simulation as the endpoint, the agent uses these gyms to execute, fail, reflect, and accumulate reusable skills. The resulting skills are re-grounded from human demonstrations and transferred to real robots through a shared control interface, enabling Real2Sim2Real self-improvement. Project page: https://real2gym.github.io/ .
Figure 2: Overview of Real2Gym. Real2Gym reconstructs aligned Blender and MuJoCo scenes from human or robot demonstrations, refines them through event-driven correction and physics validation, and augments them into diverse interactive gyms. The agent then executes operation-stage code and distills feedback into reusable skills for simulation and real-robot deployment.
Method
Content alignment ↑
Viewpoint alignment ↑
Action fidelity ↑
Simulation success score ↑
DROID
GPT-5.6 Sol xhigh
50.00
24.17
63.83
48.96
GPT-6 Astra Medium
55.00
30.00
76.08
71.08
Ours
66.25
70.83
83.17
80.88
EgoDex
GPT-5.6 Sol xhigh
55.42
34.17
54.83
46.77
GPT-6 Astra Medium
57.58
38.67
65.75
56.59
Ours
75.08
73.75
86.25
85.76
Table 1: Quantitative comparison of Real2Sim reconstruction. Results are averaged over 12 DROID and 12 EgoDex scenes. Bold and underlined denote best and second-best results.
Figure 3: Qualitative comparison of Real2Sim reconstruction. Five manipulation scenes from DROID and EgoDex are shown. Rows present GPT-5.6 Sol, GPT-6 Astra, Real2Gym, and the source video, from top to bottom.
DROID
EgoDex
Method
SR (%) ↑
Responses ↓
Tokens (M) ↓
Time (min) ↓
SR (%) ↑
Responses ↓
Tokens (M) ↓
Time (min) ↓
GPT-5.6 Sol xhigh
58.33
71.42
8.57
25.72
41.67
105.50
13.63
28.73
GPT-6 Astra Medium
75.00
28.33
1.42
7.81
66.67
33.75
2.56
10.31
Ours
75.00
16.17
0.64
8.60
83.33
16.67
0.66
9.09
Ours (w/ skills)
91.67
13.33
0.51
6.32
83.33
14.25
0.49
7.05
Table 2: Quantitative comparison of agent policies. Mean results over 12 tasks per dataset, including failures. Bold and underlined denote best and second-best results.
Figure 4: Qualitative comparison of zero-shot agent execution. Rows depict the human demonstration, GPT-6 Astra, GPT-5.6 Sol, and our agent on the bowl-stacking task. Our agent successfully completes the stack, whereas both baselines fail, with red boxes highlighting key failure cases.
Figure 5: Qualitative examples of skill reuse in cabinet manipulation and adapter placement. Earlier exploratory executions are compared with subsequent skill-conditioned runs.
Figure 6: Real-world deployment and Real2Sim2Real self-evolution. (a) Four real-world manipulation tasks and zero-shot success rates of our method versus Direct Mode. (b) Failure-driven self-evolution: reconstructing a digital twin from a human demonstration, acquiring skills through simulation, and deploying the evolved skills to improve real-world success.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
GPT-5.6 Sol xhigh
GPT-6 Astra Medium
Real2Gym (Ours)
Task
Content alignment
Viewpoint alignment
Action fidelity
Simulation success score
Content alignment
Viewpoint alignment
Action fidelity
Simulation success score
Content alignment
Viewpoint alignment
Action fidelity
Simulation success score
DROID
D01
55.00
30.00
85.00
76.50
60.00
30.00
80.00
80.00
80.00
60.00
93.00
93.00
D02
55.00
25.00
85.00
85.00
60.00
30.00
85.00
85.00
70.00
80.00
88.00
88.00
D03
55.00
20.00
88.00
88.00
60.00
30.00
80.00
80.00
70.00
80.00
88.00
88.00
D04
60.00
35.00
90.00
90.00
60.00
30.00
78.00
78.00
75.00
80.00
86.00
86.00
Appendix
Table 3: Per-task Real2Sim evaluation on DROID and EgoDex. Methods are evaluated across content alignment, viewpoint alignment, action fidelity, and simulation success score (0–100 scale, ↑ ). Bold indicates the top performance per task and metric, including ties.
GPT-5.6 Sol xhigh
DROID
EgoDex
Task
Success
Resp. ↓
Tokens (M) ↓
Time (min) ↓
Task
Success
Resp. ↓
Tokens (M) ↓
Time (min) ↓
D1
✓
52
6.88
25.09
E1
✓
75
9.12
36.88
D2
✓
60
6.25
27.78
E2
✓
59
6.12
14.76
D3
✓
81
9.32
29.77
E3
×
62
8.25
13.15
D4
×
31
4.10
11.60
E4
✓
42
4.27
15.44
Appendix
Table 4: Per-task efficiency and success results on DROID and EgoDex. Responses denote model responses, and tokens are reported in millions. Lower is better for response count, token usage, and time.
Figure 7: Real-world manipulation trajectories. The baseline and our method are shown across four tasks, with wrist-view insets. The shelf task shows our method before and after skill evolution.
Figure 8: Qualitative comparison across models and reasoning effort levels on an EgoDex assembly task. The top row depicts the source human demonstration, while subsequent rows show the 12 experimental configurations. Columns are indexed by source-video timestamps. Successful trajectories are aligned using recorded event correspondences, whereas failed attempts are approximately aligned by the attempted action phase. Repeated frames marked with FAILED retain a frozen state for visual comparison and do not imply continued execution.