Real2Gym: Building Gyms from Videos, Bringing Skills to Robots
Authors: Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai, Mulin Yu, Tao Lu
Organizations: Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Zhejiang University · University of Science and Technology of China · The Chinese University of Hong Kong · The University of Hong Kong
Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.
Figures & tables
Figure 1: Real2Gym converts human or robot videos into visually aligned and physically executable Blender and MuJoCo environments, where reconstructed interactions are validated under native physics and expanded into feasible task variations. Rather than treating simulation as the endpoint, the agent uses these gyms to execute, fail, reflect, and accumulate reusable skills. The resulting skills are re-grounded from human demonstrations and transferred to real robots through a shared control interface, enabling Real2Sim2Real self-improvement. Project page: https://real2gym.github.io/ .
Figure 2: Overview of Real2Gym. Real2Gym reconstructs aligned Blender and MuJoCo scenes from human or robot demonstrations, refines them through event-driven correction and physics validation, and augments them into diverse interactive gyms. The agent then executes operation-stage code and distills feedback into reusable skills for simulation and real-robot deployment.
Method
Content alignment ↑
Viewpoint alignment ↑
Action fidelity ↑
Simulation success score ↑
DROID
GPT-5.6 Sol xhigh
50.00
24.17
63.83
48.96
GPT-6 Astra Medium
55.00
30.00
76.08
71.08
Ours
66.25
70.83
83.17
80.88
EgoDex
GPT-5.6 Sol xhigh
55.42
34.17
54.83
46.77
GPT-6 Astra Medium
57.58
38.67
65.75
56.59
Ours
75.08
73.75
86.25
85.76
Table 1: Quantitative comparison of Real2Sim reconstruction. Results are averaged over 12 DROID and 12 EgoDex scenes. Bold and underlined denote best and second-best results.
Figure 3: Qualitative comparison of Real2Sim reconstruction. Five manipulation scenes from DROID and EgoDex are shown. Rows present GPT-5.6 Sol, GPT-6 Astra, Real2Gym, and the source video, from top to bottom.
DROID
EgoDex
Method
SR (%) ↑
Responses ↓
Tokens (M) ↓
Time (min) ↓
SR (%) ↑
Responses ↓
Tokens (M) ↓
Time (min) ↓
GPT-5.6 Sol xhigh
58.33
71.42
8.57
25.72
41.67
105.50
13.63
28.73
GPT-6 Astra Medium
75.00
28.33
1.42
7.81
66.67
33.75
2.56
10.31
Ours
75.00
16.17
0.64
8.60
83.33
16.67
0.66
9.09
Ours (w/ skills)
91.67
13.33
0.51
6.32
83.33
14.25
0.49
7.05
Table 2: Quantitative comparison of agent policies. Mean results over 12 tasks per dataset, including failures. Bold and underlined denote best and second-best results.
Figure 4: Qualitative comparison of zero-shot agent execution. Rows depict the human demonstration, GPT-6 Astra, GPT-5.6 Sol, and our agent on the bowl-stacking task. Our agent successfully completes the stack, whereas both baselines fail, with red boxes highlighting key failure cases.
Figure 5: Qualitative examples of skill reuse in cabinet manipulation and adapter placement. Earlier exploratory executions are compared with subsequent skill-conditioned runs.
Figure 6: Real-world deployment and Real2Sim2Real self-evolution. (a) Four real-world manipulation tasks and zero-shot success rates of our method versus Direct Mode. (b) Failure-driven self-evolution: reconstructing a digital twin from a human demonstration, acquiring skills through simulation, and deploying the evolved skills to improve real-world success.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
GPT-5.6 Sol xhigh
GPT-6 Astra Medium
Real2Gym (Ours)
Task
Content alignment
Viewpoint alignment
Action fidelity
Simulation success score
Content alignment
Viewpoint alignment
Action fidelity
Simulation success score
Content alignment
Viewpoint alignment
Action fidelity
Simulation success score
DROID
D01
55.00
30.00
85.00
76.50
60.00
30.00
80.00
80.00
80.00
60.00
93.00
93.00
D02
55.00
25.00
85.00
85.00
60.00
30.00
85.00
85.00
70.00
80.00
88.00
88.00
D03
55.00
20.00
88.00
88.00
60.00
30.00
80.00
80.00
70.00
80.00
88.00
88.00
D04
60.00
35.00
90.00
90.00
60.00
30.00
78.00
78.00
75.00
80.00
86.00
86.00
Appendix
Table 3: Per-task Real2Sim evaluation on DROID and EgoDex. Methods are evaluated across content alignment, viewpoint alignment, action fidelity, and simulation success score (0–100 scale, ↑ ). Bold indicates the top performance per task and metric, including ties.
GPT-5.6 Sol xhigh
DROID
EgoDex
Task
Success
Resp. ↓
Tokens (M) ↓
Time (min) ↓
Task
Success
Resp. ↓
Tokens (M) ↓
Time (min) ↓
D1
✓
52
6.88
25.09
E1
✓
75
9.12
36.88
D2
✓
60
6.25
27.78
E2
✓
59
6.12
14.76
D3
✓
81
9.32
29.77
E3
×
62
8.25
13.15
D4
×
31
4.10
11.60
E4
✓
42
4.27
15.44
Appendix
Table 4: Per-task efficiency and success results on DROID and EgoDex. Responses denote model responses, and tokens are reported in millions. Lower is better for response count, token usage, and time.
Figure 7: Real-world manipulation trajectories. The baseline and our method are shown across four tasks, with wrist-view insets. The shelf task shows our method before and after skill evolution.
Figure 8: Qualitative comparison across models and reasoning effort levels on an EgoDex assembly task. The top row depicts the source human demonstration, while subsequent rows show the 12 experimental configurations. Columns are indexed by source-video timestamps. Successful trajectories are aligned using recorded event correspondences, whereas failed attempts are approximately aligned by the attempted action phase. Repeated frames marked with FAILED retain a frozen state for visual comparison and do not imply continued execution.
Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception errors and embodiment gap. To address this, we introduce Video2Sim2Real, a full-stack framework for autonomous skill acquisition from a single human manipulation video. Our framework first uses off-the-shelf foundation models to reconstruct a simulator-ready digital twin and extract robot and object motion priors. Rather than treating the extracted robot motion as a reliable reference throughout execution, our key idea is to recover and leverage the most fundamental sources of supervision from the demonstrated skill: We identify object-centric keyframes to optimize the corresponding robot configurations using object information from the simulator, and use these configurations as anchors that refine the robot motion such that it ultimately has the desired impact on the environment. To bridge the remaining sim-to-real gap, we introduce a sim-to-real strategy that decouples robustness to noisy and incomplete perception from variations in hand-object interaction dynamics. Specifically, we learn to recalibrate robot configurations from noisy real-world point clouds via IL, and leverage residual RL to perform local finger-level adaptations to ensure for robust and effective interactions. Finally, a collision-aware motion planning module enables spatial generalization to novel object configurations. Across several everyday manipulation tasks, Video2Sim2Real improves simulated task success, safety, and trajectory coherence over numerous baselines, and achieves better sim-to-real transfer than existing techniques. These results demonstrate a promising path toward autonomous dexterous skill acquisition from human videos.
Yunhai Han, Jianuo Qiu, Linhao Bai +14
1Georgia Institute of Technology · University of Pennsylvania · 3Toyota Research Institute +1
Real-to-sim conversion for robotic interaction with objects remains labor-intensive because it requires more than visual reconstruction: a streamlined real2sim process must recover scene geometries and object states, infer physical parameters, and assemble actors, objects, cameras, poses, and trajectories into a runnable physical simulation. Today this process still depends on brittle workflow glue across visual perception tools and simulators: manual tuning of visual foundation models, mesh cleanup, coordinate frame alignments, etc. We introduce \textit{Agentic Real2Sim}, a framework for generalized physical world modeling with vision-language agents that converts a real-world recording of object-robot interaction into a simulatable episodic twin, and connects the resulting twin to downstream policy fine-tuning and evaluation. We evaluate Agentic Real2Sim on rigid-object manipulation, deformable-object interaction, and humanoid motion scenes, spanning domains that are usually handled by separate Real2Sim pipelines. The framework's agentic decisions can be driven by an open-weight VLM backend at a small fraction of the cost of frontier models, while attaining a comparable conversion success rate. The framework further supports custom scene conversion, fine-tuning of a pretrained policy with data generated from converted episodes, and works effectively as a surrogate for real-world policy evaluation. The project site, including code is available at https://agentic-real2sim.github.io.
Guanxiong Chen, Qianjun Xia, Jiawei Peng +24
University of British Columbia · National University of Singapore · Johns Hopkins University +4
We present ZeroBot, a real2sim framework for learning a robot manipulation task from scratch in minutes under challenging conditions: zero human demonstrations, zero policy pre-training, and zero known object models. Given only a single view of an object and a goal pose for that object, ZeroBot uses image-to-3D generative models to obtain a complete object mesh, which is used in simulation for large-scale parallel reinforcement learning. To accelerate training, we introduce an action space which leverages the generated geometry and learned value function to sample states involving robot-object contact. When evaluated on real-world tasks including grasping, pushing, articulated object interaction, and multi-stage manipulation, ZeroBot achieves an 87% success rate with an average training time of 119 seconds. These results show the value of using image-to-3D models in a real2sim framework for rapid, autonomous robot learning.
Ivan Kapelyukh, Xiaohan Zhang, Stephen James +2
Imperial College London · Robotics and AI Institute