cs.CVSep 29, 2026

Real2Gym: Building Gyms from Videos, Bringing Skills to Robots

Authors: Kerui Ren, Yingxiang Xu, Kaiwen Song, Linning Xu, Bo Dai, Mulin Yu, Tao Lu

Organizations: Shanghai Artificial Intelligence Laboratory · Shanghai Jiao Tong University · Zhejiang University · University of Science and Technology of China · The Chinese University of Hong Kong · The University of Hong Kong

Abstract

Real-world videos provide rich demonstrations of manipulation, but turning them into reusable robot skills requires visually aligned environments, executable physical interactions, and mechanisms for learning from experience. We introduce Real2Gym, an agentic Real2Sim2Real framework that turns human and robot demonstrations into interactive simulation gyms and brings skills acquired in simulation to physical robots. The Real2Sim module reconstructs editable scenes, aligns objects and cameras with the input, validates demonstrated or retargeted actions through native physics execution, and generates task-conditioned variations with action-feasibility checks. Within these environments, the agent generates executable code for manipulation stages, observes their outcomes, and distills successful attempts and failures into reusable task procedures, object-relative motions, and recovery strategies. Through a shared perception-and-control interface, these skills guide subsequent execution in simulation and on real robots, with motions adapted to current observations and no updates to the underlying model weights. Extensive evaluations demonstrate that Real2Gym enables high-fidelity simulation environment reconstruction, outperforming GPT-6 Astra Direct Mode by 16.7% in success rate with approximately 74.9% fewer policy-execution tokens across these environments, while exceeding it by 33.3% in physical robot execution success rate across four tasks on a real Franka robot.

Figures & tables

Appendix figures & tables4 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Video2Sim2Real: Full-Stack Autonomous Dexterous Skill Acquisition from a Single Human Video

    Jun 7, 2026Yunhai Han, Jianuo Qiu, Linhao Bai +14Human-To-Robot TransferSim-To-Real Gap

  2. Agentic Real2Sim: Physics-based World Modeling with Vision-Language Agents

    Jul 21, 2026Guanxiong Chen, Qianjun Xia, Jiawei Peng +24Physics SimulationDiffusion-Based Vision-Language-Actions

  3. ZeroBot: Learning from Scratch in Minutes with Generative Real2Sim

    Sep 27, 2026Ivan Kapelyukh, Xiaohan Zhang, Stephen James +2Sim-To-Real Reinforcement LearningMulti-Turn Reinforcement Learning