Human videos offer a scalable source of demonstrations for dexterous robot manipulation. However, existing human-to-simulation-to-robot (Human2Sim2Robot) pipelines rely on predefined procedures that struggle to accommodate diverse object properties and interactions, particularly those involving articulated and deformable objects. We introduce DexAgent, an agentic Human2Sim2Robot framework that converts a single egocentric human video and a task prompt into physically grounded robot trajectories for policy training. It operates through four stages: semantic understanding of human videos, property-based simulation reconstruction, robot trajectory optimization, and robot data generation. At each stage, DexAgent adapts its approach to the task and object properties by selecting suitable skills from its tool library or developing new ones when needed. Property-specific verifiers assess stage outcomes for physical validity and task-specific requirements and provide feedback for refinement, preventing error propagation through the workflow. This adaptive, verification-guided process allows DexAgent to process diverse objects and long-horizon tasks. In the final stage, DexAgent varies object and robot states in simulation to generate diverse robot trajectories from a single human video, then retextures the rendered observations to facilitate sim-to-real transfer. Newly developed skills and verifiers are retained in its tool library, making it self-evolving to accumulate reusable capabilities. This reduces processing time as DexAgent encounters more human videos. Across eleven real-world tasks, policies trained with DexAgent-generated data achieve a 3.5x higher success rate than competing baselines. Project website: https://dexagent123.github.io/.
Figures & tables
Fig. 2: DexAgent converts a single egocentric human video and a task prompt into robot training data through four stages. Stage 1 decomposes the task into subgoals and identifies the manipulated objects and their key properties. Stage 2 reconstructs the objects in simulation using property-specific skills (blue) from the tool library, creating new skills when needed. Stage 3 generates robot-hand trajectories by combining optimization guided by human motion priors with code-generated motion. Stage 4 varies viewpoints and object states to produce diverse simulation data, inpaints the original human video, and retextures the augmented simulation videos for policy training.
Rigid
Articulated
Method
F-5 ↑
F-10 ↑
CD ↓
F-5 ↑
F-10 ↑
CD ↓
HO [ 2 ]
0.28
0.51
3.86
0.29
0.47
1.30
IHOI [ 33 ]
0.42
0.70
2.70
0.32
0.47
1.47
HORSE [ 21 ]
0.26
0.45
6.69
0.19
0.34
1.91
MCC-HO [ 29 ]
0.52
0.78
1.36
0.35
0.55
1.21
G-HOP [ 34 ]
0.69
0.91
0.63
0.07
0.09
1.23
TABLE I: Reconstruction accuracy on HOI4D, scored separately for rigid and articulated objects. Evaluation metrics are F-5 and F-10, which are F-scores at two distance thresholds, and CD, which is Chamfer distance. Best entry in each column in bold.
Method
Success % ↑
Epos (m) ↓
Erot (rad) ↓
Dex-retargeting [ 22 ]
28.6
0.08
0.62
SPIDER [ 19 ] (mjwp)
71.4
0.04
0.57
SPIDER [ 19 ] (mjwp_act)
77.1
0.04
0.42
Do-as-I-Do [ 18 ] (Sharpa hand)
81.0
0.03
0.15
without transition reward
79.0
0.03
0.14
annealed sampling only
72.0
0.08
0.32
TABLE II: Human-to-robot retargeting quality on OakInk. Success is the share of clips whose converted trajectory passes the physics check. Epos and Erot are the residual position and rotation gap from the demonstration, lower is better. The two indented rows ablate the strongest baseline.
Fig. 3: Initial object-state distributions for policy evaluation. Object placements are randomized across trials for each task.
Data
Cup
Giftbox
Drawer
Rope Knot
Scissors
Drawing
Computer
Bottle
Biology
Battery
Multi-object
Total
Dex-retargeting [ 22 ]
✓
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
1/11
Do as I Do [ 18 ]
✓
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
1/11
Spider [ 19 ]
✓
✓
✓
✗
✓
✗
✓
✗
✗
✓
✗
6/11
TopoRetarget [ 30 ]
✓
✓
✓
✗
✗
✗
✗
✗
✗
✗
✗
3/11
Egoinfinity [ 27 ]
✓
✗
✗
✗
✗
✗
✗
✗
✗
✗
✗
1/11
V2D [ 14 ]
✓
✓
✓
✗
✓
✗
✗
✗
✗
✗
✓
5/11
TABLE III: Real-world replay success across eleven tasks. Each converted trajectory is executed open loop, which scores the conversion rather than a learned policy. A task is marked successful if at least one of ten replay trials succeeds. GPT-6 Astra joined the benchmark after the first evaluation round.
Fig. 4: Processing cost falls as the library grows. Over 100 EgoDex samples the library reaches 85 skills and 168 verifiers. In terms of cost, V2D converting one sample takes 3.7 hours on average and GPT 6: Astra takes 3.3 hours on average at a flat rate, whereas DexAgent takes 2.1 hours on average in a fresh run, a 36.4% reduction comparing to GPT 6: Astra and a 43.2% reduction comparing to V2D.
Fig. 5: Real experiments results on four selected tasks. Each row shows the original human demonstration ( Column 1 ), its reconstructed simulation scene( Column 2 ), and representative frames from the resulting real-robot rollout ( Columns 3–6 ). The examples span diverse manipulation behaviors, including deformable-object manipulation, multi-object interaction, precise insertion, and articulated object manipulation.
Fig. 7: Additional examples of property-based simulation reconstruction in Stage 2 and robot trajectory optimization in simulation in Stage 3. Column 1 shows the recorded human demonstration, and Column 2 shows the corresponding simulation scene reconstructed in Stage 2. Column 3-5 show representative frames t0 – t3 from the robot trajectory optimized in simulation from Stage 3.
Human manipulation videos are a convenient and intuitive source for robot learning. However, directly transferring human dexterity to robots remains challenging due to perception errors and embodiment gap. To address this, we introduce Video2Sim2Real, a full-stack framework for autonomous skill acquisition from a single human manipulation video. Our framework first uses off-the-shelf foundation models to reconstruct a simulator-ready digital twin and extract robot and object motion priors. Rather than treating the extracted robot motion as a reliable reference throughout execution, our key idea is to recover and leverage the most fundamental sources of supervision from the demonstrated skill: We identify object-centric keyframes to optimize the corresponding robot configurations using object information from the simulator, and use these configurations as anchors that refine the robot motion such that it ultimately has the desired impact on the environment. To bridge the remaining sim-to-real gap, we introduce a sim-to-real strategy that decouples robustness to noisy and incomplete perception from variations in hand-object interaction dynamics. Specifically, we learn to recalibrate robot configurations from noisy real-world point clouds via IL, and leverage residual RL to perform local finger-level adaptations to ensure for robust and effective interactions. Finally, a collision-aware motion planning module enables spatial generalization to novel object configurations. Across several everyday manipulation tasks, Video2Sim2Real improves simulated task success, safety, and trajectory coherence over numerous baselines, and achieves better sim-to-real transfer than existing techniques. These results demonstrate a promising path toward autonomous dexterous skill acquisition from human videos.
Yunhai Han, Jianuo Qiu, Linhao Bai +14
1Georgia Institute of Technology · University of Pennsylvania · 3Toyota Research Institute +1
Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.
Ruoqu Chen, Feixiang Ruan, Liu Cao +9
Tsinghua University · Shanghai Qizhi Institute · Sharpa +2
Sim-to-real transfer remains a critical bottleneck for deploying dexterous manipulation policies learned in simulation to real-world robots. Existing approaches rely on manually designed domain randomization or task-specific adaptation, limiting their generalizability across diverse manipulation scenarios. We present DexSim2Real, an integrated framework that leverages vision-language foundation models to bridge the sim-to-real gap for dexterous manipulation. Our system combines three components: (1) Foundation Model-Guided Domain Randomization (FM-DR), which uses a vision-language model as a visual realism critic to optimize simulation parameters via closed-loop CMA-ES, complementing text-based approaches like DrEureka with direct visual feedback; (2) a Tactile-Visual Cross-Attention Policy (TVCAP) that adapts cross-attention visuo-tactile fusion to zero-shot sim-to-real RL; and (3) a Progressive Skill Curriculum (PSC) that builds on LLM-based task decomposition with a difficulty scheduler tailored to contact-rich dexterous tasks. Extensive experiments on six challenging manipulation tasks with blinded evaluation demonstrate that DexSim2Real achieves a 78.2% average real-world success rate, outperforming DrEureka and DeXtreme while reducing the sim-to-real performance gap to only 8.3%.
Zijian Zeng, Fei Ding, Huiming Yang +2
Tsinghua University · Alibaba Group · Bengbu University +1