cs.ROSep 30, 2026

IronMind: Scaling Humanoid Dexterous Manipulation via Camera-Space Ego-Centric Pretraining

Authors: Huimin Pan, Yufan Ren, Kunpeng Song, Siyang Wang, Xiwen Zhang, Xiaoyun Hu, Zhuoxu Duan, Hanrui Zheng, +11 more

Organizations: World Model Team, XPENG Robotics

Abstract

Egocentric human video offers a scalable data source for dexterous manipulation, yet using it to train humanoid robots presents two challenges: (1) an embodiment gap, as human hands differ structurally from robot end-effectors and low-cost egocentric recordings lack the torso kinematics required by conventional retargeting; and (2) heterogeneous data quality, including noisy hand-pose tracking and weakly aligned text annotations. We introduce IronMind, a vision-language-action (VLA) model that uses egocentric human video and heterogeneous robot data to pretrain policies for humanoid dexterous manipulation. To bridge the embodiment gap, IronMind bypasses explicit body-retargeting by using a camera-space action representation, the native reference space of egocentric video, and semantically aligning robot and human action dimensions. Across total pretraining budgets from 250 to 10,000 hours, validation loss decreases approximately log-linearly with data scale. Larger pretraining budgets also improve out-of-distribution real-robot manipulation after post-training: across six challenging tasks with unseen objects, affordances, and reasoning prompts, the 10,000-hour model achieves a 55.0% success rate, compared with at most 11.7% for every pretraining budget up to 5,000 hours and 5.0% without pretraining. At the same pretraining budget, the camera-space action representation also outperforms the torso-frame baseline. Together, these findings support pretraining with a camera-space action representation on large-scale human egocentric data as a scalable foundation for humanoid robot manipulation.

Figures & tables

Explore similar work

CardsList
  1. EgoEngine: From Egocentric Human Videos to High-Fidelity Dexterous Robot Demonstrations

    Jun 10, 2026Yangcen Liu, Shuo Cheng, Xinchen Yin +6Egocentric VideoRobotic Data Acquisition

  2. ACE-Ego-0: Unifying Egocentric Human and Robotic Data for VLA Pretraining

    Jun 15, 2026Hao Li, Ganlong Zhao, Yufei Liu +8Robotic Data AcquisitionEgocentric Video

  3. Wh0: Generative World Models as Scalable Sources of Egocentric Human Hand Manipulation Data

    Jun 20, 2026Yangtao Chen, Zixuan Chen, Peiyang Wang +4Contact-Rich Dexterous ManipulationEmbodiment