Egocentric human video offers a scalable data source for dexterous manipulation, yet using it to train humanoid robots presents two challenges: (1) an embodiment gap, as human hands differ structurally from robot end-effectors and low-cost egocentric recordings lack the torso kinematics required by conventional retargeting; and (2) heterogeneous data quality, including noisy hand-pose tracking and weakly aligned text annotations. We introduce IronMind, a vision-language-action (VLA) model that uses egocentric human video and heterogeneous robot data to pretrain policies for humanoid dexterous manipulation. To bridge the embodiment gap, IronMind bypasses explicit body-retargeting by using a camera-space action representation, the native reference space of egocentric video, and semantically aligning robot and human action dimensions. Across total pretraining budgets from 250 to 10,000 hours, validation loss decreases approximately log-linearly with data scale. Larger pretraining budgets also improve out-of-distribution real-robot manipulation after post-training: across six challenging tasks with unseen objects, affordances, and reasoning prompts, the 10,000-hour model achieves a 55.0% success rate, compared with at most 11.7% for every pretraining budget up to 5,000 hours and 5.0% without pretraining. At the same pretraining budget, the camera-space action representation also outperforms the torso-frame baseline. Together, these findings support pretraining with a camera-space action representation on large-scale human egocentric data as a scalable foundation for humanoid robot manipulation.
Figures & tables
Figure 1: Overview of IronMind. (1) Pretraining Corpus: Over 10,000 hours in total, combining egocentric human video with heterogeneous robot data from non-target embodiments. (2) Curation & Representation: Data filtered, re-annotated, and unified into a shared camera-space action representation. (3) Architecture: A Mixture-of-Transformers policy with flow-matching action experts and optional multimodal priors. (4) Scaling: Validation loss scales log-linearly with pretraining hours without saturation up to 10,000 hours. (5) Zero-shot Evaluation: Open-loop trajectory prediction from a single unseen image and instruction. (6) Real-Robot Deployment: Closed-loop policy execution on IRON-R01.
Figure 2: Scene distribution of the IronMind pretraining corpus. The dataset comprises over 10,000 total hours of egocentric human video and heterogeneous robot manipulation data spanning diverse real-world environments, all expressed using a camera-space action representation. The donut chart shows the proportion of recording hours in each scene category, alongside representative frames. Home environments account for the largest share; the remaining data cover food and beverage settings, laboratories and workshops, hospitality venues, retail spaces, educational facilities, offices, and other environments.
Figure 3: Overview of the data curation pipeline with representative examples. Rule-based filtering rejects recordings that fail the quality criteria (two examples shown). Atomic Re-annotation: Retained recordings are segmented into clips, each corresponding to a single task intent, and assigned refined captions. Per-frame Quality Weighting: Frames receive quality scores according to Eq. 1 (plotted for a sample recording) to determine their sampling weights during training.
Figure 4: IronMind architecture. A vision-language understanding expert takes an egocentric image, a language instruction, and optionally a set of query tokens as input. During training, the query tokens produce world prior features (geometry, semantics, future frames) and are supervised using targets from pretrained models in each domain; the prior branches and prediction heads are removed at inference. The action expert attends to the understanding expert at shared attention layers and generates actions in the camera-space representation through flow-matching denoising, conditioned on robot state and action noise.
Figure 5: Evaluation protocol. Left: Quantitative open-loop evaluation. With the camera-space action representation, target approach is measured via dp10 and grasp execution via finger aperture error eFA . Middle: Evaluation loop for pretraining. Held-out RealSense tabletop scenes with language prompts are used both to iterate pretraining and to score policy-in-the-loop rollouts. Right: Policy-in-the-loop rollout. Chunk t0 is predicted from the observation; subsequent chunks t1,t2 re-render the predicted hand while the camera and physical scene stay fixed.
Figure 6: Open-loop capability evaluation on OOD scenes. Given an unseen RGB image and a text instruction, IronMind predicts open-loop hand trajectories (rendered hand meshes colored chronologically from light skin tone to purple). Visualizations illustrate fine-grained spatial grounding, including adjustments of grasp aperture and wrist orientation to target geometry.
Figure 7: Pretraining data scaling. Pretraining validation loss exhibits a smooth log-linear scaling trend across total pretraining budgets (250 h to 10,000 h) with a slope of −0.058 per decade of hours, with larger budgets also improving open-loop target approach and grasp closure.
Table 8
Pretraining Budget
Votes
Win Rate ↑
250 h
23
32.9%
10,000 h
47
67.1%
Table 3: Human preference win rate vs. pretraining data budget. The blind user study in Section 4.2.1 covers 20 held-out trials: 100 votes: 30 ties and 70 decisions; win rate is computed over decided votes.
Figure 8: What the world prior branches predict from one observation. The left panel shows observation t ; for each branch the prediction (top) sits above the ground truth (bottom) in the same rendering. The panels show future-frame dynamics, 3D motion flow, semantic features, and geometry targets over a 0.4 s horizon.
Prior Configuration
dp10 (cm) ↓
Prog. (%) ↑
θalign (°) ↓
Reference configurations
no world-prior supervision (baseline)
27.6
40.9
34.8
three priors (default: future target, 0.4 s)
22.6
53.4
26.5
Prior composition
without the semantics prior
24.1
50.2
29.2
without the depth prior
25.7
46.1
31.1
Table 4: World-prior ablation on the open-loop approaching set. Results are averaged over several checkpoints from each pretraining run. Lower is better for dp10 and θalign ; higher is better for progress.
Figure 9: Real-robot out-of-distribution (OOD) evaluation on IRON-R01. Success and progress rates are evaluated across three out-of-distribution categories (Semantic, Affordance, and Reasoning OOD).
Figure 10: Scaling effect on real-robot OOD success and progress rates. Success (blue) and milestone progress (gray) are reported for six OOD tasks grouped into Semantic, Affordance, and Reasoning categories. At smaller budgets (0–2,000 h), performance is generally low and non-monotonic across tasks. The clearest improvement occurs between 5,000 and 10,000 h, when success increases on all six tasks—for example, from 10% to 90% for grapes to the left/right basket and from 0% to 50% for the blue/yellow cup—showing an overall positive scaling trend despite fluctuations at smaller scales.
Model variant
Success ↑
Progress ↑
IronMind (ours, camera-space, no world-prior supervision)
55.0%
71.3%
IronMind (ours, torso-frame baseline)
26.7% ( ↓ 28.3)
46.2% ( ↓ 25.1)
IronMind (ours, with world-prior supervision)
31.7% ( ↓ 23.3)
53.3% ( ↓ 18.0)
InternVLA-A1.5 [ Ma et al., 2026a ]
30.0% ( ↓ 25.0)
47.8% ( ↓ 23.5)
Table 5: Real-robot ablations. Results are averaged over the six OOD tasks of Figure 9 (10 trials each), using the same post-training data and compute budget. Parentheses report percentage-point drops relative to the first row.
Dexterous manipulation is limited by the cost of collecting large-scale robot demonstrations. Egocentric human videos offer a scalable source of diverse manipulation behaviors, but directly using them for robot learning requires bridging two gaps: the visual gap between human and robot observations, and the action gap between human motion and robot-executable action. We propose EgoEngine, a scalable framework for transforming egocentric human manipulation videos into high-fidelity robot data. Given an egocentric RGB video, EgoEngine produces: (i) a high-fidelity robot observation video replacing human with robot while preserving scene context and temporal alignment, and (ii) a task-aligned, executable robot action trajectory under feasibility constraints. Experiments in simulation and on real robots show that EgoEngine enables scalable conversion of human videos into robot data and, to our knowledge, demonstrates the first zero-shot visuomotor dexterous policy learning from egocentric human videos without real-robot demonstrations. Project website: https://egoengine.github.io.
Yangcen Liu, Shuo Cheng, Xinchen Yin +6
Georgia Institute of Technology · Tsinghua University
Vision-Language-Action (VLA) models benefit from large-scale and diverse embodied data, yet scaling robot trajectory collection is costly and labor-intensive. Recent advances show that large-scale egocentric human videos provide complementary real-world supervision in pretraining. However, joint training on human and robot data remains challenging due to divergences in action spaces, embodiment structures, temporal dynamics, and supervision quality. We introduce ACE-EGO-0, a unified VLA pretraining framework jointly leveraging heterogeneous data sources. To extract large-scale pretraining supervision from egocentric human videos, we build a scalable egocentric video-to-action pipeline that converts raw human videos into robot-format pseudo-action trajectories. To make these labels comparable with robot demonstrations, ACE-EGO-0 uses a unified action representation based on camera-space actions, morphology conditioning, and time-aligned action chunking. To robustly leverage noisy pseudo-action supervision from egocentric human videos, we formulate a reliability-aware training objective with a human auxiliary loss that concentrates supervision on reliable signals. We instantiate ACE-EGO-0 on 4.53K hours of robot and simulation data, together with 1.48K hours of pseudo-action-labeled egocentric human data. Experiments show that incorporating large-scale human supervision under reliability-aware weighting consistently improves both unified joint pretraining and supervised fine-tuning. ACE-EGO-0 achieves state-of-the-art performance on RoboCasa GR1 TableTop and RoboTwin 2.0, while demonstrating strong transfer to real-world bimanual manipulation.
Scaling dexterous manipulation requires generalization across objects, scenes, and tasks, yet existing data sources face a trade-off between scale and scene/embodiment alignment: teleoperation data is well aligned with robot deployment but expensive to collect; simulation is scalable but limited by the sim-to-real gap; and real egocentric videos scale effectively but remain misaligned with robot deployment. We propose Wh0, a framework that uses generative video world models as scalable and controllable sources of egocentric human-hand manipulation data to unlock the manipulation capabilities of pretrained dexterous VLA models. Conditioned on language, objects, and scenes, Wh0 uses a generative world model to produce WM-H, a 50k-episode dataset of egocentric human-object interaction videos. Wh0 then converts the generated videos into robot-trainable supervision through hand motion reconstruction and visual editing. Co-trained with a limited amount of real robot data, WM-H adapts pretrained VLA models to dexterous manipulation deployment. Across 18 real-world dexterous manipulation tasks, compared with a model post-trained only on robot data, Wh0 improves zero-shot success on unseen tasks from 8.3% to 38.9%. Ablation studies further show that scalable generation and scene/embodiment alignment are key drivers of performance gains. Videos and open-source code can be found on our project website: https://chenyt31.github.io/wh0.github.io/.
Yangtao Chen, Zixuan Chen, Peiyang Wang +4
Shanghai Innovation Institute · Nanjing University · Shanghai Jiaotong University