Egocentric human video offers a scalable data source for dexterous manipulation, yet using it to train humanoid robots presents two challenges: (1) an embodiment gap, as human hands differ structurally from robot end-effectors and low-cost egocentric recordings lack the torso kinematics required by conventional retargeting; and (2) heterogeneous data quality, including noisy hand-pose tracking and weakly aligned text annotations. We introduce IronMind, a vision-language-action (VLA) model that uses egocentric human video and heterogeneous robot data to pretrain policies for humanoid dexterous manipulation. To bridge the embodiment gap, IronMind bypasses explicit body-retargeting by using a camera-space action representation, the native reference space of egocentric video, and semantically aligning robot and human action dimensions. Across total pretraining budgets from 250 to 10,000 hours, validation loss decreases approximately log-linearly with data scale. Larger pretraining budgets also improve out-of-distribution real-robot manipulation after post-training: across six challenging tasks with unseen objects, affordances, and reasoning prompts, the 10,000-hour model achieves a 55.0% success rate, compared with at most 11.7% for every pretraining budget up to 5,000 hours and 5.0% without pretraining. At the same pretraining budget, the camera-space action representation also outperforms the torso-frame baseline. Together, these findings support pretraining with a camera-space action representation on large-scale human egocentric data as a scalable foundation for humanoid robot manipulation.
Figures & tables
Figure 1: Overview of IronMind. (1) Pretraining Corpus: Over 10,000 hours in total, combining egocentric human video with heterogeneous robot data from non-target embodiments. (2) Curation & Representation: Data filtered, re-annotated, and unified into a shared camera-space action representation. (3) Architecture: A Mixture-of-Transformers policy with flow-matching action experts and optional multimodal priors. (4) Scaling: Validation loss scales log-linearly with pretraining hours without saturation up to 10,000 hours. (5) Zero-shot Evaluation: Open-loop trajectory prediction from a single unseen image and instruction. (6) Real-Robot Deployment: Closed-loop policy execution on IRON-R01.
Figure 2: Scene distribution of the IronMind pretraining corpus. The dataset comprises over 10,000 total hours of egocentric human video and heterogeneous robot manipulation data spanning diverse real-world environments, all expressed using a camera-space action representation. The donut chart shows the proportion of recording hours in each scene category, alongside representative frames. Home environments account for the largest share; the remaining data cover food and beverage settings, laboratories and workshops, hospitality venues, retail spaces, educational facilities, offices, and other environments.
Figure 3: Overview of the data curation pipeline with representative examples. Rule-based filtering rejects recordings that fail the quality criteria (two examples shown). Atomic Re-annotation: Retained recordings are segmented into clips, each corresponding to a single task intent, and assigned refined captions. Per-frame Quality Weighting: Frames receive quality scores according to Eq. 1 (plotted for a sample recording) to determine their sampling weights during training.
Figure 4: IronMind architecture. A vision-language understanding expert takes an egocentric image, a language instruction, and optionally a set of query tokens as input. During training, the query tokens produce world prior features (geometry, semantics, future frames) and are supervised using targets from pretrained models in each domain; the prior branches and prediction heads are removed at inference. The action expert attends to the understanding expert at shared attention layers and generates actions in the camera-space representation through flow-matching denoising, conditioned on robot state and action noise.
Figure 5: Evaluation protocol. Left: Quantitative open-loop evaluation. With the camera-space action representation, target approach is measured via dp10 and grasp execution via finger aperture error eFA . Middle: Evaluation loop for pretraining. Held-out RealSense tabletop scenes with language prompts are used both to iterate pretraining and to score policy-in-the-loop rollouts. Right: Policy-in-the-loop rollout. Chunk t0 is predicted from the observation; subsequent chunks t1,t2 re-render the predicted hand while the camera and physical scene stay fixed.
Figure 6: Open-loop capability evaluation on OOD scenes. Given an unseen RGB image and a text instruction, IronMind predicts open-loop hand trajectories (rendered hand meshes colored chronologically from light skin tone to purple). Visualizations illustrate fine-grained spatial grounding, including adjustments of grasp aperture and wrist orientation to target geometry.
Figure 7: Pretraining data scaling. Pretraining validation loss exhibits a smooth log-linear scaling trend across total pretraining budgets (250 h to 10,000 h) with a slope of −0.058 per decade of hours, with larger budgets also improving open-loop target approach and grasp closure.
Table 8
Pretraining Budget
Votes
Win Rate ↑
250 h
23
32.9%
10,000 h
47
67.1%
Table 3: Human preference win rate vs. pretraining data budget. The blind user study in Section 4.2.1 covers 20 held-out trials: 100 votes: 30 ties and 70 decisions; win rate is computed over decided votes.
Figure 8: What the world prior branches predict from one observation. The left panel shows observation t ; for each branch the prediction (top) sits above the ground truth (bottom) in the same rendering. The panels show future-frame dynamics, 3D motion flow, semantic features, and geometry targets over a 0.4 s horizon.
Prior Configuration
dp10 (cm) ↓
Prog. (%) ↑
θalign (°) ↓
Reference configurations
no world-prior supervision (baseline)
27.6
40.9
34.8
three priors (default: future target, 0.4 s)
22.6
53.4
26.5
Prior composition
without the semantics prior
24.1
50.2
29.2
without the depth prior
25.7
46.1
31.1
Table 4: World-prior ablation on the open-loop approaching set. Results are averaged over several checkpoints from each pretraining run. Lower is better for dp10 and θalign ; higher is better for progress.
Figure 9: Real-robot out-of-distribution (OOD) evaluation on IRON-R01. Success and progress rates are evaluated across three out-of-distribution categories (Semantic, Affordance, and Reasoning OOD).
Figure 10: Scaling effect on real-robot OOD success and progress rates. Success (blue) and milestone progress (gray) are reported for six OOD tasks grouped into Semantic, Affordance, and Reasoning categories. At smaller budgets (0–2,000 h), performance is generally low and non-monotonic across tasks. The clearest improvement occurs between 5,000 and 10,000 h, when success increases on all six tasks—for example, from 10% to 90% for grapes to the left/right basket and from 0% to 50% for the blue/yellow cup—showing an overall positive scaling trend despite fluctuations at smaller scales.
Model variant
Success ↑
Progress ↑
IronMind (ours, camera-space, no world-prior supervision)
55.0%
71.3%
IronMind (ours, torso-frame baseline)
26.7% ( ↓ 28.3)
46.2% ( ↓ 25.1)
IronMind (ours, with world-prior supervision)
31.7% ( ↓ 23.3)
53.3% ( ↓ 18.0)
InternVLA-A1.5 [ Ma et al., 2026a ]
30.0% ( ↓ 25.0)
47.8% ( ↓ 23.5)
Table 5: Real-robot ablations. Results are averaged over the six OOD tasks of Figure 9 (10 trials each), using the same post-training data and compute budget. Parentheses report percentage-point drops relative to the first row.