World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains. Finally, our conclusions transfer to downstream humanoid modeling. Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.
Figures & tables
Figure 1: Limitations of world model scaling. Left: Ego-centric data scaling can achieve high agent modeling fidelity, but the effect of agent’s actions on the world lags behind. Right: The same failure at the level of a single prediction. Our best model places the hands almost exactly where they belong, while the paper they are folding is rendered in the wrong configuration throughout. See the project website for more video-level results.
Figure 2: Cosmos 3 with skeleton conditioning. The projected skeleton sequence S0:T , shown in bottom left, is encoded by the same frozen VAE E as the video, and its latent patches are projected through a zero-initialized Ws and added to the corresponding video embeddings. The skeleton is derived from the camera intrinsics to better align agent’s actions with the visual token grid.
Figure 3: Dataset overview. Top: distribution of clips over environments, scenes, and tasks (ten most frequent categories each; the gray wedge aggregates the long tail). Bottom: representative first-person frames from four environments, overlaid with the dataset’s per-frame keypoints.
Figure 4: Scaling reaches the agent’s limit, not the object’s. SCS for the agent (left) and manipulated objects (right) versus training data, with and without skeleton conditioning. Skeleton conditioning reaches the agent’s limit with far less data, while object fidelity converges far below.
Figure 5
Variant
Contribution
Metric
Cosmos 3
+Human
Ours
Human pre-train
Ours
Hand SCS
0.63
0.70
0.88
+0.07
+0.17
Object SCS
0.70
0.73
0.79
+0.03
+0.06
Success agreement
0.67
0.71
0.81
+0.04
+0.10
Table 2: Transfer to humanoid manipulation. SCS on the humanoid benchmark, and agreement with the simulator’s rendering on policy-rollout success. Contribution columns give the difference between adjacent variants. Higher is better.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Value
Clips (recordings)
1,146,100
Total duration
30,012 hours
Total frames ( 30 fps)
3,241,303,112
Mean / median clip length
94.3 s / 94.5 s
Clip length range
2.6 – 240.0 s
Unique environments
116
Appendix
Table 3: Training set summary statistics (full corpus).
Figure 6: Left: clip-length distribution (mean 94.3 s, median 94.5 s; 99.9% of clips are ≤123 s). The axis is truncated at 135 s; a sparse tail ( 0.04% of clips) extends to a 240 s cap and is omitted. Right: histogram of clips per participant (log y -axis), showing a heavy tail across 14,019 contributors.
Figure 7: Most frequent environments (left, top 30 of 116 ) and scenes (right, top 30 of 1,106 , near-duplicates merged), as a percentage of all clips.
Figure 8: Task composition. Each task label decomposes into an action (inner ring) applied to a manipulated object (outer ring); together the corpus spans 874 distinct actions and 9,397 distinct objects ( 32,051 action–object tasks). For legibility the inner ring shows the 12 most frequent actions with wedge angles proportional to their share of all clips; the remaining 862 actions ( 41% of clips) are grouped into the grey “other” wedge. The outer ring shows the top objects within each action, with that action’s remaining objects aggregated into a lighter “other” sub-wedge. While a few actions (packaging, cleaning, making) are common, each fans out over a broad range of objects.
Training budget (hours)
Optimizer steps
300
655
1,000
2,182
3,000
6,545
10,000
21,818
30,000
65,454
Appendix
Table 4: Checkpoints in the balanced data ladder. Each row is a cumulative budget along one training run per model configuration.
Figure 9: Improvement and remaining failures in human manipulation. (a) Our model correctly captures the right hand’s motion in the video, whereas vanilla Cosmos 3 fails. (b) The spoon is not lifted. (c) The rigid board stretches. (d) The upper cucumber is missing after the hand moves away. (e) The paper does not open. Frames and crops are matched across methods; f0 is the observed frame. Outlines mark the compared regions. Best viewed on the project website .
Figure 10: One successful humanoid rollout and three object interaction failures. Each row shows one example, with simulator’s frames on the left and the corresponding predictions of our world model on the right. (a) Both rollouts unload the can within the shown window. (b) In the simulator the box turns, while in the world model it stays in place. (c) In the simulator the laptop opens, while the world model leaves it closed. (d) In the simulator the can drops in the bin, while in the world model’s predictions it remains stuck near the bin’s edge. Frame indices are aligned across each pair, with a common crop throughout. Success labels indicate whether the world model’s prediction matches the simulator’s object response. Best viewed on the project website .
Training condition
Agent SCS ↑
Object SCS ↑
Skeleton conditioning
0.779
0.513
+ 25% skeleton dropout
0.778
0.509
Appendix
Table 5: Skeleton dropout ablation. Cosmos 3 Nano at 10,000 training hours. Dropout is applied only during training; both variants use full skeleton conditioning at evaluation. The results indicate that skeleton conditioning does not introduce shortcuts for learning object interaction dynamics.
Figure 11: Perceptual quality across the data-scaling ladder. LPIPS (lower is better) versus training exposure. Further scaling of ego-centric data is unlikely to improve the perceptual quality of the model’s predictions either.
Figure 12: Dynamic-region supervision. (a) D4RT tracks queried points in 3D to separate scene motion from camera motion. We combine 3D motion, tracking reliability, and a soft spatial prior around the hands to obtain the dynamic region map Mdyn . (b) Dynamic-region supervision combines dynamic noise scheduling and loss reweighting in selected regions.
metric
variant
2,000
3,000
9,000
hand SCS
Ours
0.88
0.88
0.88
+Human
0.70
0.70
0.69
Cosmos 3
0.44
0.56
0.64
object SCS
Ours
0.77
0.77
0.76
+Human
0.70
0.71
0.72
Cosmos 3
0.56
0.67
0.67
Appendix
Table 6: SCS at fixed iterations, averaged over all 11 tasks. The ego-centrically pre-trained variants are converged throughout this range; Cosmos 3 is not, which is why Table 2 selects per variant rather than fixing a shared iteration.
hand SCS
object SCS
task
Cosmos 3
+Human
Ours
Cosmos 3
+Human
Ours
Close-Drawer
0.60
0.65
0.88
0.92
0.86
0.97
Flip-Mug
0.69
0.75
0.90
0.70
0.77
0.79
Insert-Cans
0.59
0.64
0.86
0.60
0.61
0.62
Open-Drawer
0.53
0.65
0.86
0.58
0.49
0.62
Pour-Balls
0.71
0.77
0.89
0.75
0.78
0.81
Appendix
Table 7: Per-task SCS on the humanoid benchmark, at the checkpoints reported in Table 2 .
variant
success
agreement
fail recall
success recall
simulator rendering
0.86
—
—
—
Cosmos 3
0.67
0.67
0.50
0.69
+Human
0.67
0.71
0.67
0.72
Ours
0.76
0.81
0.67
0.83
Appendix
Table 8: Agreement with the simulator’s rendering on the seven eligible tasks. Agreement is the fraction of clips on which the two verdicts coincide; recalls are conditioned on the verdict for the simulator’s rendering.
Figure 13: Rectifying the evaluation set onto the training camera model. Left: the raw capture from the stereo rig’s left camera, a wide-angle double-sphere fisheye ( 1920×1200 ); the green box marks the region that maps into the rectified image. Middle: that region undistorted to a 908×512 pinhole. Right: a training frame, shown for comparison.
Embodied foundation models are expected to benefit from data scaling like large language models, but face a much tighter data bottleneck. Teleoperated real-robot trajectories remain the dominant pretraining source due to their precise action supervision and embodiment alignment, yet their scalability is limited by high collection cost, acquisition difficulty, and low behavioral and environmental diversity. These limitations have sparked interest in egocentric human video as a scalable, substantially lower-cost, and more diverse alternative for embodied model pretraining. However, its effectiveness compared to teleoperated real-robot data remains underexplored. To address this question, we conduct a systematic study comparing egocentric human video and teleoperated real-robot trajectories as pretraining data sources for embodied foundation models, under fixed post-training and validation protocols. Surprisingly, we find that egocentric data, when processed through a carefully designed filtering and labeling pipeline, is not merely a viable substitute for model pretraining but can lead to superior performance. With the same amount of pretraining data, models pretrained on egocentric data achieve a 24% lower validation loss on real-robot action prediction, as well as 52.5% and 90% higher success rates on in-distribution and out-of-distribution real-robot task execution, respectively. This finding verifies a scalable paradigm for embodied foundation models: pretrain on egocentric human video to learn diverse world representations, then adapt with a small amount of labeled real-robot data for action-space alignment. We hope this study encourages broader exploration of egocentric data and offers guidance for data quality assessment before costly robot data collection.
Building interactive simulators from real-world observations is a promising way to scale embodied data, but current pipelines still rely heavily on manual environment construction and calibration. We study whether frontier foundation models and coding agents can automate this process end to end. We formulate \emph{autonomous video-to-simulation} as a software engineering task in which an agent observes an embodied video, constructs the corresponding simulated environment and robot behavior, and iteratively refines the result through execution feedback. To evaluate this capability, we introduce \textbf{Video2World}, a benchmark comprising 222 reconstruction instances derived from 189 robot and human demonstration videos. Video2World measures reconstructed worlds along geometric fidelity, dynamic fidelity, and functional correctness, capturing spatial perception, physical reasoning, and executable interaction. Evaluating 9 frontier coding-agent systems reveals a sharp improvement in Task success beginning with Claude Opus 5, rising from below 5% to over 15%, while substantial gaps to human-assisted reconstruction remain. We further find that worlds that look better could work worse: better visual fidelity does not always lead to higher task success. This echoes the broader gap between perceptual realism and factual correctness observed in generative models.
Progress in embodied intelligence increasingly depends on scalable data infrastructure. While vision and language have scaled with internet corpora, learning physical interaction remains constrained by the lack of large, diverse, and richly annotated human activity data. We present HumanNet, a one-million-hour human-centric video corpus that captures how humans interact with the physical world at scale. HumanNet spans both first-person and third-person perspectives and covers fine-grained activities, human-object interactions, tool use, and long-horizon behaviors across diverse real-world environments. Beyond raw video, the dataset provides interaction-centric annotations, including captions, motion descriptions, and hand and body-related signals, enabling motion-aware and interaction-aware learning. Beyond scale, HumanNet introduces a systematic data curation paradigm for embodied learning, where human-centric filtering, temporal structuring, viewpoint diversity, and annotation enrichment are treated as first-class design principles. This design transforms unstructured internet video into a scalable substrate for representation learning, activity understanding, motion generation, and human-to-robot transfer. We conduct a first-step validation on the value of this design through controlled vision-language-action ablation: under a fixed set of validation data, continued training from the Qwen VLM model with 1000 hours of egocentric video drawn from HumanNet surpasses the continued training with 100 hours of real-robot data from Magic Cobot, indicating that egocentric human video could be a scalable and cost-effective substitute for robot data. By building this project, we aim to explore the opportunity to scale embodied foundation models using human-centric videos, rather than relying solely on robot-specific data.
Yufan Deng, Daquan Zhou
DAGroup · SimpleSilicon Innovation Team Peking University