World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then introduce a supervision scheme that shifts capacity from scene appearance toward object dynamics, improving object fidelity though a substantial gap remains. Finally, our conclusions transfer to downstream humanoid modeling. Overall, our results suggest that scaling ego-centric data brings agent modeling close to its limit while leaving its effects on the world far behind, and that closing this gap will depend on how models are trained, not only on how much data they see.
Figures & tables
Figure 1: Limitations of world model scaling. Left: Ego-centric data scaling can achieve high agent modeling fidelity, but the effect of agent’s actions on the world lags behind. Right: The same failure at the level of a single prediction. Our best model places the hands almost exactly where they belong, while the paper they are folding is rendered in the wrong configuration throughout. See the project website for more video-level results.
Figure 2: Cosmos 3 with skeleton conditioning. The projected skeleton sequence S0:T , shown in bottom left, is encoded by the same frozen VAE E as the video, and its latent patches are projected through a zero-initialized Ws and added to the corresponding video embeddings. The skeleton is derived from the camera intrinsics to better align agent’s actions with the visual token grid.
Figure 3: Dataset overview. Top: distribution of clips over environments, scenes, and tasks (ten most frequent categories each; the gray wedge aggregates the long tail). Bottom: representative first-person frames from four environments, overlaid with the dataset’s per-frame keypoints.
Figure 4: Scaling reaches the agent’s limit, not the object’s. SCS for the agent (left) and manipulated objects (right) versus training data, with and without skeleton conditioning. Skeleton conditioning reaches the agent’s limit with far less data, while object fidelity converges far below.
Figure 5
Variant
Contribution
Metric
Cosmos 3
+Human
Ours
Human pre-train
Ours
Hand SCS
0.63
0.70
0.88
+0.07
+0.17
Object SCS
0.70
0.73
0.79
+0.03
+0.06
Success agreement
0.67
0.71
0.81
+0.04
+0.10
Table 2: Transfer to humanoid manipulation. SCS on the humanoid benchmark, and agreement with the simulator’s rendering on policy-rollout success. Contribution columns give the difference between adjacent variants. Higher is better.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Statistic
Value
Clips (recordings)
1,146,100
Total duration
30,012 hours
Total frames ( 30 fps)
3,241,303,112
Mean / median clip length
94.3 s / 94.5 s
Clip length range
2.6 – 240.0 s
Unique environments
116
Appendix
Table 3: Training set summary statistics (full corpus).
Figure 6: Left: clip-length distribution (mean 94.3 s, median 94.5 s; 99.9% of clips are ≤123 s). The axis is truncated at 135 s; a sparse tail ( 0.04% of clips) extends to a 240 s cap and is omitted. Right: histogram of clips per participant (log y -axis), showing a heavy tail across 14,019 contributors.
Figure 7: Most frequent environments (left, top 30 of 116 ) and scenes (right, top 30 of 1,106 , near-duplicates merged), as a percentage of all clips.
Figure 8: Task composition. Each task label decomposes into an action (inner ring) applied to a manipulated object (outer ring); together the corpus spans 874 distinct actions and 9,397 distinct objects ( 32,051 action–object tasks). For legibility the inner ring shows the 12 most frequent actions with wedge angles proportional to their share of all clips; the remaining 862 actions ( 41% of clips) are grouped into the grey “other” wedge. The outer ring shows the top objects within each action, with that action’s remaining objects aggregated into a lighter “other” sub-wedge. While a few actions (packaging, cleaning, making) are common, each fans out over a broad range of objects.
Training budget (hours)
Optimizer steps
300
655
1,000
2,182
3,000
6,545
10,000
21,818
30,000
65,454
Appendix
Table 4: Checkpoints in the balanced data ladder. Each row is a cumulative budget along one training run per model configuration.
Figure 9: Improvement and remaining failures in human manipulation. (a) Our model correctly captures the right hand’s motion in the video, whereas vanilla Cosmos 3 fails. (b) The spoon is not lifted. (c) The rigid board stretches. (d) The upper cucumber is missing after the hand moves away. (e) The paper does not open. Frames and crops are matched across methods; f0 is the observed frame. Outlines mark the compared regions. Best viewed on the project website .
Figure 10: One successful humanoid rollout and three object interaction failures. Each row shows one example, with simulator’s frames on the left and the corresponding predictions of our world model on the right. (a) Both rollouts unload the can within the shown window. (b) In the simulator the box turns, while in the world model it stays in place. (c) In the simulator the laptop opens, while the world model leaves it closed. (d) In the simulator the can drops in the bin, while in the world model’s predictions it remains stuck near the bin’s edge. Frame indices are aligned across each pair, with a common crop throughout. Success labels indicate whether the world model’s prediction matches the simulator’s object response. Best viewed on the project website .
Training condition
Agent SCS ↑
Object SCS ↑
Skeleton conditioning
0.779
0.513
+ 25% skeleton dropout
0.778
0.509
Appendix
Table 5: Skeleton dropout ablation. Cosmos 3 Nano at 10,000 training hours. Dropout is applied only during training; both variants use full skeleton conditioning at evaluation. The results indicate that skeleton conditioning does not introduce shortcuts for learning object interaction dynamics.
Figure 11: Perceptual quality across the data-scaling ladder. LPIPS (lower is better) versus training exposure. Further scaling of ego-centric data is unlikely to improve the perceptual quality of the model’s predictions either.
Figure 12: Dynamic-region supervision. (a) D4RT tracks queried points in 3D to separate scene motion from camera motion. We combine 3D motion, tracking reliability, and a soft spatial prior around the hands to obtain the dynamic region map Mdyn . (b) Dynamic-region supervision combines dynamic noise scheduling and loss reweighting in selected regions.
metric
variant
2,000
3,000
9,000
hand SCS
Ours
0.88
0.88
0.88
+Human
0.70
0.70
0.69
Cosmos 3
0.44
0.56
0.64
object SCS
Ours
0.77
0.77
0.76
+Human
0.70
0.71
0.72
Cosmos 3
0.56
0.67
0.67
Appendix
Table 6: SCS at fixed iterations, averaged over all 11 tasks. The ego-centrically pre-trained variants are converged throughout this range; Cosmos 3 is not, which is why Table 2 selects per variant rather than fixing a shared iteration.
hand SCS
object SCS
task
Cosmos 3
+Human
Ours
Cosmos 3
+Human
Ours
Close-Drawer
0.60
0.65
0.88
0.92
0.86
0.97
Flip-Mug
0.69
0.75
0.90
0.70
0.77
0.79
Insert-Cans
0.59
0.64
0.86
0.60
0.61
0.62
Open-Drawer
0.53
0.65
0.86
0.58
0.49
0.62
Pour-Balls
0.71
0.77
0.89
0.75
0.78
0.81
Appendix
Table 7: Per-task SCS on the humanoid benchmark, at the checkpoints reported in Table 2 .
variant
success
agreement
fail recall
success recall
simulator rendering
0.86
—
—
—
Cosmos 3
0.67
0.67
0.50
0.69
+Human
0.67
0.71
0.67
0.72
Ours
0.76
0.81
0.67
0.83
Appendix
Table 8: Agreement with the simulator’s rendering on the seven eligible tasks. Agreement is the fraction of clips on which the two verdicts coincide; recalls are conditioned on the verdict for the simulator’s rendering.
Figure 13: Rectifying the evaluation set onto the training camera model. Left: the raw capture from the stereo rig’s left camera, a wide-angle double-sphere fisheye ( 1920×1200 ); the green box marks the region that maps into the rectified image. Middle: that region undistorted to a 908×512 pinhole. Right: a training frame, shown for comparison.