Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at project page.
Figures & tables
Figure 1: Overview of WorldLine . (a) A scalable three-stage training pipeline: Stage I learns robot dynamics from large-scale action-free videos; Stage II grounds heterogeneous robot controls through action-conditioned post-training; and Stage III distills the model for causal few-step rollout. (b) WorldLine maps an initial observation and an image-space control sequence to future videos across diverse environments and robot embodiments, including zero-shot settings. (c) WorldLine support policy evaluation and planning by predicting action outcomes before real-world execution.
Scale
Diversity
Supervision
Method
Video-only
Action data
Embod.
Tasks
Action repr.
Failure
CTRL-World Guo et al. (2026)
N/A
450 h
1
86
Native Action
60 h
GE-Sim-V2 Qiu et al. (2026)
N/A
2,500 h
1
NR
Image-space map
NR
Masked Visual Action Alzayer et al. (2026)
N/A
450 h
1
86
Masked Action
60 h
Worldline(ours)
10,000 h
2,500 h
> 10
> 3,500
Image-space map
200 h
Table 1: Training-data comparison across action-conditioned robot video models in scale, diversity, and action supervision. Embod. denotes the number of robot embodiments, Failure denotes the duration of failed trajectories, and NR denotes unreported information.
Figure 2: Data construction and three-stage training of WorldLine. Left: action-free videos are filtered for dynamics pretraining; action-labeled trajectories undergo calibration and projection checks; and failure trajectories are identified and paired with geometrically aligned controls. Right: Stage I learns manipulation dynamics, Stage II grounds heterogeneous controls with multi-view and failure-enriched training, and Stage III distills the model for efficient causal rollout.
In-Domain
DROID Evaluation
Model
AgiBotWorld (Successful)
AgiBotWorld (Failure)
DROID
PSNR ↑
SSIM ↑
LPIPS ↓
Robot IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Robot IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Robot IoU ↑
Video Generation Models
Cosmos3-Nano Agarwal et al. (2026)
12.54
0.4752
0.5654
0.2804
10.35
0.4723
0.6000
0.1444
16.97
0.6516
0.3492
0.2049
Cosmos3-Super Agarwal et al. (2026)
12.57
0.4685
0.5674
0.2975
10.73
0.4786
0.6216
0.1755
17.07
0.6524
0.3457
0.1978
Action-Conditioned Models
Table 2: Action-conditioned video prediction on held-out AgiBotWorld success/failure trajectories and DROID, averaged over five runs. † denotes DROID-trained methods, making DROID in-domain. Arrows indicate metric direction; green and blue mark the best and second-best results.
Figure 3: Qualitative action-conditioned video generation. Given initial frame and actions, each model predicts a frame on held-out AgiBotWorld (top) and out-of-domain DROID (bottom). Causal WorldLine better matches ground-truth robot motion and gripper–object configuration.
Figure 4: WorldLine for policy evaluation and selection. (a) Mean policy-evaluation accuracy over five generation seeds on RoboTwin and AgiBot. (b) RoboTwin task success rate under different rollout budgets for π0.5 and LingBot-VLA.
Figure 5: Causal distillation efficiency and quality. Left: sampling steps and per-frame latency under identical settings. Right: Stage-II and causal predictions on AgiBot. The causal model accelerates rollout while preserving appearance, robot geometry, and object evolution.
Figure 6: Qualitative comparison of WorldLine ablations. We compare the full model with variants that remove key data, representation, and training components, highlighting their effects on robot motion and robot–object interaction outcomes.
AgiBotWorld
DROID
RoboTwin: π0.5
RoboTwin: LingBot-VLA
Variant
LPIPS ↓
Robot IoU ↑
LPIPS ↓
Robot IoU ↑
Success ↑
Gain ↑
Success ↑
Gain ↑
Full WorldLine
0.1804
0.5622
0.2968
0.2739
78.0±1.4%
+19.1%
62.4±3.6%
+21.4%
Data
No action-free data (compute-matched)
0.2924
0.5277
0.3028
0.2614
73.2±3.3%
+14.3%
60.4±3.0%
+19.4%
Head-view-only data
–
0.4348
–
0.2665
72.8±1.10 %
+13.9%
62.0 ±3.2%
+21.0%
Success-only trajectories
0.2401
0.4721
0.3028
0.2566
77.2±2.3%
+18.3%
51.6±3.0%
+10.6%
Table 3: Ablation of WorldLine components. LPIPS and robot-mask IoU are evaluated on held-out AgiBotWorld and out-of-domain DROID. RoboTwin planning uses 32 rollouts; gains are over direct execution. LPIPS is omitted for head-view-only. The no-action-free variant matches the full model’s training-compute budget. Green and blue mark best and second-best results.
Figure 7: Matched action-encoding comparison on RoboTwin.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Handover Block
Lift Pot
Open Laptop
Stack Two Blocks
Turn Switch
Average
π0.5
Direct Execution
7.5
83.1
85.6
37.5
52.5
53.3
GE-Sim-V2
32.0
96.0
84.0
56.0
48.0
63.2
OpenDW-0.5
32.0
84.0
80.0
48.0
52.0
59.2
Masked Visual Actions
0.0
100.0
80.0
68.0
56.0
60.8
WorldLine
56.0
80.0
100.0
92.0
60.0
77.6
Appendix
Table 4: Planning success rates on a representative subset of RoboTwin tasks (%). The selected tasks cover different manipulation types. All methods evaluate the same candidate trajectories.
Figure 8: Additional qualitative comparisons on AgiBotWorld and DROID. Each example shows the initial observation, action condition, ground-truth future, predictions from comparison methods, and the WorldLine prediction. AgiBotWorld examples are drawn from held-out scenes and include both successful and unsuccessful interactions, while DROID examples test transfer to unseen environments, viewpoints, and robot embodiments. Enlarged regions focus on commanded robot motion, contact, and the resulting object-state changes.
Figure 9: Three-view comparison between WorldLine predictions and ground-truth AgiBotWorld trajectories. Each example presents synchronized head, left-wrist, and right-wrist observations at multiple time steps from a held-out AgiBotWorld scene. The comparison shows robot motion, contact, and object-state evolution across the global and wrist-mounted camera views.
Figure 10: Three-view comparison between WorldLine predictions and ground-truth DROID trajectories. Each example presents synchronized head, left-wrist, and right-wrist observations at multiple time steps. WorldLine is evaluated directly on DROID without training or adaptation on DROID data. The comparison exposes the temporal alignment of robot motion, contact, and object-state evolution across global and wrist-mounted camera views.
Figure 11: Long-horizon three-view rollouts generated by Causal WorldLine. We compare the causal predictions with the corresponding ground-truth trajectories using synchronized head, left-wrist, and right-wrist observations sampled across the rollout. The visualization highlights long-term robot motion, object-state evolution, temporal continuity, and cross-view consistency under block-autoregressive generation.
Figure 12: Qualitative effects of the main WorldLine components on AgiBotWorld. Across three examples, we compare the full model with representative ablations of the training data, action-conditioning interface, and relational regularization. All variants receive the same initial observation and action sequence. Enlarged regions highlight differences in robot pose, contact, object motion, and the generation of unsuccessful interaction outcomes.
Figure 13: Selected failure cases of WorldLine. Each pair compares the ground-truth (GT) future with the WorldLine prediction. The top row shows two low-visual-quality examples from out-of-domain DROID evaluation; the middle row shows action misalignment; and the bottom row shows interaction failures involving inconsistent object or container states.