Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at project page.
Figures & tables
Figure 1: Overview of WorldLine . (a) A scalable three-stage training pipeline: Stage I learns robot dynamics from large-scale action-free videos; Stage II grounds heterogeneous robot controls through action-conditioned post-training; and Stage III distills the model for causal few-step rollout. (b) WorldLine maps an initial observation and an image-space control sequence to future videos across diverse environments and robot embodiments, including zero-shot settings. (c) WorldLine support policy evaluation and planning by predicting action outcomes before real-world execution.
Scale
Diversity
Supervision
Method
Video-only
Action data
Embod.
Tasks
Action repr.
Failure
CTRL-World Guo et al. (2026)
N/A
450 h
1
86
Native Action
60 h
GE-Sim-V2 Qiu et al. (2026)
N/A
2,500 h
1
NR
Image-space map
NR
Masked Visual Action Alzayer et al. (2026)
N/A
450 h
1
86
Masked Action
60 h
Worldline(ours)
10,000 h
2,500 h
> 10
> 3,500
Image-space map
200 h
Table 1: Training-data comparison across action-conditioned robot video models in scale, diversity, and action supervision. Embod. denotes the number of robot embodiments, Failure denotes the duration of failed trajectories, and NR denotes unreported information.
Figure 2: Data construction and three-stage training of WorldLine. Left: action-free videos are filtered for dynamics pretraining; action-labeled trajectories undergo calibration and projection checks; and failure trajectories are identified and paired with geometrically aligned controls. Right: Stage I learns manipulation dynamics, Stage II grounds heterogeneous controls with multi-view and failure-enriched training, and Stage III distills the model for efficient causal rollout.
In-Domain
DROID Evaluation
Model
AgiBotWorld (Successful)
AgiBotWorld (Failure)
DROID
PSNR ↑
SSIM ↑
LPIPS ↓
Robot IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Robot IoU ↑
PSNR ↑
SSIM ↑
LPIPS ↓
Robot IoU ↑
Video Generation Models
Cosmos3-Nano Agarwal et al. (2026)
12.54
0.4752
0.5654
0.2804
10.35
0.4723
0.6000
0.1444
16.97
0.6516
0.3492
0.2049
Cosmos3-Super Agarwal et al. (2026)
12.57
0.4685
0.5674
0.2975
10.73
0.4786
0.6216
0.1755
17.07
0.6524
0.3457
0.1978
Action-Conditioned Models
Table 2: Action-conditioned video prediction on held-out AgiBotWorld success/failure trajectories and DROID, averaged over five runs. † denotes DROID-trained methods, making DROID in-domain. Arrows indicate metric direction; green and blue mark the best and second-best results.
Figure 3: Qualitative action-conditioned video generation. Given initial frame and actions, each model predicts a frame on held-out AgiBotWorld (top) and out-of-domain DROID (bottom). Causal WorldLine better matches ground-truth robot motion and gripper–object configuration.
Figure 4: WorldLine for policy evaluation and selection. (a) Mean policy-evaluation accuracy over five generation seeds on RoboTwin and AgiBot. (b) RoboTwin task success rate under different rollout budgets for π0.5 and LingBot-VLA.
Figure 5: Causal distillation efficiency and quality. Left: sampling steps and per-frame latency under identical settings. Right: Stage-II and causal predictions on AgiBot. The causal model accelerates rollout while preserving appearance, robot geometry, and object evolution.
Figure 6: Qualitative comparison of WorldLine ablations. We compare the full model with variants that remove key data, representation, and training components, highlighting their effects on robot motion and robot–object interaction outcomes.
AgiBotWorld
DROID
RoboTwin: π0.5
RoboTwin: LingBot-VLA
Variant
LPIPS ↓
Robot IoU ↑
LPIPS ↓
Robot IoU ↑
Success ↑
Gain ↑
Success ↑
Gain ↑
Full WorldLine
0.1804
0.5622
0.2968
0.2739
78.0±1.4%
+19.1%
62.4±3.6%
+21.4%
Data
No action-free data (compute-matched)
0.2924
0.5277
0.3028
0.2614
73.2±3.3%
+14.3%
60.4±3.0%
+19.4%
Head-view-only data
–
0.4348
–
0.2665
72.8±1.10 %
+13.9%
62.0 ±3.2%
+21.0%
Success-only trajectories
0.2401
0.4721
0.3028
0.2566
77.2±2.3%
+18.3%
51.6±3.0%
+10.6%
Table 3: Ablation of WorldLine components. LPIPS and robot-mask IoU are evaluated on held-out AgiBotWorld and out-of-domain DROID. RoboTwin planning uses 32 rollouts; gains are over direct execution. LPIPS is omitted for head-view-only. The no-action-free variant matches the full model’s training-compute budget. Green and blue mark best and second-best results.
Figure 7: Matched action-encoding comparison on RoboTwin.
Appendix figures & tables7 assets
Supplementary material from the paper’s appendix.
Appendix
Method
Handover Block
Lift Pot
Open Laptop
Stack Two Blocks
Turn Switch
Average
π0.5
Direct Execution
7.5
83.1
85.6
37.5
52.5
53.3
GE-Sim-V2
32.0
96.0
84.0
56.0
48.0
63.2
OpenDW-0.5
32.0
84.0
80.0
48.0
52.0
59.2
Masked Visual Actions
0.0
100.0
80.0
68.0
56.0
60.8
WorldLine
56.0
80.0
100.0
92.0
60.0
77.6
Appendix
Table 4: Planning success rates on a representative subset of RoboTwin tasks (%). The selected tasks cover different manipulation types. All methods evaluate the same candidate trajectories.
Figure 8: Additional qualitative comparisons on AgiBotWorld and DROID. Each example shows the initial observation, action condition, ground-truth future, predictions from comparison methods, and the WorldLine prediction. AgiBotWorld examples are drawn from held-out scenes and include both successful and unsuccessful interactions, while DROID examples test transfer to unseen environments, viewpoints, and robot embodiments. Enlarged regions focus on commanded robot motion, contact, and the resulting object-state changes.
Figure 9: Three-view comparison between WorldLine predictions and ground-truth AgiBotWorld trajectories. Each example presents synchronized head, left-wrist, and right-wrist observations at multiple time steps from a held-out AgiBotWorld scene. The comparison shows robot motion, contact, and object-state evolution across the global and wrist-mounted camera views.
Figure 10: Three-view comparison between WorldLine predictions and ground-truth DROID trajectories. Each example presents synchronized head, left-wrist, and right-wrist observations at multiple time steps. WorldLine is evaluated directly on DROID without training or adaptation on DROID data. The comparison exposes the temporal alignment of robot motion, contact, and object-state evolution across global and wrist-mounted camera views.
Figure 11: Long-horizon three-view rollouts generated by Causal WorldLine. We compare the causal predictions with the corresponding ground-truth trajectories using synchronized head, left-wrist, and right-wrist observations sampled across the rollout. The visualization highlights long-term robot motion, object-state evolution, temporal continuity, and cross-view consistency under block-autoregressive generation.
Figure 12: Qualitative effects of the main WorldLine components on AgiBotWorld. Across three examples, we compare the full model with representative ablations of the training data, action-conditioning interface, and relational regularization. All variants receive the same initial observation and action sequence. Enlarged regions highlight differences in robot pose, contact, object motion, and the generation of unsuccessful interaction outcomes.
Figure 13: Selected failure cases of WorldLine. Each pair compares the ground-truth (GT) future with the WorldLine prediction. The top row shows two low-visual-quality examples from out-of-domain DROID evaluation; the middle row shows action misalignment; and the bottom row shows interaction failures involving inconsistent object or container states.
We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on thousands of hours of real-world robot data spanning teleoperation, contact-rich interaction, and on-robot policy deployment, substantially improving action-following fidelity and trajectory coverage. On top of this foundation, three new modules close the loop from video simulation to policy learning: a state expert that decodes proprioceptive state from video latents to support next-chunk prediction by downstream VLA policies; a world judge that scores generated rollouts against task instructions, yielding machine-verifiable success signals and rewards in place of manual inspection; and an acceleration framework that delivers a 25-frame rollout in 2.3 seconds on a single H100, with up to 4* frame skipping at inference for long-horizon evaluation. GE-Sim 2.0 tops the public WorldArena leaderboard at only 2B parameters, outperforming both dedicated robotic world models and closed-source general video generators, and policies trained against its rollouts and rewards translate into measurable real-world gains, establishing GE-Sim 2.0 as a practical platform for scalable evaluation and closed-loop learning of manipulation policies.
In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-rendered action videos bridge actions and pixels, giving markedly better controllability across embodiments, scenes, and tasks (PSNR +0.904 over alternative fusion baselines). (3) Sparse mixture-of-experts (MoE): sparse MoE layers add capacity for heterogeneous dynamics and absorb the action modality while reducing inter-modality conflict (FVD -6.530 vs. the dense backbone). (4) Efficient rollout generation: causal adaptation and few-step distillation yield a four-step autoregressive simulator, achieving a 5.67-fold speedup over the 35-step model. Benefiting from these designs, we train on approximately one million real-world and simulated trajectories and obtain large gains in action controllability and video quality: PSNR improves over the strongest evaluated baselines by 4.636 on AgiBotWorld Beta, 2.080 on RoboMIND, and 10.343 on RoboTwin, with the adapted EWMBench DYN score up 0.426 on RoboTwin. Relying on this, four downstream applications on RoboTwin succeed: 500 generated trajectories added to 50 demonstrations per task raise policy success from 70% to 93%; policy evaluation reaches a Pearson correlation of 0.994 across five checkpoints; and relative success gains reach 47.7% for action selection and 20.3% for policy improvement. Qualitative generalization across trajectory, scene, object, embodiment, and viewpoint shifts highlights its potential as a general-purpose world model simulator.
Shilong Zou, Shilin Zhang, Yingji Zhang +8
Beijing Innovation Center of Humanoid Robotics (X-Humanoid) WFM System Group
World models are increasingly used as policy-in-the-loop imagination environments, where reliable rollouts require fine-grained controllability with respect to low-level robot actions. A key obstacle to scaling such models in robotics is that actions are not a universal language in pixel space: changes in visual environment, camera view, robot placement, or embodiment alter how the same numerical action manifests visually, leading to conflicting supervision under mixed training and brittle generalization at deployment. We introduce SyncWorld, an action-conditioned world model that serves as a zero-shot simulator across unseen environments without any additional training. SyncWorld leverages a visual calibration episode---paired frames and actions that showcase all the controllable degrees of freedom---to specify the setup-specific Action--Visual Mapping in context. Training with visual calibration contexts teaches the model to interpret actions through visual evidence and to leverage interaction history when explicit calibration is unavailable. Experiments show that SyncWorld can accurately simulate action outcomes in previously unseen settings, and that its capability of simulating rollouts enables test-time policy improvement without training.