Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. https://github.com/Wang-Xiaodong1899/SimForcing
Figures & tables
Figure 1 : Comparison of training paradigms for action-conditioned robot world models. Left: Direct training on real videos and corresponding actions ( Zhu et al., 2025b ; Guo et al., 2026c ) . Middle: Training on real videos augmented with explicit visual priors, such as simulated videos, scene priors, and object priors ( Zhang et al., 2026a ; Ye et al., 2026a ) . Right: SimForcing learns a simulation world model from synthetic rollouts and distills its motion priors into a real-world model through latent-space motion supervision, without requiring additional scene or object decomposition.
Figure 2 : Overview of SimForcing. Simulation world model training: A mixture-of-transformers world model learns action-conditioned motion from augmented synthetic rollouts. Sim-to-real distillation: The pretrained model initializes a student and serves as a frozen teacher. Latent-motion distillation internalizes motion priors, while multi-block dropout conditioning provides simulation guidance. Joint training enables the final student to generate both simulation conditions and real videos, with CFG balancing explicit guidance and internalized knowledge at inference.
Model
Time (s) ↓
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
With embodied pretraining
IRASim ( Zhu et al., 2025b )
25.9
18.062
0.753
47.27
333.92
0.2102
Cosmos-Predict2.5 ( Ali et al., 2025 )
20.5
27.216
0.880
23.74
119.38
0.0996
Without embodied pretraining
Ctrl-World ( Guo et al., 2026c )
16.9
20.928
0.761
46.67
235.57
0.1674
EnerVerse-AC ( Jiang et al., 2025 )
16.6
22.599
0.797
28.67
181.55
0.0933
Table 1 : Action-conditioned video prediction on the Bridge ( Walke et al., 2023 ) validation set. Gray rows denote embodied pretrained models. Time denotes inference time per sample. Best results among models without embodied pretraining are in bold.
Figure 3 : Left: Qualitative comparison across methods. The first two rows show examples from Bridge ( Walke et al., 2023 ) , and the last two from InternData-A1 ( Tian et al., 2026 ) . Our method produces fewer visual hallucinations and more accurate action-conditioned motion. Right: Effect of simulation prediction quality on our method’s real-video predictions on Bridge.
Training components
Evaluation metrics
ID
FM
TA
MD
SC
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
Wan2.2 (SFT) ( Wan et al., 2025 )
22.751
0.838
23.97
168.71
0.0926
(a)
✓
✗
✗
✗
24.128
0.851
22.03
141.81
0.0751
(b)
✓
✓
✗
✗
24.523
0.854
21.52
134.56
0.0702
(c)
✓
✗
✓
✗
24.166
0.849
27.46
154.69
0.0768
(d)
✓
✓
✓
✗
24.310
0.852
19.18
100.27
0.0769
Table 2 : Component ablations on Bridge. FM, TA, MD, and SC denote joint flow matching, trajectory augmentation, motion distillation, and simulation conditioning, respectively. Best are in bold.
(a) Distillation target
(b) CFG weight
Target
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
w
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
Latent value
19.232
0.803
26.38
193.15
0.0972
0.0
24.620
0.8549
25.08
145.96
0.0716
0.3
25.014
0.8578
26.68
144.19
0.0687
Latent motion
24.166
0.849
27.46
154.69
0.0768
0.6
25.077
0.8583
25.21
137.49
0.0673
1.0
24.857
0.8567
24.70
133.66
0.0681
Table 3 : Ablation studies on distillation targets and CFG weights on the Bridge validation set. Bold indicates the best result within each comparison.
Figure 4 : Distillation loss curves for latent value (a–b) and latent motion (c–d) on the real and synthetic branches. Faint curves show logged losses, and bold curves show trailing averages over 250 training steps. The real-branch panels use different vertical scales; the synthetic-branch panels share the same scale.
Model
FM
MD
SC
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
EnerVerse-AC ( Jiang et al., 2025 )
–
–
–
23.780
0.782
48.29
378.88
0.1056
Ctrl-World ( Guo et al., 2026c )
–
–
–
24.063
0.786
44.42
363.15
0.1099
Wan2.2 (SFT) ( Wan et al., 2025 )
–
–
–
24.738
0.808
30.94
335.20
0.0827
GeniWorld ( Gu et al., 2026 )
–
–
–
24.234
0.804
41.54
355.52
0.0962
SimForcing component ablations
Joint FM only
✓
✗
✗
25.382
0.813
30.30
317.05
0.0742
Table 4 : Video prediction and component ablations on InternData-A1. Best results are in bold. FM, MD, SC denote joint flow matching, motion distillation, and simulation conditioning, respectively.
Init.
Spatial
Object
Goal
Long
Avg.
FastWAM
97.8
99.2
97.8
95.6
97.6
Action-only training
Wan2.2
88.0
95.6
82.8
64.0
82.6
Our WM
89.4
95.6
92.4
79.2
89.2 ↑
Unconditional training
Wan2.2
96.4
99.6
97.6
92.2
96.5
Table 5 : LIBERO success rates (%). Arrows indicate improvements over Wan2.2 within each training setting.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : Qualitative comparison on Bridge: example 1. From top to bottom: Baseline ( Wan et al., 2025 ) , EnerVerse-AC ( Jiang et al., 2025 ) , GeniWorld ( Gu et al., 2026 ) , our simulator-domain predictions, our real-world predictions, and ground truth. The seven columns show aligned temporal snapshots.
Figure A2 : Qualitative comparison on Bridge: example 2. Row order and temporal sampling are identical to Figure A1 .
Figure A3 : Qualitative comparison on Bridge: example 3. Row order and temporal sampling are identical to Figure A1 .
Figure A4 : Qualitative comparison on InternData-A1: example 1. From top to bottom: Baseline ( Wan et al., 2025 ) , EnerVerse-AC ( Jiang et al., 2025 ) , GeniWorld ( Gu et al., 2026 ) , our simulator-domain predictions, our real-world predictions, and ground truth. The seven columns show aligned temporal snapshots.
Figure A5 : Qualitative comparison on InternData-A1: example 2. Row order and temporal sampling are identical to Figure A4 .
Figure A6 : Qualitative comparison on InternData-A1: example 3. Row order and temporal sampling are identical to Figure A4 .
w=0
w=0.3
w=0.6
w=1.0
p
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
0.0
24.441
0.8524
24.441
0.8524
24.441
0.8524
24.441
0.8524
0.3
24.675
0.8547
24.825
0.8561
24.890
0.8567
24.896
0.8566
0.7
24.620
0.8549
25.014
0.8578
25.077
0.8583
24.857
0.8567
1.0
21.832
0.8408
23.570
0.8502
24.387
0.8536
24.009
0.8508
Appendix
Table A1: Joint evaluation of training-time conditioning probability p and inference-time guidance weight w . Higher PSNR (dB) and SSIM are better. Bold indicates the best result in the entire table for each metric.
Inference steps
CFG weight w
Ns
Nr
0
0.3
0.6
1
1
1
22.785
23.116
23.301
23.206
1
5
24.263
24.697
24.851
24.629
1
10
24.576
24.958
25.063
24.838
1
20
24.620
24.983
25.063
24.843
5
1
22.785
23.130
23.314
23.177
Appendix
Table A2: PSNR (dB) across inference-step budgets and CFG weights on the 100-sample Bridge evaluation. Ns and Nr denote simulation and real-video generation steps. Bold indicates the best result within each row.
Inference steps
CFG weight w
Ns
Nr
0
0.3
0.6
1
1
1
0.8231
0.8264
0.8280
0.8271
1
5
0.8483
0.8518
0.8528
0.8510
1
10
0.8539
0.8569
0.8575
0.8559
1
20
0.8549
0.8576
0.8582
0.8567
5
1
0.8231
0.8266
0.8281
0.8268
Appendix
Table A3: SSIM across inference-step budgets and CFG weights on the 100-sample Bridge evaluation. Ns and Nr denote simulation and real-video generation steps. Bold indicates the best result within each row.
Inference steps
CFG weight w
Ns
Nr
0
0.3
0.6
1
1
1
0.644
0.649
0.649
0.571
1
5
1.251
1.258
1.259
0.879
1
10
2.002
2.008
2.008
1.245
1
20
3.519
3.546
3.539
2.012
5
1
0.934
0.941
0.940
0.862
Appendix
Table A4: Mean inference latency (s/sample) across inference-step budgets and CFG weights.
Setting
BridgeData V2
InternData-A1
Video frames per window
21
33
Actions per window
20
128
Actions per video transition
1
4
Batch size per device
16
10
λFMr
1.0
1.0
λFMs
0.5
0.5
Appendix
Table A5 : Dataset-specific training settings. Batch sizes count paired simulation–real samples per device.
Real-world robot learning is constrained by the cost of collecting experience and evaluating candidate behaviors. Video generation models offer a scalable foundation for visual simulators that predict action outcomes before physical execution. Yet they often favor visual plausibility over accurate action following and coherent robot--object dynamics, while action-conditioned simulators depend on scarce, embodiment-specific data that are difficult to share across incompatible control spaces. We introduce WorldLine, an action-driven visual simulator that decouples transferable dynamics learning from heterogeneous action grounding. WorldLine learns manipulation dynamics from more than 10,000 hours of action-free robot videos and grounds them using over 2,000 hours of action trajectories across more than ten embodiments. An image-space action representation provides a shared control interface across embodiments, while multi-view and failure-enriched training with relational regularization improves interaction-sensitive prediction. Robot-focused few-step distillation enables efficient causal rollout while preserving action-critical motion. Across held-out and out-of-domain settings, WorldLine maintains strong visual quality and robot-motion agreement; on failed trajectories, it improves robot-mask IoU by 0.1626 over the strongest baseline. It predicts trajectory success with 74% mean accuracy across RoboTwin and AgiBot, one percentage point above the strongest baseline. Without RoboTwin training or adaptation, its rollouts improve task success by up to 21.4 percentage points over direct policy execution. Together, these capabilities make WorldLine a scalable and efficient visual simulator for policy evaluation and embodied planning. More results are available at project page.
Shenghe Zheng, Wenbo Li, Jiyao Zhang +4
The Hong Kong University of Science and Technology · Joy Future Academy · Peking University +1
We study action-conditioned world modeling as a scalable way to learn transferable dynamics priors for robot learning. By pretraining a model to predict how actions drive visual scene evolution, the resulting world model captures reusable interaction dynamics beyond appearance-level video generation. Concretely, we pretrain a multi-view interactive base diffusion world model, A2World, on large-scale robot manipulation data with real action annotations. We validate the learned dynamics priors from two complementary perspectives. First, we adapt A2World into a task- or scene-specialized real-world simulator, A2World-sim, whose long-horizon rollouts support simulator-based policy evaluation and scalable what-if analysis by replacing real-robot rollouts with world model rollouts. Second, starting from the same pretrained weights, we adapt A2World into a video-action joint prediction model, A2World-policy, that predicts actions under visual and instruction conditioning. Experiments across simulation benchmarks and real-robot settings demonstrate that action-conditioned world model pretraining yields transferable dynamics priors that benefit both simulator-centric and policy-centric robot learning.
Ze Huang, Jiahui Zhang, Hairuo Liu +3
School of Data Science, Fudan University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China +1
Video generation models have emerged as a promising paradigm for embodied world simulation. However, both general-domain video generators and robot-specific data fine-tuned models can still produce physically implausible manipulations, including discontinuous motion trajectories and inconsistent robot-object interactions, which limits their reliability as world simulators. Through extensive experiments, we find that such physical instability mainly arises from two factors: deformation of moving objects and implausible spatio-temporal correlations among interacting entities, particularly during contact. Building on this observation, we propose PhysisForcing, a scalable training framework that strengthens physical consistency by focusing supervision on physics-informative regions through joint optimization of pixel-level and semantic-level features. The framework consists of a pixel-level trajectory alignment loss, which supervises DiT features using reference point trajectories, and a semantic-level relational alignment loss, which aligns DiT features with inter-region relations extracted from a frozen video understanding encoder. Extensive experiments on R-Bench, PAI-Bench, and EZS-Bench show that PhysisForcing consistently improves embodied video generation over strong baselines, improving the Wan2.2-I2V-A14B and Cosmos3-Nano base models on R-Bench by 22.3% and 9.2% (7.1% and 3.7% over vanilla finetuning), with the Cosmos3-Nano variant attaining the best overall score. Beyond generation, as a world model under the WorldArena action-planner protocol it raises the closed-loop success rate from 16.0% to 24.0% and further improves downstream policy success, indicating that physically aligned video models yield stronger representations for robotic manipulation.