Action-conditioned robot world models must respond precisely to robot trajectories while preserving realistic visual dynamics, yet learning both from heterogeneous robot videos remains challenging. Simulation offers structured motion supervision, but appearance differences hinder direct transfer, and inaccurate simulation predictions can misguide real-video generation. We present SimForcing, a simulation-guided framework that uses simulation both as a source of transferable motion knowledge and as a controllable reference for prediction. First, we transfer motion knowledge from a simulation teacher through latent-motion distillation, aligning temporal changes in latent space to internalize motion priors while mitigating the influence of appearance differences. Second, we introduce multi-block simulation conditioning with condition dropout to exploit predicted simulation trajectories without relying excessively on their accuracy. Our simulation-conditioning classifier-free guidance scheme unifies these two ideas by balancing predictions based on internalized motion knowledge with those additionally guided by simulation latents. The jointly trained student generates both simulation conditions and real-domain videos, requiring no additional world model at inference. On Bridge, SimForcing achieves the best PSNR, SSIM, LPIPS, and FVD among the compared methods without external embodied pretraining. Evaluation on InternData-A1 further supports its applicability across robot datasets. Moreover, using our trained world model to initialize a vision-language-action model improves LIBERO success, suggesting its utility for downstream policy learning. https://github.com/Wang-Xiaodong1899/SimForcing
Figures & tables
Figure 1 : Comparison of training paradigms for action-conditioned robot world models. Left: Direct training on real videos and corresponding actions ( Zhu et al., 2025b ; Guo et al., 2026c ) . Middle: Training on real videos augmented with explicit visual priors, such as simulated videos, scene priors, and object priors ( Zhang et al., 2026a ; Ye et al., 2026a ) . Right: SimForcing learns a simulation world model from synthetic rollouts and distills its motion priors into a real-world model through latent-space motion supervision, without requiring additional scene or object decomposition.
Figure 2 : Overview of SimForcing. Simulation world model training: A mixture-of-transformers world model learns action-conditioned motion from augmented synthetic rollouts. Sim-to-real distillation: The pretrained model initializes a student and serves as a frozen teacher. Latent-motion distillation internalizes motion priors, while multi-block dropout conditioning provides simulation guidance. Joint training enables the final student to generate both simulation conditions and real videos, with CFG balancing explicit guidance and internalized knowledge at inference.
Model
Time (s) ↓
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
With embodied pretraining
IRASim ( Zhu et al., 2025b )
25.9
18.062
0.753
47.27
333.92
0.2102
Cosmos-Predict2.5 ( Ali et al., 2025 )
20.5
27.216
0.880
23.74
119.38
0.0996
Without embodied pretraining
Ctrl-World ( Guo et al., 2026c )
16.9
20.928
0.761
46.67
235.57
0.1674
EnerVerse-AC ( Jiang et al., 2025 )
16.6
22.599
0.797
28.67
181.55
0.0933
Table 1 : Action-conditioned video prediction on the Bridge ( Walke et al., 2023 ) validation set. Gray rows denote embodied pretrained models. Time denotes inference time per sample. Best results among models without embodied pretraining are in bold.
Figure 3 : Left: Qualitative comparison across methods. The first two rows show examples from Bridge ( Walke et al., 2023 ) , and the last two from InternData-A1 ( Tian et al., 2026 ) . Our method produces fewer visual hallucinations and more accurate action-conditioned motion. Right: Effect of simulation prediction quality on our method’s real-video predictions on Bridge.
Training components
Evaluation metrics
ID
FM
TA
MD
SC
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
Wan2.2 (SFT) ( Wan et al., 2025 )
22.751
0.838
23.97
168.71
0.0926
(a)
✓
✗
✗
✗
24.128
0.851
22.03
141.81
0.0751
(b)
✓
✓
✗
✗
24.523
0.854
21.52
134.56
0.0702
(c)
✓
✗
✓
✗
24.166
0.849
27.46
154.69
0.0768
(d)
✓
✓
✓
✗
24.310
0.852
19.18
100.27
0.0769
Table 2 : Component ablations on Bridge. FM, TA, MD, and SC denote joint flow matching, trajectory augmentation, motion distillation, and simulation conditioning, respectively. Best are in bold.
(a) Distillation target
(b) CFG weight
Target
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
w
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
Latent value
19.232
0.803
26.38
193.15
0.0972
0.0
24.620
0.8549
25.08
145.96
0.0716
0.3
25.014
0.8578
26.68
144.19
0.0687
Latent motion
24.166
0.849
27.46
154.69
0.0768
0.6
25.077
0.8583
25.21
137.49
0.0673
1.0
24.857
0.8567
24.70
133.66
0.0681
Table 3 : Ablation studies on distillation targets and CFG weights on the Bridge validation set. Bold indicates the best result within each comparison.
Figure 4 : Distillation loss curves for latent value (a–b) and latent motion (c–d) on the real and synthetic branches. Faint curves show logged losses, and bold curves show trailing averages over 250 training steps. The real-branch panels use different vertical scales; the synthetic-branch panels share the same scale.
Model
FM
MD
SC
PSNR ↑
SSIM ↑
FID ↓
FVD ↓
LPIPS ↓
EnerVerse-AC ( Jiang et al., 2025 )
–
–
–
23.780
0.782
48.29
378.88
0.1056
Ctrl-World ( Guo et al., 2026c )
–
–
–
24.063
0.786
44.42
363.15
0.1099
Wan2.2 (SFT) ( Wan et al., 2025 )
–
–
–
24.738
0.808
30.94
335.20
0.0827
GeniWorld ( Gu et al., 2026 )
–
–
–
24.234
0.804
41.54
355.52
0.0962
SimForcing component ablations
Joint FM only
✓
✗
✗
25.382
0.813
30.30
317.05
0.0742
Table 4 : Video prediction and component ablations on InternData-A1. Best results are in bold. FM, MD, SC denote joint flow matching, motion distillation, and simulation conditioning, respectively.
Init.
Spatial
Object
Goal
Long
Avg.
FastWAM
97.8
99.2
97.8
95.6
97.6
Action-only training
Wan2.2
88.0
95.6
82.8
64.0
82.6
Our WM
89.4
95.6
92.4
79.2
89.2 ↑
Unconditional training
Wan2.2
96.4
99.6
97.6
92.2
96.5
Table 5 : LIBERO success rates (%). Arrows indicate improvements over Wan2.2 within each training setting.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1 : Qualitative comparison on Bridge: example 1. From top to bottom: Baseline ( Wan et al., 2025 ) , EnerVerse-AC ( Jiang et al., 2025 ) , GeniWorld ( Gu et al., 2026 ) , our simulator-domain predictions, our real-world predictions, and ground truth. The seven columns show aligned temporal snapshots.
Figure A2 : Qualitative comparison on Bridge: example 2. Row order and temporal sampling are identical to Figure A1 .
Figure A3 : Qualitative comparison on Bridge: example 3. Row order and temporal sampling are identical to Figure A1 .
Figure A4 : Qualitative comparison on InternData-A1: example 1. From top to bottom: Baseline ( Wan et al., 2025 ) , EnerVerse-AC ( Jiang et al., 2025 ) , GeniWorld ( Gu et al., 2026 ) , our simulator-domain predictions, our real-world predictions, and ground truth. The seven columns show aligned temporal snapshots.
Figure A5 : Qualitative comparison on InternData-A1: example 2. Row order and temporal sampling are identical to Figure A4 .
Figure A6 : Qualitative comparison on InternData-A1: example 3. Row order and temporal sampling are identical to Figure A4 .
w=0
w=0.3
w=0.6
w=1.0
p
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
PSNR ↑
SSIM ↑
0.0
24.441
0.8524
24.441
0.8524
24.441
0.8524
24.441
0.8524
0.3
24.675
0.8547
24.825
0.8561
24.890
0.8567
24.896
0.8566
0.7
24.620
0.8549
25.014
0.8578
25.077
0.8583
24.857
0.8567
1.0
21.832
0.8408
23.570
0.8502
24.387
0.8536
24.009
0.8508
Appendix
Table A1: Joint evaluation of training-time conditioning probability p and inference-time guidance weight w . Higher PSNR (dB) and SSIM are better. Bold indicates the best result in the entire table for each metric.
Inference steps
CFG weight w
Ns
Nr
0
0.3
0.6
1
1
1
22.785
23.116
23.301
23.206
1
5
24.263
24.697
24.851
24.629
1
10
24.576
24.958
25.063
24.838
1
20
24.620
24.983
25.063
24.843
5
1
22.785
23.130
23.314
23.177
Appendix
Table A2: PSNR (dB) across inference-step budgets and CFG weights on the 100-sample Bridge evaluation. Ns and Nr denote simulation and real-video generation steps. Bold indicates the best result within each row.
Inference steps
CFG weight w
Ns
Nr
0
0.3
0.6
1
1
1
0.8231
0.8264
0.8280
0.8271
1
5
0.8483
0.8518
0.8528
0.8510
1
10
0.8539
0.8569
0.8575
0.8559
1
20
0.8549
0.8576
0.8582
0.8567
5
1
0.8231
0.8266
0.8281
0.8268
Appendix
Table A3: SSIM across inference-step budgets and CFG weights on the 100-sample Bridge evaluation. Ns and Nr denote simulation and real-video generation steps. Bold indicates the best result within each row.
Inference steps
CFG weight w
Ns
Nr
0
0.3
0.6
1
1
1
0.644
0.649
0.649
0.571
1
5
1.251
1.258
1.259
0.879
1
10
2.002
2.008
2.008
1.245
1
20
3.519
3.546
3.539
2.012
5
1
0.934
0.941
0.940
0.862
Appendix
Table A4: Mean inference latency (s/sample) across inference-step budgets and CFG weights.
Setting
BridgeData V2
InternData-A1
Video frames per window
21
33
Actions per window
20
128
Actions per video transition
1
4
Batch size per device
16
10
λFMr
1.0
1.0
λFMs
0.5
0.5
Appendix
Table A5 : Dataset-specific training settings. Batch sizes count paired simulation–real samples per device.
School of Data Science, Fudan University, Shanghai, China · Shanghai Innovation Institute, Shanghai, China · Shanghai Jiao Tong University, Shanghai, China +1