Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: https://github.com/Alexander-wu/FutureWorlds.
Figures & tables
Figure 1: Learning from alternative futures. RT-1 comparison at +32 ; metrics use full images. FutureWorlds constructs diverse futures, maintains their histories, and learns through MemSPO by comparing trajectory rewards.
Figure 2: Overview of FutureWorlds . Supervised initialization trains the multimodal world model. MemSPO searches alternative futures, maintains separate histories, and learns from relative trajectory quality. Generation and policy scoring use matching histories.
Figure 3Figure 4
Figure 5: Model size and quality. Bubble area denotes peak GPU memory on one H20.
Figure 6: Motion quality and optical-flow comparisons. (a) Full-cohort results at 32 frames. (b) Selected examples with a shared flow-color scale within each dataset.
Figure 7: Memory ablation at 32 frames. w/o Historical Memory , w/o Initial Anchor , and Full Model . Gains are versus w/o Historical Memory; PSNR and SSIM axes are truncated.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Action-conditioned Wan2.2 adaptation. Action chunks modulate time embeddings, and a history adapter supplies spatial context. Adapters and LoRA parameters are trainable; pretrained weights remain frozen.
Temperature T
Top- p
RT-1
BridgeV2
1.0
1.00
4.02
3.45
0.8
1.00
3.03
2.97
1.2
1.00
4.32
2.60
1.0
0.95
3.70
3.84
1.0
0.90
4.04
3.53
Appendix
Table 3: LPIPS reduction (%) across decoding settings.
Method
RT-1
BridgeV2
RoboCasa
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Video Generation Models
FitVid
19.89
0.6268
0.4529
17.86
0.6241
0.5079
14.29
0.5640
0.5373
iVideoGPT
16.62
0.6142
0.3333
21.91
0.8142
0.1544
17.56
0.7513
0.2009
Action-Conditioned World Models
IRASim
18.64
0.6771
0.2747
18.65
0.7442
0.2120
14.75
0.6495
0.3069
Appendix
Table 4: Ten-frame video prediction on 128 trajectories per reported dataset. Metrics: PSNR ↑ , SSIM ↑ , and LPIPS ↓ .
Figure 9: Policy feedback in a learned world model. A selected example for close middle drawer with a fixed RT-1 policy. The two blocks show the same interaction at seven time points; gold boxes mark matched details at step 112.
Figure 10: Additional comparisons of object placement and robot pose. Selected RT-1, BridgeV2, and RoboCasa examples at frame 32. Gold boxes identify matched detail views; model names are abbreviated in the figure.
Figure 11: Additional comparisons of articulated motion and object appearance. Three further selected cases at frame 32, using the same models and display convention as Figure 10 .
Figure 12: Temporal comparison on RT-1. The instruction is move green can near sponge . Columns show frames +2, +8, +16, +24, and +32 from the same saved prediction sequence for each method. FutureWorlds follows the can’s displacement toward the sponge more closely, while the displayed baselines largely retain its earlier position.
Figure 13: Temporal comparison on BridgeV2. A selected lid-manipulation sequence at the same five time points as Figure 12 . Full frames reveal the evolving gripper–object configuration and remaining differences from the recorded trajectory.
Figure 14: Temporal comparison on RoboCasa. A selected sink-interaction sequence under the same display convention. FutureWorlds more closely follows the changing arm configuration, while deviations in pose, geometry, and scene appearance are visible in the baseline sequences. These selected examples illustrate prediction behavior, not measured task success.
Figure 15: Candidate trajectories and relative learning signals. Ordinary and diverse beam search from the same SFT model and input. All four retained candidates appear in their original order at frames +2, +6, and +10. Annotations give the original ten-frame reward and group-relative advantage A . The illustrated input has the largest diverse-search reward range among the 16 inputs in the initial batch.
Method (updates)
10 frames
20 frames
32 frames
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
RT-1
SFT (50k)
23.76
0.8298
0.1308
22.49
0.8075
0.1527
21.37
0.7860
0.1734
GRPO (200)
23.97
0.8319
0.1289
22.60
0.8089
0.1517
21.39
0.7858
0.1739
Ordinary beam (200)
24.03
0.8329
0.1279
22.85
0.8124
0.1479
21.70
0.7908
0.1690
MemSPO (200)
24.33
0.8357
0.1240
23.42
0.8192
0.1397
22.45
0.8016
0.1565
Appendix
Table 5: Post-training with different candidate construction strategies. Matched GPU evaluation on 128 trajectories per dataset. PSNR ↑ , SSIM ↑ , and LPIPS ↓ ; bold marks the best score. Reduction rows compare MemSPO with ordinary beam search using unrounded LPIPS.
Memory configuration
10 frames
20 frames
32 frames
PSNR
SSIM
MSE
PSNR
SSIM
MSE
PSNR
SSIM
MSE
RT-1
w/o Historical Memory
22.45
81.81
7.37
20.81
78.83
11.00
19.89
76.73
13.21
w/o Initial Anchor
24.32
83.56
4.57
23.39
81.85
6.12
22.43
80.07
7.99
Full Model
24.33
83.57
4.56
23.42
81.92
6.01
22.45
80.16
7.82
Appendix
Table 6: Memory ablation across prediction horizons. PSNR ↑ (dB), SSIM ↑ ( ×100 ), and MSE ↓ ( ×103 ), averaged over 128 trajectories per dataset. Bold marks the best unrounded value within each dataset and horizon; display values use two decimals.
Model
Parameters (B)
Peak memory (GiB)
Latency (s)
Inference
Loaded
Allocated
Reserved
Median
Range
iVideoGPT
0.449
0.449
5.85
6.53
4.14
4.12–4.16
WEAVER
0.998
1.134
7.52
9.02
6.31
6.29–6.31
PersistWorld
1.624
2.320
6.68
7.84
22.18
22.16–22.18
Wan2.2-TI2V-5B
5.754
5.754
23.06
23.72
11.18
11.18–11.19
Cosmos-Predict2.5-2B
2.254
2.254
5.97
6.85
12.83
12.77–12.83
Appendix
Table 7: Native inference resource measurements on BridgeV2. One H20 per model, batch size one, and 32 predicted frames. Latency summarizes three fixed cases after one complete warmup. Peak memory is the maximum across the three measured runs.
Model
Resolution
Precision
Generation configuration
iVideoGPT
256×256
FP32; BF16 autocast
Sampling, T=1 , top- k=100 ; chunks 10+10+10+2 .
WEAVER
192×320
FP32; BF16 autocast
Native full-memory generation, 16 denoising steps; history 2, memory 6.
Table 8: Native settings for the resource benchmark. Resolution is internal height × width. Chunk lengths count retained future frames; Cosmos and DreamDojo generate a full final chunk and retain its first eight frames.
World models are increasingly used in embodied intelligence and generative simulation, yet their scope remains ambiguous across communities. This tutorial presents a design-space view of world models as action-conditioned predictive models that estimate the future evolution of task-relevant observations or states. We categorize existing methods into observation-space and state-space world models, comparing their trade-offs in visual fidelity, spatial structure, physical interpretability, and control usability. We further introduce world action models, which connect predicted futures with executable robot actions, and summarize four representative paradigms: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and auxiliary video prediction for policy learning. The goal of this tutorial is to clarify the conceptual scope of world (action) models and provide a structured taxonomy for embodied prediction and control.
Xiaoxiong Zhang, Xiong Zeng, Wei Zhang
School of Automation and Intelligent Manufacturing, Southern University of Science and Technology, Shenzhen, China · LimX Dynamics, Shenzhen, China
A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100x fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin 2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.
Bang Du, Yichen Xie, Shuqi Zhao +3
University of California, Berkeley · Southern University of Science and Technology · Xi’an Jiaotong University
We introduce Qwen-RobotWorld, a language-conditioned video world model for embodied intelligence. With natural language as a unified action interface, it predicts physically grounded future visual trajectories from current observations across robotic manipulation, autonomous driving, indoor navigation, and human-to-robot transfer. This unified formulation provides three promising application directions: synthetic data generation for policy training augmentation, scalable virtual environments for policy evaluation, and language-guided planning signals for downstream robot control. This is achieved through a three-part design: a) Double-Stream MMDiT with MLLM Action Encoding, where a 60-layer double-stream diffusion transformer couples frozen Qwen2.5-VL semantics with video-VAE latents through layer-wise joint attention; b) Embodied World Knowledge (EWK), an 8.6M video-text corpus (200M+ frames) with action-language mapping over 20+ embodiments and 500+ action categories; and c) General+Expert Progressive Curriculum, a two-stage training strategy that first learns general visual priors and then injects embodied specialization under a shared language interface. Extensive results show strong competitiveness: ranks 1st overall on EWMBench and DreamGen Bench, outperforms all open-source models on WorldModelBench and PBench. Additional zero-shot analyses on RoboTwin-IF benchmark further support robust generalization and multi-view consistency.