Robotic world models predict action-conditioned future scenes, providing a foundation for understanding action outcomes. However, turning alternative predictions into useful learning signals remains challenging: similar candidates limit informative quality comparisons, while diverging trajectories require persistent maintenance of their individual histories. We introduce FutureWorlds, a framework that unifies candidate construction, history maintenance, and learning from relative quality. Built on a multimodal discrete autoregressive model, FutureWorlds uses diverse beam search during reinforcement learning to construct candidate futures that balance confidence and diversity. Candidate-specific bounded memory preserves scene states and ensures that generation and policy scoring use matching histories. We further propose MemSPO (Memory-Conditioned Search-Guided Policy Optimization), which converts video trajectory rewards into group-relative advantages to optimize the world model. On RT-1, BridgeV2, and RoboCasa, FutureWorlds reduces LPIPS for 32-frame predictions by 14.78%, 20.84%, and 9.12%, respectively, relative to the strongest baseline on each dataset. Under fixed evaluation configurations, only 200 MemSPO updates further improve generation quality and support continued prediction beyond the training horizon. Memory ablations, decoding sensitivity analysis, and optical-flow evaluation show that these gains extend beyond visual quality to more accurate motion prediction and more consistent object states. Project page and code: https://github.com/Alexander-wu/FutureWorlds.
Figures & tables
Figure 1: Learning from alternative futures. RT-1 comparison at +32 ; metrics use full images. FutureWorlds constructs diverse futures, maintains their histories, and learns through MemSPO by comparing trajectory rewards.
Figure 2: Overview of FutureWorlds . Supervised initialization trains the multimodal world model. MemSPO searches alternative futures, maintains separate histories, and learns from relative trajectory quality. Generation and policy scoring use matching histories.
Figure 3Figure 4
Figure 5: Model size and quality. Bubble area denotes peak GPU memory on one H20.
Figure 6: Motion quality and optical-flow comparisons. (a) Full-cohort results at 32 frames. (b) Selected examples with a shared flow-color scale within each dataset.
Figure 7: Memory ablation at 32 frames. w/o Historical Memory , w/o Initial Anchor , and Full Model . Gains are versus w/o Historical Memory; PSNR and SSIM axes are truncated.
Appendix figures & tables14 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 8: Action-conditioned Wan2.2 adaptation. Action chunks modulate time embeddings, and a history adapter supplies spatial context. Adapters and LoRA parameters are trainable; pretrained weights remain frozen.
Temperature T
Top- p
RT-1
BridgeV2
1.0
1.00
4.02
3.45
0.8
1.00
3.03
2.97
1.2
1.00
4.32
2.60
1.0
0.95
3.70
3.84
1.0
0.90
4.04
3.53
Appendix
Table 3: LPIPS reduction (%) across decoding settings.
Method
RT-1
BridgeV2
RoboCasa
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
Video Generation Models
FitVid
19.89
0.6268
0.4529
17.86
0.6241
0.5079
14.29
0.5640
0.5373
iVideoGPT
16.62
0.6142
0.3333
21.91
0.8142
0.1544
17.56
0.7513
0.2009
Action-Conditioned World Models
IRASim
18.64
0.6771
0.2747
18.65
0.7442
0.2120
14.75
0.6495
0.3069
Appendix
Table 4: Ten-frame video prediction on 128 trajectories per reported dataset. Metrics: PSNR ↑ , SSIM ↑ , and LPIPS ↓ .
Figure 9: Policy feedback in a learned world model. A selected example for close middle drawer with a fixed RT-1 policy. The two blocks show the same interaction at seven time points; gold boxes mark matched details at step 112.
Figure 10: Additional comparisons of object placement and robot pose. Selected RT-1, BridgeV2, and RoboCasa examples at frame 32. Gold boxes identify matched detail views; model names are abbreviated in the figure.
Figure 11: Additional comparisons of articulated motion and object appearance. Three further selected cases at frame 32, using the same models and display convention as Figure 10 .
Figure 12: Temporal comparison on RT-1. The instruction is move green can near sponge . Columns show frames +2, +8, +16, +24, and +32 from the same saved prediction sequence for each method. FutureWorlds follows the can’s displacement toward the sponge more closely, while the displayed baselines largely retain its earlier position.
Figure 13: Temporal comparison on BridgeV2. A selected lid-manipulation sequence at the same five time points as Figure 12 . Full frames reveal the evolving gripper–object configuration and remaining differences from the recorded trajectory.
Figure 14: Temporal comparison on RoboCasa. A selected sink-interaction sequence under the same display convention. FutureWorlds more closely follows the changing arm configuration, while deviations in pose, geometry, and scene appearance are visible in the baseline sequences. These selected examples illustrate prediction behavior, not measured task success.
Figure 15: Candidate trajectories and relative learning signals. Ordinary and diverse beam search from the same SFT model and input. All four retained candidates appear in their original order at frames +2, +6, and +10. Annotations give the original ten-frame reward and group-relative advantage A . The illustrated input has the largest diverse-search reward range among the 16 inputs in the initial batch.
Method (updates)
10 frames
20 frames
32 frames
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
PSNR
SSIM
LPIPS
RT-1
SFT (50k)
23.76
0.8298
0.1308
22.49
0.8075
0.1527
21.37
0.7860
0.1734
GRPO (200)
23.97
0.8319
0.1289
22.60
0.8089
0.1517
21.39
0.7858
0.1739
Ordinary beam (200)
24.03
0.8329
0.1279
22.85
0.8124
0.1479
21.70
0.7908
0.1690
MemSPO (200)
24.33
0.8357
0.1240
23.42
0.8192
0.1397
22.45
0.8016
0.1565
Appendix
Table 5: Post-training with different candidate construction strategies. Matched GPU evaluation on 128 trajectories per dataset. PSNR ↑ , SSIM ↑ , and LPIPS ↓ ; bold marks the best score. Reduction rows compare MemSPO with ordinary beam search using unrounded LPIPS.
Memory configuration
10 frames
20 frames
32 frames
PSNR
SSIM
MSE
PSNR
SSIM
MSE
PSNR
SSIM
MSE
RT-1
w/o Historical Memory
22.45
81.81
7.37
20.81
78.83
11.00
19.89
76.73
13.21
w/o Initial Anchor
24.32
83.56
4.57
23.39
81.85
6.12
22.43
80.07
7.99
Full Model
24.33
83.57
4.56
23.42
81.92
6.01
22.45
80.16
7.82
Appendix
Table 6: Memory ablation across prediction horizons. PSNR ↑ (dB), SSIM ↑ ( ×100 ), and MSE ↓ ( ×103 ), averaged over 128 trajectories per dataset. Bold marks the best unrounded value within each dataset and horizon; display values use two decimals.
Model
Parameters (B)
Peak memory (GiB)
Latency (s)
Inference
Loaded
Allocated
Reserved
Median
Range
iVideoGPT
0.449
0.449
5.85
6.53
4.14
4.12–4.16
WEAVER
0.998
1.134
7.52
9.02
6.31
6.29–6.31
PersistWorld
1.624
2.320
6.68
7.84
22.18
22.16–22.18
Wan2.2-TI2V-5B
5.754
5.754
23.06
23.72
11.18
11.18–11.19
Cosmos-Predict2.5-2B
2.254
2.254
5.97
6.85
12.83
12.77–12.83
Appendix
Table 7: Native inference resource measurements on BridgeV2. One H20 per model, batch size one, and 32 predicted frames. Latency summarizes three fixed cases after one complete warmup. Peak memory is the maximum across the three measured runs.
Model
Resolution
Precision
Generation configuration
iVideoGPT
256×256
FP32; BF16 autocast
Sampling, T=1 , top- k=100 ; chunks 10+10+10+2 .
WEAVER
192×320
FP32; BF16 autocast
Native full-memory generation, 16 denoising steps; history 2, memory 6.
Table 8: Native settings for the resource benchmark. Resolution is internal height × width. Chunk lengths count retained future frames; Cosmos and DreamDojo generate a full final chunk and retain its first eight frames.