One from Infinity: Actualizing Futures from Pretrained World Models into Robot Actions
Organizations: University of California, Berkeley · Southern University of Science and Technology · Xi’an Jiaotong University
Abstract
A pretrained video world model admits many plausible futures for a scene, but a robot must realize the exact task-conditioned one. To turn world models into executable robot policies, existing methods fine-tune the heavy world model backbone using large-scale robot data and computational resources. Challenging this status quo, we argue that the expensive part has already been paid in the world model pretraining since the representation space of a video world model lays out the diverse potential futures. In this case, what remains is to select the future that accomplishes the task and to read out the actions that realize it. We formalize this task as actualization, which learns a task-conditioned selection and realization on top of a prior supplied by a frozen world model. This can be solved by a tiny actualizer model. We implement RoboActualizer with as few as 60M parameters on top of a frozen world model encoder. The actualizer is composed of two lightweight DiT experts that jointly predict future latents and actions by flow matching. The model can be trained entirely on a single GPU with 32 GB peak memory. With up to 100x fewer trainable parameters than existing WAMs and VLAs, RoboActualizer reaches great performance on simulation benchmarks including LIBERO, LIBERO-Plus, RoboTwin 2.0 and five tasks on two real-world platforms, with a low latency of 39 ms that allows real-time control.
Figures & tables
| Method | Trainable Params. | Emb. PT | Spatial | Object | Goal | Long | Overall | Latency (ms) |
| OpenVLA ( Kim et al., 2024 ) | 279M | Yes | 84.7 | 88.4 | 79.2 | 53.7 | 76.5 | 145.5 |
| ( Black et al., 2024 ) | 3.3B | Yes | 96.8 | 98.8 | 95.8 | 85.2 | 94.1 | 120.4 |
| ( Zhou et al., 2025 ) | 3.3B | Yes | 98.8 | 98.2 | 98.0 | 92.4 | 96.9 | 128.5 |
| UniVLA ( Bu et al., 2025 ) | 123M | Yes | 96.5 | 96.8 | 95.6 | 92.0 | 95.2 | 157.3 |
| Motus ( Bi et al., 2026 ) | 5.9B | Yes | 96.8 | 99.8 | 96.6 | 97.6 | 97.7 | 2230.0 |
| WorldVLA ( Cen et al., 2025 ) | 7.0B | No | 87.6 | 96.2 | 83.4 | 60.0 | 81.8 | 397.5 |
| Method | Trainable Params. | Camera | Robot | Lang. | Light | Bg. | Noise | Layout | Overall | Latency (ms) |
| OpenVLA ( Kim et al., 2024 ) | 279M | 0.8 | 3.5 | 23.0 | 8.1 | 34.8 | 15.2 | 28.5 | 15.6 | 145.5 |
| Black et al. (2024) | 3.3B | 13.8 | 6.0 | 58.8 | 85.0 | 81.4 | 79.0 | 68.9 | 53.6 | 120.4 |
| -Fast | 3.3B | 65.1 | 21.6 | 61.0 | 73.2 | 73.2 | 74.4 | 68.8 | 61.6 | 68.5 |
| UniVLA ( Bu et al., 2025 ) | 123M | 1.8 | 46.2 | 69.6 | 69.0 | 81.0 | 21.2 | 31.9 | 42.9 | 157.3 |
| WorldVLA ( Cen et al., 2025 ) | 7.04B | 0.1 | 27.9 | 41.6 | 43.7 | 17.1 | 10.9 | 38.0 | 25.0 | 397.5 |
| Fast-WAM ( Yuan et al., 2026 ) | 6.02B | 16.4 | 44.5 | 68.9 | 78.2 | 53.7 | 37.7 | 60.7 | 51.5 | 111.7 |
| Method | Trainable Params. | Success Rates (%) |
| ( Zhou et al., 2025 ) | 3.3B | 31.4 |
| LingBot-VA ( Li et al., 2026 ) | 5.3B | 17.2 |
| Fast-WAM ( Yuan et al., 2026 ) | 6.0B | 41.9 |
| RoboActualizer (ours) | 60M | 58.8 |
| Representation | Params. | Input Range | Latent Predictivity | Success Rates (%) |
| DINOv3 | 300M | 1 frame | ✗ | 39.7 |
| WAN VAE | 127M | 4 frames | ✗ | 25.9 |
| V-JEPA | 300M | 1 frame | ✗ | 46.4 |
| V-JEPA | 300M | 4 frames | ✗ | 44.8 |
| V-JEPA | 300M | 4 frames | ✓ | 60.0 |
Appendix figures & tables3 assets
Supplementary material from the paper’s appendix.
Appendix
| LIBERO | LIBERO-Plus | RoboTwin 2.0 | |
| Action chunk / executed steps | 16 / 10 | 16 / 10 | 32 / 24 |
| Euler steps | 4 | 4 | 10 |
| CFG scale | 1 | 1.5 | 1 |
| Future-image generation | off | off | off |
| Warm-up no-op steps | 30 | 30 | – |
| Task | C | R | M | Task | C | R | M |
| adjust_bottle | 100 | 100 | 100 | place_can_basket | 64 | 64 | 64 |
| beat_block_hammer | 76 | 64 | 70 | place_cans_plasticbox | 56 | 84 | 70 |
| blocks_ranking_rgb | 36 | 36 | 36 | place_container_plate | 100 | 96 | 98 |
| blocks_ranking_size | 24 | 32 | 28 | place_dual_shoes | 64 | 44 | 54 |
| click_alarmclock | 92 | 84 | 88 | place_empty_cup | 72 | 72 | 72 |
| click_bell | 100 | 92 | 96 | place_fan | 40 | 56 | 48 |