MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling
Authors: Jie Chen, Ruofei Bai, Yuxin Cai, Yifeng Zhang, Chengyang He, Jun Li, Wei-Yun Yau, Guillaume Sartoretti
Organizations: Department of Mechanical Engineering, National University of Singapore, Singapore · A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), Singapore · Nanyang Technological University, Singapore
World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65× fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8× speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.
Figures & tables
Figure 1: MiniWAM replaces costly native future prediction with compact predictive representations. (a) While conventional WAMs predict high-dimensional native future features, MiniWAM jointly predicts actions and compact PRISM targets learned from the same visual features via inverse dynamics and reconstruction. (b) MiniWAM remains competitive with substantially larger WAMs and VLAs on LIBERO-Plus . (c) In controlled comparisons across VAE and DINOv3 feature spaces, PRISM simultaneously improves success rate and reduces training time relative to native future-feature prediction.
Figure 2: Two-stage training pipeline of PRISM and MiniWAM. In Stage 1 , privileged current–future visual features are compressed by the PRISM encoder into a compact predictive representation Lt , shaped through inverse action prediction and feature reconstruction. The corresponding world tokens remain clean while the action tokens are corrupted for conditional action prediction. In Stage 2 , the PRISM encoder is frozen and provides compact world-modeling targets; MiniWAM jointly denoises the world and action tokens using its world and action experts. Solid and hatched tokens denote clean and noised tokens, respectively. Language and proprioceptive conditioning are omitted for clarity.
Visual features
Co-training target
LIBERO
LIBERO-Plus-2K
SR (%) ↑
SR (%) ↑
VAE
w/o Co-Train
62.3
25.1
w/ Native Co-Train
92.8
50.2
w/ PRISM
94.8
56.5
DINOv3
w/o Co-Train
89.0
56.5
w/ Native Co-Train
93.2
60.9
Table 1: Controlled comparison of world-modeling targets on LIBERO and LIBERO-Plus-2K . All variants use the same 0.25B world–action architecture and training protocol within each visual feature family.
Horizon 16
Horizon 32
Visual Feature
Target
Throughput
S2
Total
Throughput
S2
Total
↑
↑
↑
↑
↑
↑
LIBERO
WAN2.1 VAE
Native
30.81
1.00 ×
1.00 ×
14.20
1.00 ×
1.00 ×
PRISM
142.04
4.61 ×
3.55 ×
124.92
8.80 ×
6.77 ×
DINOv3
Native
36.17
1.00 ×
1.00 ×
15.55
1.00 ×
1.00 ×
Table 2: Training efficiency of PRISM versus native co-training. Throughput is measured in samples/s with global batch 256 under matched hardware. S2 is Stage-2 speedup; Total estimates training speedup including Stage 1, assuming equal per-update costs for the two PRISM stages (Appendix D ).
Type
Method
Scale
Embodied PT.
Backbone Init.
LIBERO SR (%) ↑
LIBERO-Plus SR (%) ↑
RoboTwin 2.0 SR (%) ↑
VLA
OpenVLA-OFT
7B
✓
VLM
96.6
71.4
-
π0
3.3B
✓
VLM
93.5
70.5
31.4
π0 -FAST †
3.3B
✓
VLM
85.5
61.6
-
UniVLA †
7B
✓
VLM
95.2
42.9
-
StarVLA-OFT †
4B
×
VLM
96.6
-
24.8
Xiaomi-Robotics-0 †
4.7B
✓
VLM
98.7
-
40.5
Table 3: Comparison with state-of-the-art robot policies. † denotes externally reported results. Scale is the reported model size; Embodied PT and Backbone Init. denote embodied pretraining and backbone initialization. “–” means no pretrained backbone and “-” means unavailable.
Bottleneck
Recon.
Inv. dyn.
Both
4×4 , D=64
69.5
70.5
74.1
2×2 , D=64
63.3
67.5
67.8
1×1 , D=16
57.0
56.8
61.7
Table 4: PRISM ablation on LIBERO-Plus-2K . Success rate (%).
Table 6: Behavioral structure on unseen LIBERO-90 tasks. Leave-task-out retrieval results are reported as Raw → PRISM.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Stage 1
Stage 2
Optimizer; peak LR; decay
AdamW; 5⋅10−4 ; 0
same
LR schedule; warmup
cosine; 5% of updates
same
Global batch; precision; seed
256; BF16; 42
same
LIBERO duration
15 epochs
50 epochs
RoboTwin duration
6 epochs
25 epochs
LIBERO action / visual horizon
16 / 16
same
Appendix
Table 7: Shared 0.25B training settings. Microbatch and gradient accumulation vary with GPU count while their product with GPU count remains 256.
Method
LIBERO
LIBERO-Plus
RoboTwin 2.0
π0 -FAST
85.5 ( Fei et al., 2026 )
61.6 ( Fei et al., 2026 )
–
UniVLA
95.2 ( Bu et al., 2025 )
42.9 ( Fei et al., 2026 )
–
StarVLA-OFT
96.6 ( Community, 2026 )
–
24.8 ( RoboTwin Team, 2026 )
Xiaomi-Robotics-0
98.7 ( Cai et al., 2026b )
–
40.5 ( RoboTwin Team, 2026 )
WorldVLA
79.1 ( Cen et al., 2025 ; Fei et al., 2026 )
25.0 ( Fei et al., 2026 )
–
AHA-WAM
–
–
33.8 ( RoboTwin Team, 2026 )
Appendix
Table 8: Sources of externally reported results in Table 3 . Values are success rates (%).
Benchmark
Visual
Native
PRISM
Pipeline
horizon
Core
Pipeline
Core
Pipeline
speedup
LIBERO
16
36.17
35.16
143.25
138.81
3.95×
32
15.55
15.26
126.09
123.08
8.07×
RoboTwin 2.0
16
21.48
21.11
93.57
92.86
4.40×
32
8.61
8.41
81.47
80.74
9.60×
Appendix
Table 9: DINOv3 training throughput through the data pipeline. Samples per second on one RTX PRO 6000 GPU at global batch 256. Core values reproduce the compute measurements in Table 2 ; pipeline values include data loading and transfer. Speedup compares PRISM and native Stage-2 pipeline throughput.
View
Objective
Action RMSE ↓
EEF Err. (m) ↓
Grip. Agr. ↑
Main
Reconstruction only
0.414
0.090
0.747
PRISM (I.D. + recon.)
0.356
0.068
0.802
Wrist
Reconstruction only
0.326
0.087
0.907
PRISM (I.D. + recon.)
0.266
0.078
0.975
Appendix
Table 10: Effect of inverse-dynamics supervision on neighborhood behavior. Matched DINOv3 bottlenecks on unseen LIBERO-90 tasks; both retain the full posterior-mean grid.
World Action Models (WAMs) extend robot policy learning by incorporating future prediction as an additional training objective, encouraging the policy to encode task-relevant temporal structure in its representations. Current WAMs often rely on large-scale generative architectures that incur high training costs and inference latency, making them difficult to deploy as efficient closed-loop policies. We propose Light-WAM, a lightweight World Action Model for efficient robot manipulation. Specifically, it is built with a compact video backbone and performs future-video supervision in a downsampled latent space, reducing the cost of video co-training while retaining its benefits for representation learning. For action prediction, Light-WAM introduces the StateFusionActionExpert, which reads adapted states from multiple backbone layers, fuses them through learned-query pooling, and directly predicts action chunks in a single forward pass. This design provides an efficient interface between video backbone representations and robot actions, avoiding the need for heavy generative action experts. Experiments demonstrate that Light-WAM maintains strong performance on LIBERO and achieves usable multi-task performance on RoboTwin 2.0, while using only 0.44B trainable parameters. It also achieves 72.03ms inference latency with 4.1GiB peak GPU memory and improved training throughput.
Ziang Li, Dongzhou Cheng, Yibin Wang +5
Wuhan University · Shanghai Innovation Institute · Southeast University +2
World Action Models (WAMs) improve robot policy learning by incorporating future dynamics, yet explicitly generating future videos at inference introduces substantial computational overhead. Removing future generation improves efficiency, but leaves future dynamics only implicitly encoded in observation features, which can limit robustness under distribution shifts. We propose MoWAM, an efficient WAM that replaces future video generation with explicit future motion prediction. Instead of reconstructing the complete future scene, MoWAM models structured robot motion as a compact abstraction of the future, capturing how the robot is expected to evolve under the current scene and interaction constraints. A Mixture-of-Transformer architecture learns future visual dynamics during training while jointly predicting motion and action, allowing video generation to be removed entirely at inference while retaining an explicit representation of the future. The compact motion representation further enables efficient inference-time scaling by sampling multiple candidates of motion and action pairs and selecting among them with a motion-aware task-progress verifier. Experiments on LIBERO, LIBERO-Plus, and real-world manipulation tasks demonstrate that MoWAM achieves strong in-distribution performance, improved out-of-distribution robustness, and higher average real-world success than representative WAM baselines. In addition, performance improves as more candidates are explored, demonstrating that explicit future motion provides an effective and efficient basis for inference-time scaling.
Jiayu Wang, Bin Zhu, Yue Yu +1
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · Singapore Management University, Singapore · Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China
World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.
Jiajun Li, Tiecheng Guo, Yifan Ye +9
1The University of Hong Kong · 2Peking University · 3Muka Robotics +2