MiniWAM: Learning Compact Future Targets for Efficient World-Action Modeling
Authors: Jie Chen, Ruofei Bai, Yuxin Cai, Yifeng Zhang, Chengyang He, Jun Li, Wei-Yun Yau, Guillaume Sartoretti
Organizations: Department of Mechanical Engineering, National University of Singapore, Singapore · A*STAR Institute of Advanced Intelligence and Computing (A*STAR IAIC), Singapore · Nanyang Technological University, Singapore
World modeling has emerged as an effective co-training objective for robot policies, giving rise to World Action Models (WAMs) that jointly predict actions and future states. However, most WAMs predict future states in the native representation space of pretrained visual backbones, resulting in high-dimensional targets with substantial training cost. We introduce MiniWAM, which instead predicts compact future representations learned from privileged current-future transitions. To construct these targets, we propose Predictive Representations via Inverse Spatiotemporal Modeling (PRISM), which combines inverse-dynamics supervision with feature reconstruction to emphasize control-relevant transition information while preserving useful future-state information. With the learned PRISM encoder frozen, MiniWAM is trained to jointly predict the resulting targets and robot actions from current observations. With 65× fewer native future feature tokens, MiniWAM consistently outperforms native future-feature prediction with both DINOv3 and WAN2.1 VAE features, while achieving up to an 8× speedup in world-action training. At 0.25B parameters, MiniWAM is already competitive with substantially larger WAMs on LIBERO, LIBERO-Plus, and RoboTwin 2.0 simulation benchmarks. Representation analyses further show that PRISM contributes behavioral structure beyond feature reconstruction alone. These results demonstrate that effective world-action modeling does not require predicting native visual futures, and that compact predictive representations provide a strong and substantially more efficient target for policy learning. The project page is available at: https://j1dan.github.io/MiniWAM.
Figures & tables
Figure 1: MiniWAM replaces costly native future prediction with compact predictive representations. (a) While conventional WAMs predict high-dimensional native future features, MiniWAM jointly predicts actions and compact PRISM targets learned from the same visual features via inverse dynamics and reconstruction. (b) MiniWAM remains competitive with substantially larger WAMs and VLAs on LIBERO-Plus . (c) In controlled comparisons across VAE and DINOv3 feature spaces, PRISM simultaneously improves success rate and reduces training time relative to native future-feature prediction.
Figure 2: Two-stage training pipeline of PRISM and MiniWAM. In Stage 1 , privileged current–future visual features are compressed by the PRISM encoder into a compact predictive representation Lt , shaped through inverse action prediction and feature reconstruction. The corresponding world tokens remain clean while the action tokens are corrupted for conditional action prediction. In Stage 2 , the PRISM encoder is frozen and provides compact world-modeling targets; MiniWAM jointly denoises the world and action tokens using its world and action experts. Solid and hatched tokens denote clean and noised tokens, respectively. Language and proprioceptive conditioning are omitted for clarity.
Visual features
Co-training target
LIBERO
LIBERO-Plus-2K
SR (%) ↑
SR (%) ↑
VAE
w/o Co-Train
62.3
25.1
w/ Native Co-Train
92.8
50.2
w/ PRISM
94.8
56.5
DINOv3
w/o Co-Train
89.0
56.5
w/ Native Co-Train
93.2
60.9
Table 1: Controlled comparison of world-modeling targets on LIBERO and LIBERO-Plus-2K . All variants use the same 0.25B world–action architecture and training protocol within each visual feature family.
Horizon 16
Horizon 32
Visual Feature
Target
Throughput
S2
Total
Throughput
S2
Total
↑
↑
↑
↑
↑
↑
LIBERO
WAN2.1 VAE
Native
30.81
1.00 ×
1.00 ×
14.20
1.00 ×
1.00 ×
PRISM
142.04
4.61 ×
3.55 ×
124.92
8.80 ×
6.77 ×
DINOv3
Native
36.17
1.00 ×
1.00 ×
15.55
1.00 ×
1.00 ×
Table 2: Training efficiency of PRISM versus native co-training. Throughput is measured in samples/s with global batch 256 under matched hardware. S2 is Stage-2 speedup; Total estimates training speedup including Stage 1, assuming equal per-update costs for the two PRISM stages (Appendix D ).
Type
Method
Scale
Embodied PT.
Backbone Init.
LIBERO SR (%) ↑
LIBERO-Plus SR (%) ↑
RoboTwin 2.0 SR (%) ↑
VLA
OpenVLA-OFT
7B
✓
VLM
96.6
71.4
-
π0
3.3B
✓
VLM
93.5
70.5
31.4
π0 -FAST †
3.3B
✓
VLM
85.5
61.6
-
UniVLA †
7B
✓
VLM
95.2
42.9
-
StarVLA-OFT †
4B
×
VLM
96.6
-
24.8
Xiaomi-Robotics-0 †
4.7B
✓
VLM
98.7
-
40.5
Table 3: Comparison with state-of-the-art robot policies. † denotes externally reported results. Scale is the reported model size; Embodied PT and Backbone Init. denote embodied pretraining and backbone initialization. “–” means no pretrained backbone and “-” means unavailable.
Bottleneck
Recon.
Inv. dyn.
Both
4×4 , D=64
69.5
70.5
74.1
2×2 , D=64
63.3
67.5
67.8
1×1 , D=16
57.0
56.8
61.7
Table 4: PRISM ablation on LIBERO-Plus-2K . Success rate (%).
Table 6: Behavioral structure on unseen LIBERO-90 tasks. Leave-task-out retrieval results are reported as Raw → PRISM.
Appendix figures & tables4 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
Stage 1
Stage 2
Optimizer; peak LR; decay
AdamW; 5⋅10−4 ; 0
same
LR schedule; warmup
cosine; 5% of updates
same
Global batch; precision; seed
256; BF16; 42
same
LIBERO duration
15 epochs
50 epochs
RoboTwin duration
6 epochs
25 epochs
LIBERO action / visual horizon
16 / 16
same
Appendix
Table 7: Shared 0.25B training settings. Microbatch and gradient accumulation vary with GPU count while their product with GPU count remains 256.
Method
LIBERO
LIBERO-Plus
RoboTwin 2.0
π0 -FAST
85.5 ( Fei et al., 2026 )
61.6 ( Fei et al., 2026 )
–
UniVLA
95.2 ( Bu et al., 2025 )
42.9 ( Fei et al., 2026 )
–
StarVLA-OFT
96.6 ( Community, 2026 )
–
24.8 ( RoboTwin Team, 2026 )
Xiaomi-Robotics-0
98.7 ( Cai et al., 2026b )
–
40.5 ( RoboTwin Team, 2026 )
WorldVLA
79.1 ( Cen et al., 2025 ; Fei et al., 2026 )
25.0 ( Fei et al., 2026 )
–
AHA-WAM
–
–
33.8 ( RoboTwin Team, 2026 )
Appendix
Table 8: Sources of externally reported results in Table 3 . Values are success rates (%).
Benchmark
Visual
Native
PRISM
Pipeline
horizon
Core
Pipeline
Core
Pipeline
speedup
LIBERO
16
36.17
35.16
143.25
138.81
3.95×
32
15.55
15.26
126.09
123.08
8.07×
RoboTwin 2.0
16
21.48
21.11
93.57
92.86
4.40×
32
8.61
8.41
81.47
80.74
9.60×
Appendix
Table 9: DINOv3 training throughput through the data pipeline. Samples per second on one RTX PRO 6000 GPU at global batch 256. Core values reproduce the compute measurements in Table 2 ; pipeline values include data loading and transfer. Speedup compares PRISM and native Stage-2 pipeline throughput.
View
Objective
Action RMSE ↓
EEF Err. (m) ↓
Grip. Agr. ↑
Main
Reconstruction only
0.414
0.090
0.747
PRISM (I.D. + recon.)
0.356
0.068
0.802
Wrist
Reconstruction only
0.326
0.087
0.907
PRISM (I.D. + recon.)
0.266
0.078
0.975
Appendix
Table 10: Effect of inverse-dynamics supervision on neighborhood behavior. Matched DINOv3 bottlenecks on unseen LIBERO-90 tasks; both retain the full posterior-mean grid.
College of Computer Science and Artificial Intelligence, Fudan University, Shanghai, China · Singapore Management University, Singapore · Institute of Trustworthy Embodied AI, Fudan University, Shanghai, China