Organizations: Shenyang Institute of Automation, Chinese Academy of Sciences · Mohamed bin Zayed University of Artificial Intelligence · Anhui University · Xiaomi Corporation · Fudan University · University of Trento
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{https://github.com/JiahuaDong/AED}{AED}.
Figures & tables
Figure 1: (a) Example of reusable action skills. (b) Visualization of reusable action experience in the action experience dictionary (AED). f⋆[0]–f⋆[2] indicate the transporting action pattern, while f⋆[3]–f⋆[7] encode the dipping action pattern in the AED. The visualization of the cosine similarities among action embeddings in the AED shows that learning to place the bowl on the plate involves substantial reuse of both the transporting and dipping action patterns. (c) Gray and green lines show gripper height and its relative changes, respectively, during training to place the bowl on the plate, linking action patterns (e.g., transporting and dipping ) to physical height.
Figure 2: Overview of the proposed AED. After defining learnable action experience dictionary (AED) shared across tasks, we utilize the visually conditioned action embeddings to encode the task-relevant information and employ a motion-aware transition loss to encode action-relevant motion.
Figure 3: Visualization of manipulation tasks performed by our model in OOD settings.
Figure 4: Results on real-world manipulation tasks across robotic embodiments under OOD settings.
Table 5
Figure 5: Analysis of reusing action experience embeddings across different manipulation tasks.
Figure 6: (a) Quantitative performance evaluation of the proposed AED on LIBERO Goal. (b) Visualization of direct and composed transition predictions under supervision from the MT loss.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
▹ Training
1:
Set k=H/V ; compute ah[j] by Eq. ( 2 ), tokenize ζj=Φ(ah[j]) , and retrieve the temporally ordered action embeddings f⋆=[pj+M−1∑lD(ζj[l])]j=1V .
2:
Compress historical visual latents zvh to zvh and obtain visually conditioned action embeddings ev=E(f⋆,zvh) ; prepend them to noisy action tokens (Eq. ( 6 )).
3:
Sample τ and ϵe ; set xeτ=(1−τ)xe+τϵe for e∈{v,a} and compute LFM (Eq. ( 1 )).
4:
Sample (t′,Δ) ; predict Δht′=G(ht′,ut′Δ) and compute LMT against Δht′=ht′+Δ−ht′ (Eq. ( 7 )).
5:
Update trainable parameters by L=LFM+0.01LMT ; keep the video VAE and target encoder F frozen.
▹ Inference
Appendix
Table 8
ID
Task instruction
Successes
SR (%)
LIBERO-Spatial
S0
pick up the black bowl between the plate and the ramekin and place it on the plate
50/50
100.00
S1
pick up the black bowl next to the ramekin and place it on the plate
50/50
100.00
S2
pick up the black bowl from table center and place it on the plate
50/50
100.00
S3
pick up the black bowl on the cookie box and place it on the plate
48/50
96.00
S4
pick up the black bowl in the top drawer of the wooden cabinet and place it on the plate
49/50
98.00
Appendix
Table 6: Per-task success rates on all 40 standard LIBERO tasks, grouped by suite. The same 42,000-step checkpoint is evaluated with seed 3407, 50 trials per task, and replanning after ten executed actions. SR denotes success rate in percent. Each suite average covers 500 trials; the overall average covers 2,000 trials.
Fast-WAM
Ours
Task
Clean
Random
Avg.
Clean
Random
Avg.
adjust_bottle
100.00
100.00
100.00
100.00
99.00
99.50
beat_block_hammer
99.00
97.00
98.00
98.00
99.00
98.50
blocks_ranking_rgb
100.00
100.00
100.00
100.00
99.00
99.50
blocks_ranking_size
94.00
98.00
96.00
92.00
96.00
94.00
click_alarmclock
100.00
100.00
100.00
100.00
99.00
99.50
Appendix
Table 7: Per-task success rate (%) on RoboTwin 2.0. Fast-WAM entries are reported baseline results; ours use 100 trials per task and setting with unseen instructions. Avg. averages Clean and Random, and the final row averages all 50 tasks. Bold indicates the better result between the two methods for each metric, including ties.
Figure 7: Overview of ten real-world manipulation tasks on Spirit AI MOZ1. Each row shows six frames from the high camera, covering an episode from start to finish.
Figure 8: Robot platforms and deployment pipeline. Left and center: Spirit AI MOZ1 and ROKAE AR5-5_0.7 with camera and gripper annotations. Right: camera images and language instructions are processed by the model on a host computer, and predicted actions are sent to the robot.
Figure 9: Cross-task similarity of action content embeddings. Orange and yellow compare the same skill and different skills across tasks, respectively. Each distribution sample is a weighted average cosine-similarity score for a cross-task pair; white points mark the medians of these task-pair scores. The annotated gaps for grasp, transport, dip, release, lift, and reach are +0.173 , +0.105 , +0.092 , +0.163 , +0.090 , and +0.067 , respectively.
Figure 10: Additional transition-prediction examples. Each row shows, from left to right, the observation at t3 , a PCA visualization of its ground-truth features, the endpoint features estimated by direct t1→t3 prediction, and those estimated by composing t1→t2 and t2→t3 predictions. The two prediction routes produce similar spatial feature patterns across the illustrated scenes.
Figure 11: Visual context associated with historical action embeddings. Panels are ordered from left to right and top to bottom, corresponding to f⋆[0] – f⋆[7] . Each panel pairs two camera views with overlaid visual responses and trajectory markers. The responses vary with the interaction stage and include regions around the robot gripper and nearby objects.
Figure 16
Figure 14: Temporal-window sampling probabilities for Start → Span , with V=8 . Left: marginal length distribution P(Δ) ; the upper axis counts saved action records (four per visual interval). Right: joint distribution P(r,Δ) over start offsets and lengths; blank cells are invalid windows. All values are percentages.
This work presents RepWAM, a representation-centric world action model (WAM) built on representation visual-action tokenizers. Existing WAMs typically inherit reconstruction-oriented video tokenizers from pretrained video generation models. Although these tokenizers preserve visual fidelity, pixel reconstruction alone provides limited guidance for learning instruction-following dynamics that connect future prediction with robot control. To address this, we explore a semantic visual-action latent space for representation-centric world action modeling. Specifically, we train a representation visual-action tokenizer that maps visual inputs into aligned visual and latent action tokens. We then pretrain our WAM to jointly model future visual states and the latent actions that connect them under language instructions, followed by adaptation to real robot trajectories for closed-loop manipulation. Experiments on real-world manipulation tasks and simulation benchmarks show that RepWAM delivers strong performance across diverse manipulation settings, while ablations highlight the value of semantic visual-action tokenization over reconstruction-oriented alternatives. These results establish representation visual-action tokenization as a promising foundation for world action models and a step toward generalist robot policies. Code and weights will be available at https://github.com/wdrink/RepWAM.
Junke Wang, Qihang Zhang, Shuai Yang +5
Institute of Trustworthy Embodied AI, Fudan University · 2Robbyant, Ant Group · 3Hongkong University of Science and Technology
World-action models have shown promising robot-manipulation performance by jointly predicting future visual states and actions. However, existing methods mainly rely on short-term history and short-horizon future prediction, which is insufficient for long-horizon tasks whose correct execution depends on earlier observations and task progress. Such temporally dependent tasks require effective use of complementary temporal information, including recent local context, cross-stage historical events, immediate future dynamics, and global task progress. To address long-term forgetting and poor awareness of the global task state, we introduce DiM-WAM, a memory-augmented world-action model that integrates multi-scale historical context, local future dynamics, and global task progress. The memory extracts compact visual event information from real observations, updates multiple memory banks through independent similarity-based merging, and then reads the bank-identity- and time-embedded long-term context to condition video and action denoising. A progress-supervision objective further encourages memory tokens to encode not only completed historical events but also the current task stage and its implications for the remaining task. On RMBench, DiM-WAM raises average success from 28.4% with LingBot-VA to 69.8%, exceeding the explicit-memory Mem-0 baseline at 42.0%. On four real-world Franka tasks, it improves average stage success from 70.7% to 91.5% and full-task success from 52.5% to 80.0%. Project page: https://wangkai-casia.github.io/dim-wam.
Predicting the next state is a core paradigm of World Models for modeling physical dynamics, emphasizing prediction fidelity. As World Models evolve into World-Action Models (WAMs), existing methods still fix the next state before training as RGB, a single latent feature, or a static combination of predefined targets, thereby constraining action learning to the inductive biases preserved by a particular representation. To address this limitation, we propose CF-WAM, a dynamic next-state prediction framework that samples visual, semantic, geometric, and interaction projections of the same future, standardizes them into a common video form, and supervises a unified WAM across these projections. The action-relevant constraints exposed by these projections accumulate across training steps, forcing WAM to capture the underlying state-transition structure that supports multiple projections of the same action-conditioned future. This dynamic mechanism also provides a natural cross-embodiment dynamics reference frame for Human and Robot learning. By jointly learning across different next-state parameterizations, heterogeneous Human and Robot experience can bypass appearance differences and directly contribute to shared state-transition learning, improving cross-embodiment generalization. Experiments show that CF-WAM improves both training efficiency and final control performance, while translating Human experience effectively into policy gains. CF-WAM achieves state-of-the-art performance on RoboCasa-GR1 with an average success rate of 82.50%, while reaching 82.65% on LIBERO-Plus and up to 84.00% in real-world evaluations.