Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models
Organizations: Shenyang Institute of Automation, Chinese Academy of Sciences · Mohamed bin Zayed University of Artificial Intelligence · Anhui University · Xiaomi Corporation · Fudan University · University of Trento
Abstract
World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{https://github.com/JiahuaDong/AED}{AED}.
Figures & tables
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
| Training | |
|---|---|
| 1: | Set ; compute by Eq. ( 2 ), tokenize , and retrieve the temporally ordered action embeddings . |
| 2: | Compress historical visual latents to and obtain visually conditioned action embeddings ; prepend them to noisy action tokens (Eq. ( 6 )). |
| 3: | Sample and ; set for and compute (Eq. ( 1 )). |
| 4: | Sample ; predict and compute against (Eq. ( 7 )). |
| 5: | Update trainable parameters by ; keep the video VAE and target encoder frozen. |
| Inference | |
| ID | Task instruction | Successes | SR (%) |
|---|---|---|---|
| LIBERO-Spatial | |||
| S0 | pick up the black bowl between the plate and the ramekin and place it on the plate | 50/50 | 100.00 |
| S1 | pick up the black bowl next to the ramekin and place it on the plate | 50/50 | 100.00 |
| S2 | pick up the black bowl from table center and place it on the plate | 50/50 | 100.00 |
| S3 | pick up the black bowl on the cookie box and place it on the plate | 48/50 | 96.00 |
| S4 | pick up the black bowl in the top drawer of the wooden cabinet and place it on the plate | 49/50 | 98.00 |
| Fast-WAM | Ours | |||||
|---|---|---|---|---|---|---|
| Task | Clean | Random | Avg. | Clean | Random | Avg. |
| adjust_bottle | 100.00 | 100.00 | 100.00 | 100.00 | 99.00 | 99.50 |
| beat_block_hammer | 99.00 | 97.00 | 98.00 | 98.00 | 99.00 | 98.50 |
| blocks_ranking_rgb | 100.00 | 100.00 | 100.00 | 100.00 | 99.00 | 99.50 |
| blocks_ranking_size | 94.00 | 98.00 | 96.00 | 92.00 | 96.00 | 94.00 |
| click_alarmclock | 100.00 | 100.00 | 100.00 | 100.00 | 99.00 | 99.50 |