cs.CVSep 30, 2026

Learning Skills from Historical Action Trajectories: Action Experience Dictionary for World Action Models

Authors: Qi Lyu, Jiahua Dong, Hao Shen, Xudong Wang, Hongyuan Yu, Baichen Liu, Henghui Ding, Zhi Han, +4 more

Organizations: Shenyang Institute of Automation, Chinese Academy of Sciences · Mohamed bin Zayed University of Artificial Intelligence · Anhui University · Xiaomi Corporation · Fudan University · University of Trento

Abstract

World Action Models (WAMs) couple visual dynamics prediction with action generation, yet they do not explicitly support the reuse of action experience across manipulation tasks. Furthermore, existing WAMs struggle to capture underlying cross-task semantic relationships that could guide target action prediction, as redundant background elements interfere with the extraction of key visual information. To address these challenges, we develop a novel Action Experience Dictionary (AED) that encodes historical physical action trajectories into shared action embeddings to support skill reuse and model cross-task relationships. Specifically, we first aggregate historical actions to align with visual observations and retrieve action embeddings from the AED using a pretrained action tokenizer. Subsequently, we visually condition the pooled embeddings through cross-attention and prepend them to noisy action tokens, providing interaction context and action intent for prediction. To model action-related motion and reduce reliance on irrelevant background cues, we introduce a motion-aware transition loss that supervises visual feature change prediction over random temporal intervals. Experiments on simulation benchmarks and in real-world cross-embodiment settings verify the effectiveness of our AED. The anonymous project website is available at \href{https://github.com/JiahuaDong/AED}{AED}.

Figures & tables

Appendix figures & tables12 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

    Jun 11, 2026Junke Wang, Qihang Zhang, Shuai Yang +5Efficient World-Action ModelDiscrete Action Tokenizers

  2. From World Models to World Action Models: Rethinking Next-State Prediction

    Sep 28, 2026Tingyu Yuan, Ziming Ji, Biaoliang Guan +9Efficient World-Action ModelWorld Models