cs.ROAug 23, 2026

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

Authors: Chunkai Yang, Andong Yang, Di Huang, Chao Gao, Guyue Zhou

Organizations: School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China. · Department of Electronic Engineering, Tsinghua University, Beijing, China. · Institute for AI Industry Research, Tsinghua University, Beijing, China.

Abstract

Vision-language-action policies inherit both capabilities and input representations from pretrained vision-language models, showing great potential for robotic manipulation across diverse industrial settings. As these policies increasingly use interaction history, organizing the representation of historical observations determines how experiences enter temporal context and how relationships across time are modeled, which is a generally ignored challenge in previous works. In this work, we believe solving this challenge requires an architectural reconstruction and propose \textbf{WorldToken}. WorldToken encodes each timestep's observations into one world token, processes the resulting history with a causal Transformer, and generates action chunks with a diffusion action head. Unified token enables long horizon tasks while relieving the memory requirement of the temporal backbone, hence improving performance. This design also offers high interpretability and allows advances in language modeling, such as pre-training and scaling, to be transferred to robot interaction policies. In RoboCasa experiments, WorldToken successfully handles most tasks with 85M parameters and achieves 59.4% mean closed-loop success close to π0.5π_{0.5} with 3.35B parameters. On the memory benchmark RMBench Blocks Ranking, WorldToken can reach the context of two minutes and achieves a success rate of 95%. In addition, holdout action RMSE is well described by power-law fits, and its closed-loop success rate improves consistently with increasing training data size in a study involving approximately 350,000 closed-loop evaluation episodes across 50 trained policies on RoboCasa, showing its scaling potential. We also conduct experiments to analyze information preservation and history use in WorldToken, providing empirical grounding for future work.

Figures & tables

Appendix figures & tables23 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation

    Sep 30, 2026Chuyao Fu, Xiaowei Chi, Yuhan Rui +14Efficient World-Action ModelWorld Models

  2. World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

    Aug 10, 2026Qu Tang, Benhui Zhuang, Bo Yuan +3Efficient World-Action ModelWorld Models

  3. RepWAM: World Action Modeling with Representation Visual-Action Tokenizers

    Jun 11, 2026Junke Wang, Qihang Zhang, Shuai Yang +5Efficient World-Action ModelDiscrete Action Tokenizers