cs.ROSep 30, 2026

Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation

Authors: Chuyao Fu, Xiaowei Chi, Yuhan Rui, Yu-kai Wang, Zezhong Qian, Xiaojie Zhang, Yunfan Lou, Kevin Zhang, +9 more

Organizations: Southern University of Science and Technology · MUKA Robotics · Hong Kong University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · University of Pennsylvania · Institute of Automation, Chinese Academy of Sciences

Abstract

A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World (r=0.794r=0.794 vs.\ 0.5830.583), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.

Figures & tables

Explore similar work

CardsList
  1. World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

    Aug 10, 2026Qu Tang, Benhui Zhuang, Bo Yuan +3Efficient World-Action ModelWorld Models

  2. WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

    Aug 23, 2026Chunkai Yang, Andong Yang, Di Huang +2Discrete Action TokenizersEfficient World-Action Model

  3. Think Like a World Model, Act Like a VLA: Distilling World-Model Representations into Compact Robot Policies

    Sep 21, 2026Trung Dao, Sankalp Yamsani, Jaden Park +2World ModelsRobot Systems