Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation
Authors: Chuyao Fu, Xiaowei Chi, Yuhan Rui, Yu-kai Wang, Zezhong Qian, Xiaojie Zhang, Yunfan Lou, Kevin Zhang, +9 more
Organizations: Southern University of Science and Technology · MUKA Robotics · Hong Kong University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · University of Pennsylvania · Institute of Automation, Chinese Academy of Sciences
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World (r=0.794 vs.\ 0.583), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.
Figures & tables
Fig. 2 : Token-World overview. (a) Compact Token Construction. Qwen3-VL visual tokens are compressed into a compact latent space. (b) Dynamics Modeling. A spatiotemporal Transformer models action-conditioned dynamics in compact token space. (c) Training Objective. We train the model using an x0 -parameterized flow objective with weighted shortcut forcing. (d) Autoregressive Rollout. Future compact tokens are generated with a sliding temporal window and recurrent state.
Fig. 3 : Real-world evaluation setup. We collect manipulation trajectories using a Franka Research 3 arm equipped with a Robotiq adaptive gripper and two Intel RealSense 435 cameras. The real-world benchmark contains six manipulation tasks (hang on M, hang on cup, stack jenga, stack ring, put jenga in drawer, put chili in drawer) with 100 trajectories per task.
Method
Venue
RoboTwin (50 Tasks)
Real Franka (6 Tasks)
VLM Feature Fidelity
Policy-Action Consistency
VLM Feature Fidelity
Policy-Action Consistency
Cosine ↑
NMSE ↓
Cosine ↑
NMSE ↓
Cosine ↑
NMSE ↓
Cosine ↑
NMSE ↓
IRASim
ICCV’25
0.7372
0.5685
0.9260
0.1533
0.7061
0.6095
0.9174
0.1702
Ctrl-World
ICLR’26
0.7097
0.6437
0.9265
0.1463
0.7213
0.5719
0.9288
0.1410
WorldGym
ICLR’26
0.7146
0.6157
0.9251
0.1602
0.6697
0.6930
0.9089
0.1981
Token-World
Ours
0.7714
0.4892
0.9439
0.1145
0.7506
0.5129
0.9402
0.1208
TABLE I : Open-loop world-model fidelity on simulated and real-world manipulation. Feature metrics compare predicted and ground-truth frozen VLM representations. Policy-action metrics compare the outputs of the same frozen policy conditioned on predicted and ground-truth representations. For RGB-generative baselines, generated frames are re-encoded by the frozen VLM before evaluation. Higher cosine similarity and lower NMSE are better.
Fig. 4 : Qualitative comparison of long-horizon open-loop rollout. We visualize seven states along the same action-replay trajectory. The top row shows the reference RGB observations, followed by predictions from Ctrl-World and Token-World. Token-World more closely preserves the task-relevant object configuration and interaction progression over long horizons. Token-World predictions are rendered to RGB using the auxiliary decoder only for visualization.
Fig. 5 : Long-horizon open-loop fidelity under chunk-based autoregressive rollout. Results are reported at replay chunks 0, 5, 10, and 15, with 16 steps a chunk. Each point denotes the chunk-wise average VLM-feature and action similarity. Token-World degrades more slowly than Ctrl-World over long-horizon rollout.
Fig. 6 : Policy evaluation fidelity and simulation efficiency. (a) Correlation between policy success rates measured in learned world models and the corresponding reference environments. Each point represents one policy–task pair. The gray dotted line indicates the oracle relation y=x , and the black dashed line denotes linear regression. (b) Average per-step simulation time of different world-model simulators. Lower is better.
d
NMSE ↓
Cos. ↑
SEC ↓
LDS ↑
8
0.1026
0.9389
0.3511
0.3100
16
0.0791
0.9528
0.3854
0.2534
24
0.0694
0.9587
0.3960
0.2383
32
0.0633
0.9623
0.4058
0.2274
48
0.0562
0.9668
0.4269
0.1969
96
0.0445
0.9739
0.4394
0.1595
TABLE II : Reconstruction and modelability of S-VAE representations across compact-state dimensions. Increasing the latent dimension improves reconstruction fidelity, whereas the modelability metrics exhibit the opposite trend.
Representation
Feat. NMSE ↓
Feat. Cos. ↑
Act. NMSE ↓
Act. Cos. ↑
S-VAE-8
0.1549
0.8254
0.0259
0.9603
S-VAE-16
0.1517
0.8306
0.0247
0.9641
S-VAE-32
0.1634
0.8182
0.0275
0.9596
S-VAE-48
0.1698
0.8109
0.0236
0.9638
TABLE III : Effect of S-VAE dimensionality on downstream dynamics prediction. The dynamics backbone, training data, and optimization protocol are fixed; only the compact-state dimension is varied.
Representation
Feat. NMSE ↓
Feat. Cos. ↑
Act. NMSE ↓
Act. Cos. ↑
Raw VLM
0.7851
0.6375
0.2180
0.8948
SDXL-VAE
0.8164
0.621
0.2586
0.875
S-VAE
0.4892
0.7714
0.1145
0.9439
TABLE IV : Controlled ablation of world-state representations. All variants use the same spatiotemporal dynamics backbone and training protocol; only the state representation and its corresponding input/output projections are changed.
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation. It uses world modeling during training to enhance the action policy while preserving efficient deployment. Specifically, the World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and simultaneously serve as the action expert's sole visual-language context. This shared conditioning allows gradients from future-video denoising to directly shape the representation used for action prediction, while exclusive routing prevents the policy from bypassing that representation. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert, with no online video-model inference. With a 2B backbone and no embodied action pretraining, World Tokens is highly competitive on LIBERO, attains the best reported averages on SIMPLER, substantially improves real-world R1 Pro success over a matched action-only baseline, and generates each action chunk at VLA-level latency.
Vision-language-action policies inherit both capabilities and input representations from pretrained vision-language models, showing great potential for robotic manipulation across diverse industrial settings. As these policies increasingly use interaction history, organizing the representation of historical observations determines how experiences enter temporal context and how relationships across time are modeled, which is a generally ignored challenge in previous works. In this work, we believe solving this challenge requires an architectural reconstruction and propose \textbf{WorldToken}. WorldToken encodes each timestep's observations into one world token, processes the resulting history with a causal Transformer, and generates action chunks with a diffusion action head. Unified token enables long horizon tasks while relieving the memory requirement of the temporal backbone, hence improving performance. This design also offers high interpretability and allows advances in language modeling, such as pre-training and scaling, to be transferred to robot interaction policies. In RoboCasa experiments, WorldToken successfully handles most tasks with 85M parameters and achieves 59.4% mean closed-loop success close to π0.5 with 3.35B parameters. On the memory benchmark RMBench Blocks Ranking, WorldToken can reach the context of two minutes and achieves a success rate of 95%. In addition, holdout action RMSE is well described by power-law fits, and its closed-loop success rate improves consistently with increasing training data size in a study involving approximately 350,000 closed-loop evaluation episodes across 50 trained policies on RoboCasa, showing its scaling potential. We also conduct experiments to analyze information preservation and history use in WorldToken, providing empirical grounding for future work.
Chunkai Yang, Andong Yang, Di Huang +2
School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China. · Department of Electronic Engineering, Tsinghua University, Beijing, China. · Institute for AI Industry Research, Tsinghua University, Beijing, China.
Vision-Language-Action (VLA) models map observations to actions with no objective that accounts for how the world responds, so their robustness is bounded primarily by data coverage. World models carry precisely that missing objective and are better grounded for it, yet rolling the future forward costs seconds per decision and rules them out of the control loop. We show the two can be separated. What a world model knows about physical scenes lives in its internal features; generating the future is merely the objective that produced them, so the grounding can be inherited while the generative machinery is left behind. We add one feature-alignment term to ordinary VLA training: a frozen world model is run over the training frames once and cached, and the student learns to agree with that cache. No teacher is loaded during training, the projector is discarded after it, and the deployed policy is identical to the undistilled baseline, running in 32ms and 1.86GB on a consumer RTX5090, so every gain is attributable to the representation rather than to added capacity or test-time compute. A 0.8B student reaches 97.9% on LIBERO, improves from 48.2% to 50.5% on RoboCasa-GR1 humanoid manipulation, and the same objective carries over to real hardware, on both a single-arm and a bimanual platform. The gain survives changes of student scale, backbone, alignment layer, and teacher, indicating a broad representational prior rather than a fragile alignment between two particular networks. Project page: https://thaw-vla.trung-dt.com/.
Trung Dao, Sankalp Yamsani, Jaden Park +2
University of Wisconsin–Madison, USA · University of Illinois Urbana-Champaign, USA