Token-World: World Modeling in Vision-Language Model Token Space for Robot Manipulation
Authors: Chuyao Fu, Xiaowei Chi, Yuhan Rui, Yu-kai Wang, Zezhong Qian, Xiaojie Zhang, Yunfan Lou, Kevin Zhang, +9 more
Organizations: Southern University of Science and Technology · MUKA Robotics · Hong Kong University of Science and Technology · State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University · University of Pennsylvania · Institute of Automation, Chinese Academy of Sciences
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World (r=0.794 vs.\ 0.583), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.
Figures & tables
Fig. 2 : Token-World overview. (a) Compact Token Construction. Qwen3-VL visual tokens are compressed into a compact latent space. (b) Dynamics Modeling. A spatiotemporal Transformer models action-conditioned dynamics in compact token space. (c) Training Objective. We train the model using an x0 -parameterized flow objective with weighted shortcut forcing. (d) Autoregressive Rollout. Future compact tokens are generated with a sliding temporal window and recurrent state.
Fig. 3 : Real-world evaluation setup. We collect manipulation trajectories using a Franka Research 3 arm equipped with a Robotiq adaptive gripper and two Intel RealSense 435 cameras. The real-world benchmark contains six manipulation tasks (hang on M, hang on cup, stack jenga, stack ring, put jenga in drawer, put chili in drawer) with 100 trajectories per task.
Method
Venue
RoboTwin (50 Tasks)
Real Franka (6 Tasks)
VLM Feature Fidelity
Policy-Action Consistency
VLM Feature Fidelity
Policy-Action Consistency
Cosine ↑
NMSE ↓
Cosine ↑
NMSE ↓
Cosine ↑
NMSE ↓
Cosine ↑
NMSE ↓
IRASim
ICCV’25
0.7372
0.5685
0.9260
0.1533
0.7061
0.6095
0.9174
0.1702
Ctrl-World
ICLR’26
0.7097
0.6437
0.9265
0.1463
0.7213
0.5719
0.9288
0.1410
WorldGym
ICLR’26
0.7146
0.6157
0.9251
0.1602
0.6697
0.6930
0.9089
0.1981
Token-World
Ours
0.7714
0.4892
0.9439
0.1145
0.7506
0.5129
0.9402
0.1208
TABLE I : Open-loop world-model fidelity on simulated and real-world manipulation. Feature metrics compare predicted and ground-truth frozen VLM representations. Policy-action metrics compare the outputs of the same frozen policy conditioned on predicted and ground-truth representations. For RGB-generative baselines, generated frames are re-encoded by the frozen VLM before evaluation. Higher cosine similarity and lower NMSE are better.
Fig. 4 : Qualitative comparison of long-horizon open-loop rollout. We visualize seven states along the same action-replay trajectory. The top row shows the reference RGB observations, followed by predictions from Ctrl-World and Token-World. Token-World more closely preserves the task-relevant object configuration and interaction progression over long horizons. Token-World predictions are rendered to RGB using the auxiliary decoder only for visualization.
Fig. 5 : Long-horizon open-loop fidelity under chunk-based autoregressive rollout. Results are reported at replay chunks 0, 5, 10, and 15, with 16 steps a chunk. Each point denotes the chunk-wise average VLM-feature and action similarity. Token-World degrades more slowly than Ctrl-World over long-horizon rollout.
Fig. 6 : Policy evaluation fidelity and simulation efficiency. (a) Correlation between policy success rates measured in learned world models and the corresponding reference environments. Each point represents one policy–task pair. The gray dotted line indicates the oracle relation y=x , and the black dashed line denotes linear regression. (b) Average per-step simulation time of different world-model simulators. Lower is better.
d
NMSE ↓
Cos. ↑
SEC ↓
LDS ↑
8
0.1026
0.9389
0.3511
0.3100
16
0.0791
0.9528
0.3854
0.2534
24
0.0694
0.9587
0.3960
0.2383
32
0.0633
0.9623
0.4058
0.2274
48
0.0562
0.9668
0.4269
0.1969
96
0.0445
0.9739
0.4394
0.1595
TABLE II : Reconstruction and modelability of S-VAE representations across compact-state dimensions. Increasing the latent dimension improves reconstruction fidelity, whereas the modelability metrics exhibit the opposite trend.
Representation
Feat. NMSE ↓
Feat. Cos. ↑
Act. NMSE ↓
Act. Cos. ↑
S-VAE-8
0.1549
0.8254
0.0259
0.9603
S-VAE-16
0.1517
0.8306
0.0247
0.9641
S-VAE-32
0.1634
0.8182
0.0275
0.9596
S-VAE-48
0.1698
0.8109
0.0236
0.9638
TABLE III : Effect of S-VAE dimensionality on downstream dynamics prediction. The dynamics backbone, training data, and optimization protocol are fixed; only the compact-state dimension is varied.
Representation
Feat. NMSE ↓
Feat. Cos. ↑
Act. NMSE ↓
Act. Cos. ↑
Raw VLM
0.7851
0.6375
0.2180
0.8948
SDXL-VAE
0.8164
0.621
0.2586
0.875
S-VAE
0.4892
0.7714
0.1145
0.9439
TABLE IV : Controlled ablation of world-state representations. All variants use the same spatiotemporal dynamics backbone and training protocol; only the state representation and its corresponding input/output projections are changed.
School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China. · Department of Electronic Engineering, Tsinghua University, Beijing, China. · Institute for AI Industry Research, Tsinghua University, Beijing, China.