WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning
Organizations: School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China. · Department of Electronic Engineering, Tsinghua University, Beijing, China. · Institute for AI Industry Research, Tsinghua University, Beijing, China.
Abstract
Vision-language-action policies inherit both capabilities and input representations from pretrained vision-language models, showing great potential for robotic manipulation across diverse industrial settings. As these policies increasingly use interaction history, organizing the representation of historical observations determines how experiences enter temporal context and how relationships across time are modeled, which is a generally ignored challenge in previous works. In this work, we believe solving this challenge requires an architectural reconstruction and propose \textbf{WorldToken}. WorldToken encodes each timestep's observations into one world token, processes the resulting history with a causal Transformer, and generates action chunks with a diffusion action head. Unified token enables long horizon tasks while relieving the memory requirement of the temporal backbone, hence improving performance. This design also offers high interpretability and allows advances in language modeling, such as pre-training and scaling, to be transferred to robot interaction policies. In RoboCasa experiments, WorldToken successfully handles most tasks with 85M parameters and achieves 59.4% mean closed-loop success close to with 3.35B parameters. On the memory benchmark RMBench Blocks Ranking, WorldToken can reach the context of two minutes and achieves a success rate of 95%. In addition, holdout action RMSE is well described by power-law fits, and its closed-loop success rate improves consistently with increasing training data size in a study involving approximately 350,000 closed-loop evaluation episodes across 50 trained policies on RoboCasa, showing its scaling potential. We also conduct experiments to analyze information preservation and history use in WorldToken, providing empirical grounding for future work.
Figures & tables
| Training seed 0 | Training seed 1 | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| D50 | D100 | D300 | D1000 | D2900 | D50 | D100 | D300 | D1000 | D2900 | ||
| SR (%) | N1 | 14.1 | 30.0 | 41.0 | 49.1 | 54.4 | 18.1 | 27.4 | 41.4 | 49.9 | 54.7 |
| N2 | 21.4 | 31.2 | 47.5 | 55.7 | 59.1 | 18.9 | 32.4 | 46.2 | 52.2 | 59.8 | |
| N3 | 22.0 | 35.8 | 48.1 | 58.5 | 59.3 | 22.7 | 33.2 | 48.4 | 57.8 | 61.2 | |
| N4 | 27.1 | 36.5 | 51.2 | 57.7 | 58.8 | 24.0 | 36.3 | 51.4 | 56.7 | 57.8 | |
| N5 | 23.1 | 33.7 | 49.8 | 56.2 | 59.7 | 21.7 | 34.6 | 49.9 | 56.4 | 59.3 | |
| SR (%) | RMSE | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Seed | Params | RelPs | D50 | D100 | D300 | D1000 | D2900 | D50 | D100 | D300 | D1000 | D2900 | |
| 1 | 0 | 85.3M | 21.4 | 31.2 | 47.5 | 55.7 | 59.1 | 0.195 | 0.172 | 0.137 | 0.109 | 0.091 | |
| 1 | 1 | 85.3M | 18.9 | 32.4 | 46.2 | 52.2 | 59.8 | 0.197 | 0.169 | 0.135 | 0.110 | 0.091 | |
| 4 | 0 | 92.4M | 20.3 | 38.1 | 48.8 | 55.6 | 59.9 | 0.202 | 0.166 | 0.138 | 0.110 | 0.090 | |
| 50 | 0 | 83.0M | 26.5 | 37.6 | 49.6 | 57.7 | 60.7 | 0.191 | 0.165 | 0.132 | 0.107 | 0.089 | |
| D50 | D100 | D300 | D1000 | D2900 | D50 | D100 | D300 | D1000 | D2900 | D50 | D100 | D300 | D1000 | D2900 | ||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Seed 0 | N1 | |||||||||||||||
| N2 | ||||||||||||||||
| N3 | ||||||||||||||||
| N4 | ||||||||||||||||
| N5 | ||||||||||||||||
| Context | RMSE, seed 0 | RMSE, seed 1 | SR (%) , seed 0 | SR (%) , seed 1 |
|---|---|---|---|---|
| 1 | 0.150260 | 0.147767 | 46.75 1.02 | 47.68 1.50 |
| 2 | 0.149383 | 0.147528 | 46.87 0.84 | 48.52 1.36 |
| 5 | 0.137165 | 0.139603 | 51.48 0.26 | 50.00 0.92 |
| 10 | 0.131710 | 0.133098 | 48.12% 1.70 pp | 48.38 1.09 |
| (episodes) | 1 (24) | 2 (14) | 3 (28) | 4 (18) | 5 (16) | Total (100) |
|---|---|---|---|---|---|---|
| 608 | 24/24/0 | 14/14/0 | 25/25/0 | 17/16/1 | 15/15/0 | 95/94/1 |
| 288 | 24/24/0 | 14/14/0 | 23/23/0 | 16/16/0 | 15/15/0 | 92/92/0 |
| 128 | 24/24/0 | 14/14/0 | 13/3/10 | 5/5/0 | 3/3/0 | 59/49/10 |
| 64 | 24/24/0 | 3/3/0 | 2/0/2 | 9/0/9 | 0/0/0 | 38/27/11 |
| 32 | 24/24/0 | 3/0/3 | 0/0/0 | 1/0/1 | 0/0/0 | 28/24/4 |
| Env. seed | Swaps | Last (s) | Env. seed | Swaps | Last (s) | Env. seed | Swaps | Last (s) | |||
|---|---|---|---|---|---|---|---|---|---|---|---|
| 100000 | 2 | 13 | 350.52 | 100004 | 3 | 8 | 222.12 | 100015 | 5 | 5 | 139.38 |
| 100001 | 3 | 5 | 139.20 | 100003 | 2 | 9 | 250.14 | 100016 | 1 | 5 | 139.32 |
| 100002 | 4 | 17 | 470.28 | 100007 | 1 | 5 | 139.08 | 100008 | 1 | 31 | 856.44 |
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
| Setting | N1 | N2 | N3 | N4 | N5 |
|---|---|---|---|---|---|
| Parameters | 44.3M | 85.3M | 218.8M | 648.9M | 1490.3M |
| Encoder/temporal LR ( ) | 6.00 | 4.25 | 3.00 | 1.50 | 0.50 |
| Action-decoder LR ( ) | 3.00 | 3.00 | 3.00 | 3.00 | 3.00 |
| Within-timestep fusion encoder (SwiGLU) | |||||
| Width | 512 | 768 | 1024 | 1536 | 2048 |
| Layers | 1 | 2 | 4 | 6 | 8 |
| Component | Protocol |
|---|---|
| Inputs | Three RGB views, 16-D proprioception, and a frozen 768-D CLIP task embedding |
| Visual stems | One five-stage CNN per camera; convolutions with channels , each followed by max pooling, group normalization, and SiLU; output grid per view |
| Token readout | 48 visual tokens plus proprioception and task tokens; four learned readouts are concatenated, RMS-normalized, and projected to one temporal world token |
| Temporal sampling | 20 Hz environment control; one observation, world token, and replanning decision every four control steps (5 Hz) |
| Actions | Same-frame-aligned 12-D commands; predicted, executed before replanning |
| Context and diffusion | world tokens; 20-stage cosine schedule and stochastic 20-step DDPM evaluation |
| Experiment | Changed / held fixed | Training seeds | Executions |
|---|---|---|---|
| Data and model scaling | Five values five capacities; | 0, 1 | 3 |
| Token count | at all five ; , | 0 | 3 |
| Same-checkpoint history | for every scaling checkpoint; | 0, 1 | 1 |
| Matched training history | ; 218.8M, | 0, 1 | 3 |
| Single execution | repeats | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Model | Seed | RMSE | 1 | 2 | 3 | Mean SD | ||||
| 50 | N1 | 0 | 0.20765 | 10.61 | 10.00 | 13.22 | 14.52 | 13.91 | 14.00 | |
| 1 | 0.20301 | 12.52 | 15.57 | 18.78 | 18.17 | 18.96 | 17.30 | |||
| N2 | 0 | 0.19514 | 14.26 | 16.09 | 21.65 | 20.96 | 20.96 | 22.26 | ||
| 1 | 0.19682 | 10.96 | 13.57 | 18.09 | 20.35 | 17.65 | 18.78 | |||
| N3 | 0 | 0.19532 | 14.70 | 17.57 | 22.26 | 22.70 | 22.00 | 21.22 | ||
| Params | 10% | 20% | 30% | 40% | 50% | 60% | 70% | 80% | 90% | 100% | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 50 | 44.3M | 0.41821 | 0.31303 | 0.27148 | 0.25973 | 0.23128 | 0.21949 | 0.21252 | 0.21222 | 0.20983 | 0.20765 |
| 85.3M | 0.41246 | 0.35694 | 0.24800 | 0.22547 | 0.21357 | 0.20352 | 0.20087 | 0.19598 | 0.19501 | 0.19514 | |
| 218.8M | 0.40335 | 0.34023 | 0.27017 | 0.23318 | 0.21947 | 0.20875 | 0.19826 | 0.19802 | 0.19648 | 0.19532 | |
| 648.9M | 0.33931 | 0.23841 | 0.22386 | 0.20937 | 0.20262 | 0.19702 | 0.19359 | 0.19386 | 0.19313 | 0.19279 | |
| 1.49B | 0.27697 | 0.21628 | 0.20776 | 0.19756 | 0.19285 | 0.19263 | 0.19124 | 0.19109 | 0.19044 | 0.18934 | |
| 100 | 44.3M | 0.39791 | 0.27210 | 0.22569 | 0.20335 | 0.19464 | 0.18634 | 0.18098 | 0.17686 | 0.17595 | 0.17477 |
| Params | 10% | 20% | 30% | 40% | 50% | 60% | 70% | 80% | 90% | 100% | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 50 | 44.3M | 0.41265 | 0.36703 | 0.30551 | 0.24915 | 0.24223 | 0.22526 | 0.21346 | 0.20761 | 0.20684 | 0.20301 |
| 85.3M | 0.41460 | 0.31348 | 0.25889 | 0.23951 | 0.22562 | 0.21208 | 0.20414 | 0.19972 | 0.19819 | 0.19682 | |
| 218.8M | 0.39254 | 0.28892 | 0.23680 | 0.21865 | 0.20694 | 0.20339 | 0.19930 | 0.19517 | 0.19357 | 0.19411 | |
| 648.9M | 0.37310 | 0.25907 | 0.22671 | 0.21701 | 0.20702 | 0.19907 | 0.20138 | 0.19804 | 0.19506 | 0.19518 | |
| 1.49B | 0.26374 | 0.21731 | 0.20156 | 0.19317 | 0.19170 | 0.18972 | 0.18823 | 0.18731 | 0.18720 | 0.18610 | |
| 100 | 44.3M | 0.28741 | 0.23559 | 0.20988 | 0.19770 | 0.18795 | 0.18145 | 0.17921 | 0.17774 | 0.17699 | 0.17701 |
| Training seed 0: | Training seed 1: | Selected run | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| Task | 50 | 100 | 300 | 1000 | 2900 | 50 | 100 | 300 | 1000 | 2900 | Successes |
| CloseDoubleDoor | 23.33 | 67.33 | 88.00 | 91.33 | 94.67 | 36.67 | 68.67 | 92.67 | 90.67 | 91.33 | 45/50 |
| CloseDrawer | 88.00 | 98.67 | 100.00 | 100.00 | 98.67 | 87.33 | 96.00 | 99.33 | 100.00 | 97.33 | 48/50 |
| CloseSingleDoor | 57.33 | 67.33 | 82.67 | 83.33 | 87.33 | 50.67 | 76.00 | 76.00 | 84.00 | 84.00 | 45/50 |
| CoffeePressButton | 44.67 | 56.67 | 78.00 | 91.33 | 94.67 | 38.00 | 51.33 | 84.67 | 86.00 | 94.67 | 49/50 |
| CoffeeServeMug | 12.67 | 20.00 | 54.00 | 62.00 | 64.67 | 6.00 | 24.00 | 61.33 | 67.33 | 69.33 | 38/50 |
| Policy | Training seed | Repeat 1 | Repeat 2 | Repeat 3 | Mean SD |
|---|---|---|---|---|---|
| BC-Transformer | 123 | 32.00 | 30.78 | 31.04 |
| Policy | Training data | Parameters | Tasks | SR (%) |
|---|---|---|---|---|
| WorldToken | 2,900 | 85.3M | 23 | 59.4 |
| 300 | 3.35B | 24 | 62.1 |
| (a) Complete token-interface results | |||||||
| Params | Rel. FLOPs | RMSE | SR: repeats 1 / 2 / 3 | SR: mean SD | |||
| 50 | 1 | 85.3M | 10 | 0.19514 | 20.96/20.96/22.26 | ||
| 4 | 92.4M | 40 | 0.20163 | 20.00/19.48/21.30 | |||
| 50 | 83.0M | 500 | 0.19086 | 25.83/27.04/26.61 | |||
| 100 | 1 | 85.3M | 10 | 0.17212 | 30.17/31.74/31.65 | ||
| 4 | 92.4M | 40 | 0.16586 | 38.17/38.00/38.26 | |||
| (a) One-observation baseline: | ||||||
| CNN | Fusion + projections | Encoder total | Temporal backbone | Action head | Total | |
| 1 | 19.1103 | 2.1189 | 21.2292 | 0.0578 | 0.4502 | 21.7372 |
| 4 | 19.1103 | 2.1330 | 21.2433 | 0.2314 | 0.4502 | 21.9250 |
| 50 | 19.1103 | 1.9606 | 21.0709 | 2.9209 | 0.4502 | 24.4420 |
| SR (%): repeats | SR (%) | |||||
|---|---|---|---|---|---|---|
| Seed | RMSE | 1 | 2 | 3 | Mean SD | |
| 1 | 0 | 0.150260 | 45.65 | 47.65 | 46.96 | |
| 1 | 0.147767 | 48.17 | 46.00 | 48.87 | ||
| 2 | 0 | 0.149383 | 47.83 | 46.52 | 46.26 | |
| 1 | 0.147528 | 47.74 | 50.09 | 47.74 | ||
| 5 | 0 | 0.137165 | 51.74 | 51.22 | 51.48 | |
| Setting | Value |
|---|---|
| Initial / continuation budget | 5,000 / 500 optimizer steps; continuation uses 45 training / 5 continuation-holdout demonstrations |
| Training context cap | Initial ; continuation |
| Microbatch / accumulation | 1 / 8 steps |
| Optimizer | AdamW, , , zero weight decay |
| Gradient clipping / precision | Unit gradient norm / BF16 |
| Reference learning rates | Within-timestep encoder and temporal backbone: ; action decoder: |
| Standard loss | Modified loss | |
| (a) Episode outcomes | ||
| Evaluator successes | 95 | |
| Evaluator failures | 5 | |
| Target block geometry reached | 98 | |
| Failures reaching target block geometry | 3 | |
| Failures without target block geometry | 2 | |
| Quantity | Operational rule |
|---|---|
| Stable order | Same left-to-right block order for eight consecutive action-level records satisfying the motion, row/height, and right-gripper criteria below |
| Motion | Maximum block displacement between consecutive sampled positions mm |
| Row and height | Each block within 6 cm of m and m |
| Right gripper | Open-command value |
| Final nominal-slot placement | For blocks , m and m; maximum absolute error cm at evaluator success |
| Sensitivity cutoffs | Recompute the final-placement classification at 2, 3, 4, and 5 cm; keep sequence and evaluator labels fixed |
| (a) History windows and strict-success durations | |||||
|---|---|---|---|---|---|
| Visible window (s) | 145.92 | 69.12 | 30.72 | 15.36 | 7.68 |
| Longest strict success (s) | 142.62 | 142.50 | 142.26 | 59.94 | 31.98 |
| (b) Block placement and swap order in successful episodes | |||||
| Strict-success placement: minimum (cm) | 0.36 | 0.36 | 0.36 | 0.45 | 0.25 |
| Strict-success placement: median (cm) | 1.63 | 1.66 | 0.84 | 0.62 | 0.53 |