Organizations: School of Remote Sensing and Information Engineering, Wuhan University, Wuhan, China. · Department of Electronic Engineering, Tsinghua University, Beijing, China. · Institute for AI Industry Research, Tsinghua University, Beijing, China.
Vision-language-action policies inherit both capabilities and input representations from pretrained vision-language models, showing great potential for robotic manipulation across diverse industrial settings. As these policies increasingly use interaction history, organizing the representation of historical observations determines how experiences enter temporal context and how relationships across time are modeled, which is a generally ignored challenge in previous works. In this work, we believe solving this challenge requires an architectural reconstruction and propose \textbf{WorldToken}. WorldToken encodes each timestep's observations into one world token, processes the resulting history with a causal Transformer, and generates action chunks with a diffusion action head. Unified token enables long horizon tasks while relieving the memory requirement of the temporal backbone, hence improving performance. This design also offers high interpretability and allows advances in language modeling, such as pre-training and scaling, to be transferred to robot interaction policies. In RoboCasa experiments, WorldToken successfully handles most tasks with 85M parameters and achieves 59.4% mean closed-loop success close to π0.5 with 3.35B parameters. On the memory benchmark RMBench Blocks Ranking, WorldToken can reach the context of two minutes and achieves a success rate of 95%. In addition, holdout action RMSE is well described by power-law fits, and its closed-loop success rate improves consistently with increasing training data size in a study involving approximately 350,000 closed-loop evaluation episodes across 50 trained policies on RoboCasa, showing its scaling potential. We also conduct experiments to analyze information preservation and history use in WorldToken, providing empirical grounding for future work.
Figures & tables
Figure 1: Structure of a language model and WorldToken. (a) An LM head maps the language model’s contextual representation to a distribution over the next text token, which is appended to the sequence. (b) WorldToken encodes each policy timestep’s multimodal observations into one world token and uses a diffusion action head to generate an action chunk from the history-conditioned state. Executed actions affect the environment, whose next observation supplies the next world token.
Figure 2: WorldToken architecture. The encoder maps each observation to one world token.
Figure 3: Examples of manipulation process.
Training seed 0
Training seed 1
D50
D100
D300
D1000
D2900
D50
D100
D300
D1000
D2900
SR (%) ↑
N1
14.1
30.0
41.0
49.1
54.4
18.1
27.4
41.4
49.9
54.7
N2
21.4
31.2
47.5
55.7
59.1
18.9
32.4
46.2
52.2
59.8
N3
22.0
35.8
48.1
58.5
59.3
22.7
33.2
48.4
57.8
61.2
N4
27.1
36.5
51.2
57.7
58.8
24.0
36.3
51.4
56.7
57.8
N5
23.1
33.7
49.8
56.2
59.7
21.7
34.6
49.9
56.4
59.3
Table 1: Performance of WorldToken across dataset sizes and model capacities on RoboCasa.
Figure 4: On RoboCasa, WorldToken’s RMSE decreases approximately as a power law with dataset size, while gains from increasing model capacity diminish. (a) RMSE versus dataset size for each model capacity. (b) RMSE versus model capacity for each dataset size. Both panels use log axes. Markers show individual training seeds, and solid lines connect the means across the two seeds. Dashed lines in (a) show fits of RMSE∝D−α to these means. The table reports the fitted scaling exponent α and the coefficient of determination R2 for each model, with R2 computed in log–log space.
SR (%) ↑
RMSE ↓
K
Seed
Params
RelPs
D50
D100
D300
D1000
D2900
D50
D100
D300
D1000
D2900
1
0
85.3M
1.0×
21.4
31.2
47.5
55.7
59.1
0.195
0.172
0.137
0.109
0.091
1
1
85.3M
1.0×
18.9
32.4
46.2
52.2
59.8
0.197
0.169
0.135
0.110
0.091
4
0
92.4M
4.0×
20.3
38.1
48.8
55.6
59.9
0.202
0.166
0.138
0.110
0.090
50
0
83.0M
55.3×
26.5
37.6
49.6
57.7
60.7
0.191
0.165
0.132
0.107
0.089
Table 2: Token-interface comparisons with a context of C=10 policy timesteps. K is the number of temporal tokens contributed by each timestep. Relative FLOPs (RelPs) are the leading analytic full-prefix temporal-backbone counts (Appendix D.3 ), normalized to the N2, K=1 reference. Encoder and action-head costs are excluded. Each SR entry is the mean over three complete evaluations.
Ctest=1
Ctest=2
Ctest=5
D50
D100
D300
D1000
D2900
D50
D100
D300
D1000
D2900
D50
D100
D300
D1000
D2900
Seed 0
N1
−3.5
−10.1
−18.1
−14.9
−13.0
−4.1
−3.3
−11.0
−8.6
−6.8
−0.9
+0.4
−0.5
−0.1
−0.6
N2
−7.1
−8.7
−17.8
−18.0
−15.3
−5.3
−6.7
−12.3
−6.5
−6.2
+0.3
−0.1
−1.1
+0.8
+1.5
N3
−7.3
−11.1
−20.2
−20.3
−18.9
−4.4
−9.9
−13.2
−11.7
−5.6
+0.3
−0.3
−1.1
−0.6
0.0
N4
−10.5
−15.0
−21.3
−18.9
−22.2
−8.9
−11.3
−11.9
−10.1
−9.2
−1.4
−1.9
−1.3
−0.6
−1.0
N5
−6.4
−13.6
−19.7
−17.3
−18.6
−5.7
−8.5
−12.7
−9.4
−8.8
+0.5
−0.4
−1.5
+1.5
−1.0
Table 3: SR changes when the same policies are evaluated with shorter histories. Each cell reports ΔSR=SR(Ctest)−SR(10) in percentage points for one, two, and five visible policy timesteps.
Context C
RMSE, seed 0
RMSE, seed 1
SR (%) , seed 0
SR (%) , seed 1
1
0.150260
0.147767
46.75 ± 1.02
47.68 ± 1.50
2
0.149383
0.147528
46.87 ± 0.84
48.52 ± 1.36
5
0.137165
0.139603
51.48 ± 0.26
50.00 ± 0.92
10
0.131710
0.133098
48.12% ± 1.70 pp
48.38 ± 1.09
Table 4: RMSE and SR at different history lengths, with Ctrain=Ctest=C .
Figure 5: History length for N3 on D300 ( Ctrain=Ctest=C ). (a) Holdout action RMSE during training. (b) Mean final SR over three fixed-seed repeats per checkpoint. Error bars show sample standard deviations.
Ctest\S (episodes)
1 (24)
2 (14)
3 (28)
4 (18)
5 (16)
Total (100)
608
24/24/0
14/14/0
25/25/0
17/16/1
15/15/0
95/94/1
288
24/24/0
14/14/0
23/23/0
16/16/0
15/15/0
92/92/0
128
24/24/0
14/14/0
13/3/10
5/5/0
3/3/0
59/49/10
64
24/24/0
3/3/0
2/0/2
9/0/9
0/0/0
38/27/11
32
24/24/0
3/0/3
0/0/0
1/0/1
0/0/0
28/24/4
Table 5: Blocks Ranking outcomes stratified by the number of swaps S in the reference solution. Each entry gives evaluator successes / strict behavioral successes / evaluator-only successes.
Figure 6: One episode evaluated with two history lengths. The checkpoint, initial state, and evaluation protocol are fixed, only Ctest changes. With Ctest=608 , the policy completes the five-swap reference sequence and succeeds at 142.6 seconds. With Ctest=64 , it completes the first swap but fails to finish the task within the 210-second horizon. Frame timestamps are simulated seconds.
Env. seed
S
Swaps
Last (s)
Env. seed
S
Swaps
Last (s)
Env. seed
S
Swaps
Last (s)
100000
2
13
350.52
100004
3
8
222.12
100015
5
5
139.38
100001
3
5
139.20
100003
2
9
250.14
100016
1
5
139.32
100002
4
17
470.28
100007
1
5
139.08
100008
1
31
856.44
Table 6: Nine exploratory Blocks Ranking stress trajectories with rollouts allowed to continue after first success. S is the number of swaps required to first reach the target. “Swaps” counts correctly ordered reference-sequence swaps. “Last” is the time of the final correct swap.
Appendix figures & tables23 assets
Supplementary material from the paper’s appendix.
Appendix
Setting
N1
N2
N3
N4
N5
Parameters
44.3M
85.3M
218.8M
648.9M
1490.3M
Encoder/temporal LR ( 10−4 )
6.00
4.25
3.00
1.50
0.50
Action-decoder LR ( 10−4 )
3.00
3.00
3.00
3.00
3.00
Within-timestep fusion encoder (SwiGLU)
Width
512
768
1024
1536
2048
Layers
1
2
4
6
8
Appendix
Table 7: RoboCasa N1–N5 architecture and peak learning rates. Parameters include all visual stems and policy modules, excluding the frozen CLIP text encoder. FF denotes feed-forward hidden width; Q and KV denote query and key/value attention heads.
Component
Protocol
Inputs
Three 128×128 RGB views, 16-D proprioception, and a frozen 768-D CLIP task embedding
Visual stems
One five-stage CNN per camera; 5×5 convolutions with channels (96,144,192,384,768) , each followed by 2×2 max pooling, group normalization, and SiLU; output grid 4×4×768 per view
Token readout
48 visual tokens plus proprioception and task tokens; four learned readouts are concatenated, RMS-normalized, and projected to one temporal world token
Temporal sampling
20 Hz environment control; one observation, world token, and replanning decision every four control steps (5 Hz)
Actions
Same-frame-aligned 12-D commands; H=10 predicted, Hexec=4 executed before replanning
Context and diffusion
Ctrain=10 world tokens; 20-stage cosine schedule and stochastic 20-step DDPM evaluation
Appendix
Table 8: RoboCasa training and closed-loop execution recipe. The main sweep uses Ctrain=10 ; separately trained context variants are identified where they are analyzed.
Experiment
Changed / held fixed
Training seeds
Executions
Data and model scaling
Five D values × five capacities; Ctrain=Ctest=10
0, 1
3
Token count
K=1,4,50 at all five D ; d=768 , C=10
0
3
Same-checkpoint history
Ctest=1,2,5 for every scaling checkpoint; Ctrain=10
0, 1
1
Matched training history
Ctrain=Ctest=1,2,5,10 ; 218.8M, D=300
0, 1
3
Appendix
Table 9: RoboCasa experiment configurations. “Executions” counts complete 1,150-episode evaluations per checkpoint and stated inference setting, not independent training seeds.
Single execution
h=10 repeats
h=10
D
Model
Seed
RMSE
h=1
h=2
h=5
1
2
3
Mean ± SD
50
N1
0
0.20765
10.61
10.00
13.22
14.52
13.91
14.00
14.14±0.33
1
0.20301
12.52
15.57
18.78
18.17
18.96
17.30
18.14±0.83
N2
0
0.19514
14.26
16.09
21.65
20.96
20.96
22.26
21.39±0.75
1
0.19682
10.96
13.57
18.09
20.35
17.65
18.78
18.93±1.35
N3
0
0.19532
14.70
17.57
22.26
22.70
22.00
21.22
21.97±0.74
Appendix
Table 10: Closed-loop SR (%) for all 50 scaling policies trained with Ctrain=10 . Each checkpoint is evaluated once at h=1,2,5 and three times at h=10 , where h=Ctest counts policy timesteps. Each evaluation contains 1,150 episodes; mean ± sample SD is reported only for h=10 . Final-checkpoint stochastic full-chunk RMSE follows Appendix B.4 . N1–N5 are defined in Table 7 .
D
Params
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%
50
44.3M
0.41821
0.31303
0.27148
0.25973
0.23128
0.21949
0.21252
0.21222
0.20983
0.20765
85.3M
0.41246
0.35694
0.24800
0.22547
0.21357
0.20352
0.20087
0.19598
0.19501
0.19514
218.8M
0.40335
0.34023
0.27017
0.23318
0.21947
0.20875
0.19826
0.19802
0.19648
0.19532
648.9M
0.33931
0.23841
0.22386
0.20937
0.20262
0.19702
0.19359
0.19386
0.19313
0.19279
1.49B
0.27697
0.21628
0.20776
0.19756
0.19285
0.19263
0.19124
0.19109
0.19044
0.18934
100
44.3M
0.39791
0.27210
0.22569
0.20335
0.19464
0.18634
0.18098
0.17686
0.17595
0.17477
Appendix
Table 11: RoboCasa holdout stochastic full-chunk RMSE at 10% training intervals, seed 0, using the metric in Appendix B.4 . Progress is optimizer steps divided by the scheduled budget: 5k/10k/30k/100k/280k for D=50/100/300/1000/2900 . Entries use exact logged steps, without interpolation.
D
Params
10%
20%
30%
40%
50%
60%
70%
80%
90%
100%
50
44.3M
0.41265
0.36703
0.30551
0.24915
0.24223
0.22526
0.21346
0.20761
0.20684
0.20301
85.3M
0.41460
0.31348
0.25889
0.23951
0.22562
0.21208
0.20414
0.19972
0.19819
0.19682
218.8M
0.39254
0.28892
0.23680
0.21865
0.20694
0.20339
0.19930
0.19517
0.19357
0.19411
648.9M
0.37310
0.25907
0.22671
0.21701
0.20702
0.19907
0.20138
0.19804
0.19506
0.19518
1.49B
0.26374
0.21731
0.20156
0.19317
0.19170
0.18972
0.18823
0.18731
0.18720
0.18610
100
44.3M
0.28741
0.23559
0.20988
0.19770
0.18795
0.18145
0.17921
0.17774
0.17699
0.17701
Appendix
Table 12: RoboCasa holdout stochastic full-chunk RMSE at 10% training intervals, seed 1, using the metric in Appendix B.4 . Progress is optimizer steps divided by the scheduled budget: 5k/10k/30k/100k/280k for D=50/100/300/1000/2900 . Entries use exact logged steps, without interpolation.
Training seed 0: D
Training seed 1: D
Selected run
Task
50
100
300
1000
2900
50
100
300
1000
2900
Successes
CloseDoubleDoor
23.33
67.33
88.00
91.33
94.67
36.67
68.67
92.67
90.67
91.33
45/50
CloseDrawer
88.00
98.67
100.00
100.00
98.67
87.33
96.00
99.33
100.00
97.33
48/50
CloseSingleDoor
57.33
67.33
82.67
83.33
87.33
50.67
76.00
76.00
84.00
84.00
45/50
CoffeePressButton
44.67
56.67
78.00
91.33
94.67
38.00
51.33
84.67
86.00
94.67
49/50
CoffeeServeMug
12.67
20.00
54.00
62.00
64.67
6.00
24.00
61.33
67.33
69.33
38/50
Appendix
Table 13: Task-level RoboCasa SR (%) for the 218.8M policy. Each data-scale entry averages three executions of 50 episodes per task, with training seeds separate. The final column records successes in the selected single execution: seed 1, D=2900 , repeat 02 (713/1,150 overall).
Policy
Training seed
Repeat 1
Repeat 2
Repeat 3
Mean ± SD
BC-Transformer
123
32.00
30.78
31.04
31.28±0.64
Appendix
Table 14: Local BC-Transformer trained with the official native recipe (Appendix C.4 ): training seed 123, D=300 , final 500k-step checkpoint. All three repeats use the same 1,150 episodes across 23 tasks. SR is in percent; SD is the sample standard deviation across repeats, in percentage points.
Policy
Training data
Parameters
Tasks
SR (%)
WorldToken
2,900
85.3M
23
59.4
π0.5
300
∼ 3.35B
24
62.1
Appendix
Table 15: WorldToken and the reported π0.5 result on RoboCasa Kitchen. Training data counts are generated demonstrations per task. The WorldToken result averages two training seeds, each with three evaluations. The π0.5 result is reported by Kim et al. (2026) .
(a) Complete token-interface results
D
K
Params
N
Rel. FLOPs
RMSE
SR: repeats 1 / 2 / 3
SR: mean ± SD
50
1
85.3M
10
1.00×
0.19514
20.96/20.96/22.26
21.39±0.75
4
92.4M
40
4.03×
0.20163
20.00/19.48/21.30
20.26±0.94
50
83.0M
500
55.31×
0.19086
25.83/27.04/26.61
26.49±0.62
100
1
85.3M
10
1.00×
0.17212
30.17/31.74/31.65
31.19±0.88
4
92.4M
40
4.03×
0.16586
38.17/38.00/38.26
38.14±0.13
Appendix
Table 16: Token-count experiments at C=10 , training seed 0. (a) Final-checkpoint results for all 15 (D,K) settings. Stochastic full-chunk RMSE follows Appendix B.4 . SR entries give three evaluation percentages (1,150 episodes each) and their mean ± sample SD. The K=1,4,50 interfaces are defined in Appendix D . (b) Temporal dimensions and analytic full-prefix FLOPs (Equation 19 ), excluding encoder and decoder. Relative FLOPs use the 85.3M, K=1 reference; parameter count and compute are not jointly matched.
(a) One-observation baseline: C=1
K
CNN
Fusion + projections
Encoder total
Temporal backbone
Action head
Total
1
19.1103
2.1189
21.2292
0.0578
0.4502
21.7372
4
19.1103
2.1330
21.2433
0.2314
0.4502
21.9250
50
19.1103
1.9606
21.0709
2.9209
0.4502
24.4420
Appendix
Table 17: Cached computation in GFLOPs per timestamp, where one timestamp denotes one policy query. (a) Baseline with one observation, C=1 . Fusion/projections include encoder readouts, and encoder total is a subtotal. Action-head costs include all 20 diffusion steps; totals are computed before rounding. (b) Additional cost per query for 100 more retained observations. Relative slopes compare the cost of added history, not total policy costs.
SR (%): repeats
SR (%)
C
Seed
RMSE
1
2
3
Mean ± SD
1
0
0.150260
45.65
47.65
46.96
46.75±1.02
1
0.147767
48.17
46.00
48.87
47.68±1.50
2
0
0.149383
47.83
46.52
46.26
46.87±0.84
1
0.147528
47.74
50.09
47.74
48.52±1.36
5
0
0.137165
51.74
51.22
51.48
51.48±0.26
Appendix
Table 18: D300 matched-context training with N3 (218.8M), Ctrain=Ctest=C . Each final 30k-step checkpoint is evaluated three times on 1,150 episodes. Stochastic full-chunk RMSE uses the context-specific crops in Appendix B.4 . SR entries give the repeated evaluation percentages and their mean ± sample SD. C=10 is the scaling reference.
Within-timestep encoder and temporal backbone: 4.25×10−4 ; action decoder: 3.0×10−4
Appendix
Table 19: Recorded Blocks Ranking continuation and action-generation settings. The same step-5,500 checkpoint is used for every visible-history intervention.
Standard loss
Modified loss
(a) Episode outcomes
Evaluator successes
13
95
Evaluator failures
87
5
Target block geometry reached
99
98
Failures reaching target block geometry
86
3
Failures without target block geometry
1
2
Appendix
Table 20: Standard-loss and modified-loss results at C=608 on the same fixed evaluation set of 100 initial conditions. Panels report episode outcomes, paired outcome counts, and button-position diagnostics (mm). Definitions and source records are given in Appendix F.4 .
Quantity
Operational rule
Stable order
Same left-to-right block order for eight consecutive action-level records satisfying the motion, row/height, and right-gripper criteria below
Motion
Maximum block displacement between consecutive sampled positions ≤1.5 mm
Row and height
Each block within 6 cm of y=−0.10 m and z∈[0.740,0.785] m
Right gripper
Open-command value >0.8
Final nominal-slot placement
For blocks 1,2,3 , x∗=(0.04,0.16,0.28) m and y∗=−0.10 m; maximum absolute x/y error ≤4 cm at evaluator success
Sensitivity cutoffs
Recompute the final-placement classification at 2, 3, 4, and 5 cm; keep sequence and evaluator labels fixed
Appendix
Table 21: Criteria for stable block order and final placement. Stability criteria are shared with the extended-rollout sequence monitor. The final-placement cutoff is used only to classify strict behavioral success.
C=608
C=288
C=128
C=64
C=32
(a) History windows and strict-success durations
Visible window (s)
145.92
69.12
30.72
15.36
7.68
Longest strict success (s)
142.62
142.50
142.26
59.94
31.98
(b) Block placement and swap order in successful episodes
Strict-success placement: minimum (cm)
0.36
0.36
0.36
0.45
0.25
Strict-success placement: median (cm)
1.63
1.66
0.84
0.62
0.53
Appendix
Table 22: Placement, sequence, and failure diagnostics for the 100 evaluation episodes at each history length. Aggregate and swap-stratified outcomes appear in Table 5 ; success criteria are defined in Appendix F.5 . Panels (a)–(e) describe history length, placement, and behavioral outcomes. Panel (f) reports swap-and-press cycle durations at C=608 .
Figure 7: Holdout stochastic RMSE versus closed-loop SR for all 50 trained policies. The x -axis is task-averaged RMSE over all 12 action dimensions and all ten steps in each sampled action chunk; the y -axis is mean SR over three complete evaluations with a ten-step context. Color, marker, and fill encode demonstrations per task D , capacity, and training seed. Error bars are fixed-seed repeatability sample standard deviations (Appendix B.4 ). Descriptive Spearman correlation is ρ=−0.980 , summarizing the sweep-level trend rather than ranking nearby checkpoints.
Figure 8: Pick and place (I): cabinet and sink transfers. The robot moves objects between the counter and a cabinet (a,b), and between the counter and a sink (c,d).
Figure 9: Pick and place (II): microwave and stove transfers. The robot places an object in the microwave (a), moves an egg from a plate to a pan (b), and moves an apple from a pan to a plate (c).
Figure 10: Opening and closing doors and drawers. Panels (a–c) show opening a cabinet door, closing a cabinet door, and closing both cabinet doors, respectively. Panels (d,e) show opening and closing a drawer.
Figure 11: Sink and stove controls. Panels (a–c) show faucet and spout control in the Turning levers family. Panels (d,e) show switching stove burners on and off in the Twisting knobs family.
Figure 12: Mug placement and retrieval at the coffee machine. The two tasks belong to the official Insertion family: placing a mug under the dispenser (a) and moving it from the dispenser to the counter (b).
Figure 13: Pressing appliance buttons. The robot presses the button on the coffee machine (a), the microwave start button (b), and the microwave stop button (c).
A common approach to world-model simulation for vision-language-action (VLA) systems is to predict future RGB observations and then re-encode them into policy inputs, introducing an indirect interface between simulation and downstream policy execution. We instead investigate whether world dynamics can be modeled in a compact, policy-oriented state derived from VLM visual tokens. A key challenge is that raw VLM visual tokens are high-dimensional, making efficient and accurate autoregressive dynamics modeling challenging. To address this, we introduce Token-World, an action-conditioned world model that compresses VLM features into a compact token state, learns future dynamics in this reduced space, and maps predicted states back to the original policy-facing representation for downstream use. Across manipulation benchmarks, Token-World improves open-loop feature fidelity and policy-action consistency over recent world-model simulators, with slower degradation over long rollout horizons. In closed-loop evaluation, its simulated policy performance correlates more strongly with reference policy performance than Ctrl-World (r=0.794 vs.\ 0.583), while requiring lower simulation latency. Ablations further show that compact-representation design and dimensionality substantially affect future-state prediction. Code will be available at https://chuyaofu.github.io/Token-World/.
Chuyao Fu, Xiaowei Chi, Yuhan Rui +14
Southern University of Science and Technology · MUKA Robotics · Hong Kong University of Science and Technology +3
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes evolve as a task unfolds. Recently emerging world-action models (WAMs) leverage pretrained video world models to capture spatiotemporal evolution, yet retaining future generation or a large video backbone in the control loop substantially increases inference cost. We introduce World Tokens, an embodied policy architecture built around a World Adapter that bridges visual-language understanding, world-dynamics modeling, and action generation. It uses world modeling during training to enhance the action policy while preserving efficient deployment. Specifically, the World Adapter transforms VLM features into a fixed set of world tokens, which condition a jointly fine-tuned future-video denoiser and simultaneously serve as the action expert's sole visual-language context. This shared conditioning allows gradients from future-video denoising to directly shape the representation used for action prediction, while exclusive routing prevents the policy from bypassing that representation. At deployment, the world-model branch is removed, leaving only the VLM, World Adapter, and action expert, with no online video-model inference. With a 2B backbone and no embodied action pretraining, World Tokens is highly competitive on LIBERO, attains the best reported averages on SIMPLER, substantially improves real-world R1 Pro success over a matched action-only baseline, and generates each action chunk at VLA-level latency.
This work presents RepWAM, a representation-centric world action model (WAM) built on representation visual-action tokenizers. Existing WAMs typically inherit reconstruction-oriented video tokenizers from pretrained video generation models. Although these tokenizers preserve visual fidelity, pixel reconstruction alone provides limited guidance for learning instruction-following dynamics that connect future prediction with robot control. To address this, we explore a semantic visual-action latent space for representation-centric world action modeling. Specifically, we train a representation visual-action tokenizer that maps visual inputs into aligned visual and latent action tokens. We then pretrain our WAM to jointly model future visual states and the latent actions that connect them under language instructions, followed by adaptation to real robot trajectories for closed-loop manipulation. Experiments on real-world manipulation tasks and simulation benchmarks show that RepWAM delivers strong performance across diverse manipulation settings, while ablations highlight the value of semantic visual-action tokenization over reconstruction-oriented alternatives. These results establish representation visual-action tokenization as a promising foundation for world action models and a step toward generalist robot policies. Code and weights will be available at https://github.com/wdrink/RepWAM.
Junke Wang, Qihang Zhang, Shuai Yang +5
Institute of Trustworthy Embodied AI, Fudan University · 2Robbyant, Ant Group · 3Hongkong University of Science and Technology