Vision-language-action (VLA) models using discrete action tokens have proven effective for controling robotic arms on manipulation tasks. For a humanoid, however, the whole-body action space -- legs, torso, arms, and hands -- is far higher-dimensional and heterogeneous, raising tokenization, training, and real-time inference challenges that the previous VLA models do not address. We present Holo-M, to our knowledge the first discrete VLA model for humanoid loco-manipulation that intrinsically exploits the language model by extending its vocabulary with action tokens. In this model, we devise a unified action tokenizer that decomposes the humanoid action space into four body-part-specific tokenizers -- end-effector, body, hand, and kinematics -- enabling training across drastically different embodiments and data sources, including humanoid teleoperation, ego-centric human video, and simulation. By extending the language model's vocabulary with these action tokens, we avoid the knowledge-insulation problem inherent to the models that use separate continuous action experts. To meet real-time control requirements, we decode each body part's action tokens through grouped discrete diffusion decoding, rather than using autoregression on the action tokens. We have conducted extensive experiments on the SIMPLE humanoid loco-manipulation benchmark, in which Holo-M achieves the highest success rates in both the generalist and specialist evaluations, leading the second best by significant margins. We will release all the code and model weights.
Figures & tables
Figure 1: Overview of the proposed Holo-M framework. A unified action tokenizer, pre-trained across heterogeneous data (teleoperated, human-centric, simulated), decomposes the humanoid action space into four body-part tokenizers (EEF, body, hand, kinematics) sharing one discrete vocabulary. At inference, the VLM backbone encodes the image, language, and proprioceptive state into this vocabulary and decodes action tokens via grouped discrete diffusion – parallel within a body part, autoregressive across parts – which are then detokenized into whole-body controller commands driving the humanoid.
Group
Representation
Dim.
Tokens
End effector
Wrist and fingertip poses
48
100
Body
Body joint targets
29
62
Hand
Hand joint targets
14
32
Kinematics
Base-motion commands
5
14
Full action
All groups
96
208
Table 1: Canonical action-space decomposition used by Holo-M.
Figure 2: Progressive training of Holo-M through cross-embodiment autoregressive pre-training, humanoid post-training, grouped discrete diffusion fine-tuning, and task-specific adaptation.
Pretraining Data
Open-loop MAE ↓
Closed-loop Success Rate ↑
No Pretraining
0.022
53/180
HE
0.020
119/180
EgoDex + HE
0.019
138/180
Table 5: Effect of scaling pre-training data on Holo-M AR.
Figure 3: Our fixed schedule along the tick axis using 8 de-masking steps as example: observation is taken every 500ms, which the model consumes to infer 1-second action chunk. During model inference, the action steps from the previous chunk are executed.
De-masking steps
Inference latency, mean ± std (ms) ↓
2
175.8±11.5
4
275.7±9.3
8
476.8±16.9
Table 6: Real-world inference latency per 30-timestep chunk.
Task
Success Rate
Tabletop Grasp
10/10
Move Pick
8/10
Table 7: Real-robot success rate over 10 rollouts per task.