UNITAS: A 3D-Native World Action Model for Embodied Manipulation
Authors: Ruixiang Wang, Yongyi Su, Wenlve Zhou, Bo Yue, Hengyan Liu, Dekun Lu, Yuxin Tian, Yihan Fang, +9 more
Organizations: The Chinese University of Hong Kong, Shenzhen · DexForce Technology Co., Ltd. · Foshan University · Sun Yat-sen University · South China University of Technology
World action models (WAMs) aim to answer a coupled physical question: given a task instruction, what motion should the robot execute, and how will that motion change the surrounding world? Most existing WAMs build on pretrained video generators and represent world evolution through images or visual latents. Robotic interaction, however, takes place in metric three-dimensional space, while images are view-dependent projections whose pixel distances do not directly encode physical distances. We introduce UNITAS, to our knowledge the first 3D-native world action model that unifies observations, actions, and scene dynamics in a shared metric 3D frame within each interaction, using a common representation across robot embodiments and human hands. Action flow represents human hands and robot grippers as 3D point trajectories, while scene flow describes scene-point displacements conditioned on these trajectories. World-aligned 3D positional embeddings ground visual tokens with or without depth input, and a physical-time trajectory tokenizer encodes each point trajectory as one token anchored at its current 3D position. This interface supports both direct action execution and action-conditioned scene prediction. With 1.7B parameters, UNITAS achieves the best action-conditioned scene prediction on RoboTwin among the compared methods, with up to 49% lower displacement errors than PointWorld, and state-of-the-art manipulation success, including 99.8% on LIBERO and an average of 85% across real-world tasks. The code is available at https://github.com/DexForce/UNITAS.
Figures & tables
Figure 1: Unitas is a 3D-native world action model. Observations, action flow, and scene flow share a metric frame within each prediction window, with a common representation across embodiments. With 1.7B parameters, Unitas achieves leading manipulation results among the compared methods and predicts action-conditioned scene trajectories.
Figure 2: Unitas architecture and attention masks. Left: The point dynamics expert models both action and scene flows, while the action expert generates the end-effector action chunk. Scene flow generation is conditioned on clean demonstrated action flow during training and generated or supplied action flow at inference. Right: Colored/white cells indicate allowed/blocked attention. Policy-only inference omits both future action-flow and scene-flow tokens.
Figure 3: Shared point-flow representations across data sources. Left: gripper/hand trajectories (purple) and scene trajectories (yellow–blue), with RGB insets. Right: RoboTwin end-effector XYZ error (mm). Thick curves show exponential moving averages (EMA) of the thin curves.
Table 4
Figure 5: Action-conditioned scene-flow predictions. Top: observed RGB images. Middle: initial point clouds. Bottom: predicted endpoint point clouds and scene trajectories (color gradient), conditioned on recorded gripper trajectories (purple). Translucent robot meshes indicate the corresponding poses.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 6: Representative observations and aligned geometry from pretraining data. Robot panels show RGB, depth where available, and scene points with robot overlays. VITRA shows egocentric RGB and hand landmarks.
Dataset
Source
f (Hz)
H
Hhist
Duration
H+1
RoboTwin
Simulation
15
20
20
1.33
21
LIBERO
Simulation
20
20
20
1.00
21
VLABench
Simulation
10
15
15
1.50
16
DROID
Real robot
15
20
20
1.33
21
AgiBot
Real robot
15
20
20
1.33
21
RoboMIND
Real robot
15
20
20
1.33
21
Appendix
Table 5: Pretraining trajectory sampling configuration. H and Hhist count future and historical intervals; duration is H/f in seconds. H+1 includes the current anchor.
Data source
Action flow
Scene flow
ADE ↓
FDE ↓
ADE ↓
FDE ↓
RoboTwin
1.96
4.62
1.01
2.32
LIBERO
2.54
6.05
0.78
1.56
VLABench
3.38
6.25
0.94
1.89
DROID
3.17
7.85
–
–
AgiBot
2.26
4.62
–
–
Appendix
Table 6: Held-out trajectory reconstruction errors of the frozen tokenizer (mm). Dashes indicate unavailable scene-flow ground truth.
Component
Configuration
Params. (M)
VLM
Florence-2-large
570.7
Point dynamics expert
28 layers; width 1,280; 20 heads
826.3
Action expert
28 layers; width 640; 10 heads
256.0
Trajectory tokenizer
Time-conditioned MLPs; latent dim. 256
1.9
Embeddings and projections
Spatial and modality interfaces
39.7
Total
1,694.5
Appendix
Table 7: Model architecture. Parameter counts include the frozen tokenizer.
Method / output mode
Params (B)
Ns
Latency (ms)
Fast-WAM (w/o video)
6.0
–
156.81
Fast-WAM (joint)
6.0
–
579.79
Motus
8.0
–
2,414.90
Unitas : actions only
1.7
–
55.77
Unitas : actions + action flow
1.7
–
122.51
Unitas : actions + both flows
1.7
1,024
305.49
Appendix
Table 8: Inference efficiency on an RTX 5090 D. Latency is measured per prediction chunk; Ns denotes the number of scene queries.
Task
π0.5
Motus
Fast-WAM
Unitas (w/o Pretrain)
Unitas
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Clean
Rand.
Adjust Bottle
100
99
89
93
100
100
100
98
100
100
Beat Block Hammer
96
93
95
88
99
97
97
92
100
98
Blocks Ranking RGB
92
85
99
97
100
100
100
99
97
94
Blocks Ranking Size
49
26
75
63
94
98
91
83
78
82
Click Alarmclock
98
89
100
100
100
100
100
100
100
100
Appendix
Table 12: Detailed RoboTwin 2.0 success rates (%, ↑ ). Clean and Rand. denote demo_clean and demo_randomized . All methods are evaluated over 100 episodes per task in each setting.
Figure 8: Additional scene-flow visualizations (2/4): bowl and mug placement, articulated motion, and gravity-driven motion after object release.
Figure 9: Additional scene-flow visualizations (3/4): book, bottle, block, and bowl manipulation across LIBERO and RoboTwin.
Figure 10: Additional scene-flow visualizations (4/4): object placement, drawer opening and closing, and microwave opening.
Figure 11: Zero-shot scene-flow predictions on real-world observations. The first four rows are from DROID and the last is from RH20T. Columns show observed RGB, initial points, and predicted midpoint and endpoint point clouds with scene trajectories (orange to purple) and ground-truth gripper trajectories used for conditioning (magenta). No scene-trajectory ground truth is shown.