Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1% in position error and 44.0% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.
Figures & tables
Figure 1 : Challenges of mobile manipulation and our approach. Existing unified policies either share one action representation across navigation and manipulation or separate the two only inside the action module, and they learn where to stop only implicitly from costly complete trajectories. UniWAM models navigation and manipulation as separate streams over a shared backbone, and MAP-Data supervises the manipulation-ready pose at scale with automatically generated trajectories.
Model
Action generator
Separate views
One behavior per sample
Video world model
Explicit terminal pose
π0.5 [ 3 ]
single action expert for all degrees of freedom
✗
✗
✗
✗
GR00T N1.7 [ 2 ]
single action head for all degrees of freedom
✗
✗
✗
✗
Xiaomi-Robotics-1 [ 32 ]
single action head for all degrees of freedom
✗
✗
✗
✗
ABot-M0.5 [ 9 ]
action stream with mobility and manipulation sub-towers
✗
✗
✓
✗
MobileWAM [ 12 ]
action expert with routed locomotion and manipulation experts
✗
✗
✓
✗
CrossFormer [ 11 ]
shared transformer with embodiment-specific action heads
✓
✓
✗
✗
Table 1 : Comparison with recent unified mobile-manipulation models and cross-embodiment policies. The first three are VLA policies, the next two are world-action models, and the last two are cross-embodiment policies, which separate different robots rather than the navigation and manipulation phases of one mobile manipulator. Separate views: navigation and manipulation receive different camera inputs. One behavior per sample: a training sample can supervise one behavior alone without padding or masking. Video world model: the model predicts future video. Explicit terminal pose: the terminal base pose for manipulation is supervised directly.
Method
Output
Label source
Supervises approach
Scale
Reachability inversion [ 26 ]
feasible base poses
arm kinematics
✗
–
MoMa-Kitchen [ 43 ]
floor affordance map
simulation
✗
127K episodes
FloAff-Kitchen [ 45 ]
floor affordance map
simulation
✗
24.8K observations
GBPP [ 8 ]
scores of candidate base poses
rule labels and simulation trials
✗
180K rule and 12K simulation labels
N2M [ 6 ]
distribution over base poses
policy rollouts
✗
12 to 15 rollouts per task
Mobi- π [ 36 ]
base pose for a given policy
test-time optimization
✗
–
Table 2 : Comparison of MAP-Data with methods for manipulation-ready base placement. Supervises approach: the labels cover the entire approach trajectory rather than only a terminal pose or a score for a given scene. Scale is the amount of labeled data reported in each paper, and a dash denotes methods that do not learn from a dataset, such as kinematic search or test-time optimization.
Figure 2 : Overview of MAP-Data and MAP-Bench. From large-scale indoor 3D scenes, S1 curates valid target objects, S2 annotates each target with a MAP and feasible observation poses, and S3 synthesizes approach trajectories with RGB video, actions, and MAP and box labels. MAP-Bench evaluates position and heading errors from held-out initial poses.
Figure 3 : Overview of UniWAM. Navigation and manipulation samples are drawn independently, encoded by stream-specific action encoders, and packed as separate batch rows through the shared video and action DiTs. Stream-specific heads predict actions and image-plane auxiliary targets, while the video branch predicts future frames. Right: attention masks with clean action prefixes for training and inference, and the camera-frame end-effector pose shared across embodiments.
Figure 4 : Mobile manipulation platform. A shallow-angle head camera with a wide field of view provides the navigation-stream input, and a steep-angle head camera with a narrow field of view, together with the wrist cameras, provides the manipulation-stream input. In Mobile Manipulation, the navigation stream also receives the wrist cameras. The two inputs form the navigation and manipulation rows of a mixed-stream batch, which UniWAM processes in one forward pass.
π0.5
GR00T N1.7
UniWAM-Base
UniWAM-Scaled
Initialization
π0.5 base checkpoint
GR00T N1.7-3B checkpoint
Wan2.2-TI2V-5B
Wan2.2-TI2V-5B
Robot pretraining
yes
yes
none
none
Frozen modules
none
none
VAE, text encoder
VAE, text encoder
Task demonstrations
24 tasks
24 tasks
24 tasks
24 tasks
Additional data
none
none
none
MAP-Data and 402 tasks
Image-plane auxiliary outputs
✗
✗
✓
✓
Table 3 : Real-robot training settings. Steps and GPUs are given per model. Each baseline trains one MAP Navigation model per mobile robot and one model for the remaining tasks of each embodiment, whereas each UniWAM variant is one model for all tasks.
Epos (cm) ↓
Overall
Ehead ( ∘ ) ↓
Overall
Model
Short
Mid
Long
Short
Mid
Long
Vision-language-action models (VLA)
GR00T N1.7 [ 2 ]
9.48
10.8
22.0
14.1
5.66
5.42
7.46
6.18
π0.5 [ 3 ]
7.98
9.85
50.8
22.9
3.26
3.51
5.55
4.11
Vision-language navigation models (VLN)
Uni-NaVid [ 42 ]
17.3
40.0
87.2
48.2
7.98
15.0
23.4
15.5
Table 4 : MAP-Bench results over the evaluation split. Bold and underline mark the best and second-best results.
Model
MAP-Data targets
Epos Short
Epos Mid
Epos Long
Epos Overall
Ehead Overall
UniWAM-Base
0%
7.71
8.85
21.0
12.5
3.24
UniWAM
25%
7.63
8.52
18.1
11.4
2.81
UniWAM
50%
7.52
8.03
15.3
10.3
2.43
UniWAM-Scaled
100%
7.48
7.86
14.2
9.85
2.30
π0.5
0%
7.98
9.85
50.8
22.9
4.11
π0.5
100%
7.42
9.45
47.8
21.6
4.07
Table 5 : MAP-Bench errors when training with a fraction of the MAP-Data targets under a fixed compute budget. The second column is the fraction of MAP-Data targets from which the MAP-Data samples are drawn. Whenever MAP-Data is used, it is mixed with the MAP-Bench training split at a fixed 7:3 ratio, so all runs with MAP-Data see the same number of MAP-Data samples. 0% denotes training on the MAP-Bench split alone. UniWAM at 0% and 100% are UniWAM-Base and UniWAM-Scaled, and π0.5 uses the same data and budget. Position errors are in centimeters and heading errors in degrees.
MAP Navigation
Mobile Manipulation
Manipulation
Model
Epos (cm) ↓
Ehead ( ∘ ) ↓
Epos (cm) ↓
Ehead ( ∘ ) ↓
SR (%) ↑
SR (%) ↑
π0.5 [ 3 ]
23.3
5.15
26.2
5.72
62.5
71.4
GR00T N1.7 [ 2 ]
18.1
6.55
20.4
7.20
39.4
36.8
UniWAM-Base
15.1
4.91
17.1
5.30
63.1
44.3
UniWAM-Scaled
11.7
4.25
13.2
4.56
80.0
68.9
Table 6 : Real-robot results averaged over the tasks in each category. Bold and underline mark the best and second-best results.
ID
Task
π0.5
GR00T N1.7
UniWAM-Base
UniWAM-Scaled
MAP Navigation Epos (cm) ↓ / Ehead ( ∘ ) ↓
1
Water dispenser
18.9 / 4.83
19.7 / 5.47
13.9 / 4.23
10.3 / 3.86
2
Humanoid robot
24.7 / 5.17
18.4 / 6.57
15.3 / 4.84
12.6 / 4.47
3
TV cabinet
20.2 / 4.56
14.8 / 5.74
12.1 / 3.76
8.94 / 3.58
4
Tool cabinet
25.3 / 5.63
20.2 / 7.18
16.5 / 7.27
11.7 / 4.06
5
Cushion
21.8 / 4.69
15.6 / 5.93
13.2 / 4.34
9.73 / 3.74
Table 7 : Per-task real-robot results with 40 trials per task on each embodiment. MAP Navigation cells report position error and heading error averaged over the Cobot Magic and the AgiBot G2. Mobile Manipulation cells additionally report the success rate in percent, and Manipulation cells report only the success rate. Avg. is the unweighted mean over the tasks in each category. Bold and underline mark the best and second-best values for each metric.
Model
Training data
Epos (cm)
Ehead ( ∘ )
π0.5
MAP-Data only
55.1
11.4
GR00T N1.7
MAP-Data only
35.0
12.9
FastWAM
MAP-Data only
30.8
8.61
UniWAM-Sim
MAP-Data only
18.7
6.34
Reference: trained on real demonstrations
π0.5
real task data
23.3
5.15
Table 8 : Zero-shot real-robot MAP Navigation of models trained only on MAP-Data, averaged over the Cobot Magic and the AgiBot G2. The bottom rows are models trained on real demonstrations, taken from Table 6 .
Model
Mobile Manipulation SR (%)
Manipulation SR (%)
π0.5
62.5 [54.8, 69.6]
71.4 [65.9, 76.4]
GR00T N1.7
39.4 [32.1, 47.1]
36.8 [31.4, 42.6]
UniWAM-Base
63.1 [55.4, 70.2]
44.3 [38.6, 50.1]
UniWAM-Scaled
80.0 [73.1, 85.5]
68.9 [63.3, 74.1]
w/o navigation auxiliary supervision
56.9 [49.1, 64.3]
43.6 [37.9, 49.4]
w/o manipulation auxiliary supervision
58.8 [51.0, 66.1]
39.3 [33.7, 45.1]
Table 9 : Real-robot success rates with 95% Wilson intervals, pooled over 160 Mobile Manipulation trials and 280 Manipulation trials per method, matching the category means of Table 6 . Brackets give the interval bounds.
Figure 5 : Unified mobile-manipulation architectures compared in Table 10 : (a) unified action tokens, (b) separate navigation and manipulation action DiTs, (c) streams concatenated along the sequence axis, and (d) Mixed-Stream (ours), with streams in separate batch rows of shared video and action DiTs and independent stream sampling.
Alternative architectures
(d) Mixed-stream
Metric
(a) Unified tokens
(b) Sep. experts
(c) Seq. concat.
Joint
Indep.
Epos (cm) ↓
22.0
19.7
22.1
18.6
17.1
Ehead ( ∘ ) ↓
7.49
6.06
6.78
5.87
5.30
SR (%) ↑
46.9
54.4
60.0
60.0
63.1
Table 10 : Architecture and sampling comparison on the four Mobile Manipulation tasks. (a) to (c) use joint sampling. (d) uses joint or independent sampling, and the latter is UniWAM-Base as reported in Table 6 .
Epos (cm) ↓
Overall
Ehead ( ∘ ) ↓
Overall
Variant
Short
Mid
Long
Short
Mid
Long
UniWAM-Base
7.71
8.85
21.0
12.5
2.37
3.06
4.29
3.24
Velocity actions (v,ω)
8.30
9.57
30.9
16.3
3.03
3.29
6.19
4.17
w/o causal mask
8.83
10.1
29.0
16.0
3.21
3.50
5.49
4.07
w/o target box
7.74
9.29
29.5
15.5
2.43
3.02
5.14
3.53
w/o image-plane MAP
8.05
8.40
22.5
13.0
2.63
2.82
4.72
3.39
Table 11 : Ablations of UniWAM-Base on MAP-Bench. The first row is the full model.
MAP Navigation
Mobile Manipulation
Manipulation
Variant
Epos (cm) ↓
Ehead ( ∘ ) ↓
Epos (cm) ↓
Ehead ( ∘ ) ↓
SR (%) ↑
SR (%) ↑
UniWAM-Base
15.1
4.91
17.1
5.30
63.1
44.3
w/o navigation auxiliary supervision
17.6
6.84
19.9
7.21
56.9
43.6
w/o manipulation auxiliary supervision
15.4
5.18
17.8
5.61
58.8
39.3
w/o independent sampling
16.8
5.62
18.6
5.87
60.0
41.8
Table 12 : Ablations of UniWAM-Base on real robots. The first row is the full model.
MAP Navigation
Mobile Manipulation
Manipulation
Tokens per bank
Epos (cm) ↓
Ehead ( ∘ ) ↓
Epos (cm) ↓
Ehead ( ∘ ) ↓
SR (%) ↑
SR (%) ↑
0
16.2
5.38
17.8
5.48
61.9
43.6
4, ours †
15.1
4.91
17.1
5.30
63.1
44.3
8
14.8
5.02
17.3
5.24
63.8
44.6
Table 13 : Real-robot results for the number of learnable tokens per bank. † marks the default.
Target frame
Cobot Magic
Dual Franka
AgileX single arm
Mean
Robot base
55.0
28.1
62.5
48.5
Camera, ours
63.1
32.5
70.0
55.2
Table 14 : Manipulation target frame in UniWAM-Base under multi-embodiment co-training. Values are success rates in percent. Cobot Magic uses the Mobile Manipulation tasks, the dual Franka uses Manipulation tasks 1 to 3 and 7, and the AgileX single arm uses Manipulation task 6. Mean is the unweighted mean over the three reported embodiment and task groups.
Figure 6 : Real-robot executions of UniWAM. A single UniWAM model performs MAP Navigation, Mobile Manipulation, and Manipulation across four embodiments. The MAP Navigation examples dock at targets of different categories and sizes, the Mobile Manipulation examples show the transition from approach to manipulation, and the Manipulation examples show object sorting, test-tube insertion, and cup stacking.
Figure 7 : MAP-Bench trajectories of all evaluated methods on 24 randomly selected scenes. Subcaptions give the target category and initial distance. Stars mark the reference MAP, circles the start pose, and arrowheads the terminal heading. In the last two rows, every episode starts at least 3.5 m from the MAP, which falls in the long range.
Figure 8 : Navigation auxiliary predictions of UniWAM on MAP-Bench, overlaid over consecutive frames of each episode. Magenta boxes are the predicted target boxes, and cyan points are the predicted MAP positions projected into the image, both predicted jointly with the actions by the navigation stream of UniWAM at every step of the approach.
Figure 9 : Alignment between predicted 2D end-effector tracks and executed 3D actions on several tasks and two embodiments. Curves show the tracks of the two arms predicted in the previous action chunk, and circles mark the positions reached after that chunk.
Figure 10 : Failure case in the color-sorting task. From left to right, a pink and a blue cube stand next to each other, and the gripper descends toward the point between them and closes on the empty gap instead of a cube.
World action models (WAMs) built on video generation backbones are a rising recipe for robot learning, yet remain confined to tabletop manipulation. Mobile manipulation demands simultaneous locomotion and whole-body manipulation amid scene-scale dynamics, yet is still dominated by dynamics-blind visual encoders with hand-crafted coordination. We bridge this gap with MobileWAM, a mixture-of-transformers architecture that fuses a pretrained video diffusion transformer with a lightweight action expert through layerwise joint attention, translating internet-scale motion priors into whole-body control. To reconcile the heterogeneous dynamics of moving and manipulating, each feed-forward layer of the action expert becomes a three-expert mixture of shared, locomotion, and manipulation experts, softly routed by the motion intent in the action tokens. To densify supervision, we further propose Chain-of-Foresight (CoF): intermediate representations sequentially predict a chain of future latent chunks, each step conditioned on its predecessor. CoF pairs naturally with our decoupled video--action denoising scheme. At deployment, the WAM serves as a pure current-frame encoder; foresight acts only through gradients, so at inference the foresight chain and video generation are discarded, leaving only policy-level cost. MobileWAM surpasses state-of-the-art mobile manipulation policies on ManiSkill-HAB and fine-tunes to a real ARX Lift2 mobile manipulator across diverse tasks with strong generalization. Code will be released soon.
Zehua Fan, Junjie He, Wenxuan Song +14
Institute for AI Industry Research (AIR), Tsinghua University · 2Shanghai Jiao Tong University · 3The Hong Kong University of Science and Technology (Guangzhou) +8
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result, they often miss fine-grained contact dynamics, suffer from action-distribution conflicts, and accumulate errors over long-horizon rollouts. We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls. To align action space, we design a dual-level Mixture-of-Transformers architecture that disentangles both modality representations and heterogeneous action subspaces such as base movement and arm manipulation. To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction. Experiments on challenging mobile and fine-grained manipulation benchmarks demonstrate that ABot-M0.5 achieves state-of-the-art performance in both long-horizon task success and finegrained control accuracy. These results highlight the critical importance of granularity-aligned, action-disentangled, and inference-consistent world-action modeling.
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an SE(3) conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5% on LIBERO and 85.7% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7% average success across six evaluation settings, compared with 60.8% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Song Liu, Linyi Li, Yanshun Zhao +11
University of Science and Technology of China, Hefei, China. · Suzhou Artificial Intelligence Laboratory, Suzhou, China.