Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1% in position error and 44.0% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.
Figures & tables
Figure 1 : Challenges of mobile manipulation and our approach. Existing unified policies either share one action representation across navigation and manipulation or separate the two only inside the action module, and they learn where to stop only implicitly from costly complete trajectories. UniWAM models navigation and manipulation as separate streams over a shared backbone, and MAP-Data supervises the manipulation-ready pose at scale with automatically generated trajectories.
Model
Action generator
Separate views
One behavior per sample
Video world model
Explicit terminal pose
π0.5 [ 3 ]
single action expert for all degrees of freedom
✗
✗
✗
✗
GR00T N1.7 [ 2 ]
single action head for all degrees of freedom
✗
✗
✗
✗
Xiaomi-Robotics-1 [ 32 ]
single action head for all degrees of freedom
✗
✗
✗
✗
ABot-M0.5 [ 9 ]
action stream with mobility and manipulation sub-towers
✗
✗
✓
✗
MobileWAM [ 12 ]
action expert with routed locomotion and manipulation experts
✗
✗
✓
✗
CrossFormer [ 11 ]
shared transformer with embodiment-specific action heads
✓
✓
✗
✗
Table 1 : Comparison with recent unified mobile-manipulation models and cross-embodiment policies. The first three are VLA policies, the next two are world-action models, and the last two are cross-embodiment policies, which separate different robots rather than the navigation and manipulation phases of one mobile manipulator. Separate views: navigation and manipulation receive different camera inputs. One behavior per sample: a training sample can supervise one behavior alone without padding or masking. Video world model: the model predicts future video. Explicit terminal pose: the terminal base pose for manipulation is supervised directly.
Method
Output
Label source
Supervises approach
Scale
Reachability inversion [ 26 ]
feasible base poses
arm kinematics
✗
–
MoMa-Kitchen [ 43 ]
floor affordance map
simulation
✗
127K episodes
FloAff-Kitchen [ 45 ]
floor affordance map
simulation
✗
24.8K observations
GBPP [ 8 ]
scores of candidate base poses
rule labels and simulation trials
✗
180K rule and 12K simulation labels
N2M [ 6 ]
distribution over base poses
policy rollouts
✗
12 to 15 rollouts per task
Mobi- π [ 36 ]
base pose for a given policy
test-time optimization
✗
–
Table 2 : Comparison of MAP-Data with methods for manipulation-ready base placement. Supervises approach: the labels cover the entire approach trajectory rather than only a terminal pose or a score for a given scene. Scale is the amount of labeled data reported in each paper, and a dash denotes methods that do not learn from a dataset, such as kinematic search or test-time optimization.
Figure 2 : Overview of MAP-Data and MAP-Bench. From large-scale indoor 3D scenes, S1 curates valid target objects, S2 annotates each target with a MAP and feasible observation poses, and S3 synthesizes approach trajectories with RGB video, actions, and MAP and box labels. MAP-Bench evaluates position and heading errors from held-out initial poses.
Figure 3 : Overview of UniWAM. Navigation and manipulation samples are drawn independently, encoded by stream-specific action encoders, and packed as separate batch rows through the shared video and action DiTs. Stream-specific heads predict actions and image-plane auxiliary targets, while the video branch predicts future frames. Right: attention masks with clean action prefixes for training and inference, and the camera-frame end-effector pose shared across embodiments.
Figure 4 : Mobile manipulation platform. A shallow-angle head camera with a wide field of view provides the navigation-stream input, and a steep-angle head camera with a narrow field of view, together with the wrist cameras, provides the manipulation-stream input. In Mobile Manipulation, the navigation stream also receives the wrist cameras. The two inputs form the navigation and manipulation rows of a mixed-stream batch, which UniWAM processes in one forward pass.
π0.5
GR00T N1.7
UniWAM-Base
UniWAM-Scaled
Initialization
π0.5 base checkpoint
GR00T N1.7-3B checkpoint
Wan2.2-TI2V-5B
Wan2.2-TI2V-5B
Robot pretraining
yes
yes
none
none
Frozen modules
none
none
VAE, text encoder
VAE, text encoder
Task demonstrations
24 tasks
24 tasks
24 tasks
24 tasks
Additional data
none
none
none
MAP-Data and 402 tasks
Image-plane auxiliary outputs
✗
✗
✓
✓
Table 3 : Real-robot training settings. Steps and GPUs are given per model. Each baseline trains one MAP Navigation model per mobile robot and one model for the remaining tasks of each embodiment, whereas each UniWAM variant is one model for all tasks.
Epos (cm) ↓
Overall
Ehead ( ∘ ) ↓
Overall
Model
Short
Mid
Long
Short
Mid
Long
Vision-language-action models (VLA)
GR00T N1.7 [ 2 ]
9.48
10.8
22.0
14.1
5.66
5.42
7.46
6.18
π0.5 [ 3 ]
7.98
9.85
50.8
22.9
3.26
3.51
5.55
4.11
Vision-language navigation models (VLN)
Uni-NaVid [ 42 ]
17.3
40.0
87.2
48.2
7.98
15.0
23.4
15.5
Table 4 : MAP-Bench results over the evaluation split. Bold and underline mark the best and second-best results.
Model
MAP-Data targets
Epos Short
Epos Mid
Epos Long
Epos Overall
Ehead Overall
UniWAM-Base
0%
7.71
8.85
21.0
12.5
3.24
UniWAM
25%
7.63
8.52
18.1
11.4
2.81
UniWAM
50%
7.52
8.03
15.3
10.3
2.43
UniWAM-Scaled
100%
7.48
7.86
14.2
9.85
2.30
π0.5
0%
7.98
9.85
50.8
22.9
4.11
π0.5
100%
7.42
9.45
47.8
21.6
4.07
Table 5 : MAP-Bench errors when training with a fraction of the MAP-Data targets under a fixed compute budget. The second column is the fraction of MAP-Data targets from which the MAP-Data samples are drawn. Whenever MAP-Data is used, it is mixed with the MAP-Bench training split at a fixed 7:3 ratio, so all runs with MAP-Data see the same number of MAP-Data samples. 0% denotes training on the MAP-Bench split alone. UniWAM at 0% and 100% are UniWAM-Base and UniWAM-Scaled, and π0.5 uses the same data and budget. Position errors are in centimeters and heading errors in degrees.
MAP Navigation
Mobile Manipulation
Manipulation
Model
Epos (cm) ↓
Ehead ( ∘ ) ↓
Epos (cm) ↓
Ehead ( ∘ ) ↓
SR (%) ↑
SR (%) ↑
π0.5 [ 3 ]
23.3
5.15
26.2
5.72
62.5
71.4
GR00T N1.7 [ 2 ]
18.1
6.55
20.4
7.20
39.4
36.8
UniWAM-Base
15.1
4.91
17.1
5.30
63.1
44.3
UniWAM-Scaled
11.7
4.25
13.2
4.56
80.0
68.9
Table 6 : Real-robot results averaged over the tasks in each category. Bold and underline mark the best and second-best results.
ID
Task
π0.5
GR00T N1.7
UniWAM-Base
UniWAM-Scaled
MAP Navigation Epos (cm) ↓ / Ehead ( ∘ ) ↓
1
Water dispenser
18.9 / 4.83
19.7 / 5.47
13.9 / 4.23
10.3 / 3.86
2
Humanoid robot
24.7 / 5.17
18.4 / 6.57
15.3 / 4.84
12.6 / 4.47
3
TV cabinet
20.2 / 4.56
14.8 / 5.74
12.1 / 3.76
8.94 / 3.58
4
Tool cabinet
25.3 / 5.63
20.2 / 7.18
16.5 / 7.27
11.7 / 4.06
5
Cushion
21.8 / 4.69
15.6 / 5.93
13.2 / 4.34
9.73 / 3.74
Table 7 : Per-task real-robot results with 40 trials per task on each embodiment. MAP Navigation cells report position error and heading error averaged over the Cobot Magic and the AgiBot G2. Mobile Manipulation cells additionally report the success rate in percent, and Manipulation cells report only the success rate. Avg. is the unweighted mean over the tasks in each category. Bold and underline mark the best and second-best values for each metric.
Model
Training data
Epos (cm)
Ehead ( ∘ )
π0.5
MAP-Data only
55.1
11.4
GR00T N1.7
MAP-Data only
35.0
12.9
FastWAM
MAP-Data only
30.8
8.61
UniWAM-Sim
MAP-Data only
18.7
6.34
Reference: trained on real demonstrations
π0.5
real task data
23.3
5.15
Table 8 : Zero-shot real-robot MAP Navigation of models trained only on MAP-Data, averaged over the Cobot Magic and the AgiBot G2. The bottom rows are models trained on real demonstrations, taken from Table 6 .
Model
Mobile Manipulation SR (%)
Manipulation SR (%)
π0.5
62.5 [54.8, 69.6]
71.4 [65.9, 76.4]
GR00T N1.7
39.4 [32.1, 47.1]
36.8 [31.4, 42.6]
UniWAM-Base
63.1 [55.4, 70.2]
44.3 [38.6, 50.1]
UniWAM-Scaled
80.0 [73.1, 85.5]
68.9 [63.3, 74.1]
w/o navigation auxiliary supervision
56.9 [49.1, 64.3]
43.6 [37.9, 49.4]
w/o manipulation auxiliary supervision
58.8 [51.0, 66.1]
39.3 [33.7, 45.1]
Table 9 : Real-robot success rates with 95% Wilson intervals, pooled over 160 Mobile Manipulation trials and 280 Manipulation trials per method, matching the category means of Table 6 . Brackets give the interval bounds.
Figure 5 : Unified mobile-manipulation architectures compared in Table 10 : (a) unified action tokens, (b) separate navigation and manipulation action DiTs, (c) streams concatenated along the sequence axis, and (d) Mixed-Stream (ours), with streams in separate batch rows of shared video and action DiTs and independent stream sampling.
Alternative architectures
(d) Mixed-stream
Metric
(a) Unified tokens
(b) Sep. experts
(c) Seq. concat.
Joint
Indep.
Epos (cm) ↓
22.0
19.7
22.1
18.6
17.1
Ehead ( ∘ ) ↓
7.49
6.06
6.78
5.87
5.30
SR (%) ↑
46.9
54.4
60.0
60.0
63.1
Table 10 : Architecture and sampling comparison on the four Mobile Manipulation tasks. (a) to (c) use joint sampling. (d) uses joint or independent sampling, and the latter is UniWAM-Base as reported in Table 6 .
Epos (cm) ↓
Overall
Ehead ( ∘ ) ↓
Overall
Variant
Short
Mid
Long
Short
Mid
Long
UniWAM-Base
7.71
8.85
21.0
12.5
2.37
3.06
4.29
3.24
Velocity actions (v,ω)
8.30
9.57
30.9
16.3
3.03
3.29
6.19
4.17
w/o causal mask
8.83
10.1
29.0
16.0
3.21
3.50
5.49
4.07
w/o target box
7.74
9.29
29.5
15.5
2.43
3.02
5.14
3.53
w/o image-plane MAP
8.05
8.40
22.5
13.0
2.63
2.82
4.72
3.39
Table 11 : Ablations of UniWAM-Base on MAP-Bench. The first row is the full model.
MAP Navigation
Mobile Manipulation
Manipulation
Variant
Epos (cm) ↓
Ehead ( ∘ ) ↓
Epos (cm) ↓
Ehead ( ∘ ) ↓
SR (%) ↑
SR (%) ↑
UniWAM-Base
15.1
4.91
17.1
5.30
63.1
44.3
w/o navigation auxiliary supervision
17.6
6.84
19.9
7.21
56.9
43.6
w/o manipulation auxiliary supervision
15.4
5.18
17.8
5.61
58.8
39.3
w/o independent sampling
16.8
5.62
18.6
5.87
60.0
41.8
Table 12 : Ablations of UniWAM-Base on real robots. The first row is the full model.
MAP Navigation
Mobile Manipulation
Manipulation
Tokens per bank
Epos (cm) ↓
Ehead ( ∘ ) ↓
Epos (cm) ↓
Ehead ( ∘ ) ↓
SR (%) ↑
SR (%) ↑
0
16.2
5.38
17.8
5.48
61.9
43.6
4, ours †
15.1
4.91
17.1
5.30
63.1
44.3
8
14.8
5.02
17.3
5.24
63.8
44.6
Table 13 : Real-robot results for the number of learnable tokens per bank. † marks the default.
Target frame
Cobot Magic
Dual Franka
AgileX single arm
Mean
Robot base
55.0
28.1
62.5
48.5
Camera, ours
63.1
32.5
70.0
55.2
Table 14 : Manipulation target frame in UniWAM-Base under multi-embodiment co-training. Values are success rates in percent. Cobot Magic uses the Mobile Manipulation tasks, the dual Franka uses Manipulation tasks 1 to 3 and 7, and the AgileX single arm uses Manipulation task 6. Mean is the unweighted mean over the three reported embodiment and task groups.
Figure 6 : Real-robot executions of UniWAM. A single UniWAM model performs MAP Navigation, Mobile Manipulation, and Manipulation across four embodiments. The MAP Navigation examples dock at targets of different categories and sizes, the Mobile Manipulation examples show the transition from approach to manipulation, and the Manipulation examples show object sorting, test-tube insertion, and cup stacking.
Figure 7 : MAP-Bench trajectories of all evaluated methods on 24 randomly selected scenes. Subcaptions give the target category and initial distance. Stars mark the reference MAP, circles the start pose, and arrowheads the terminal heading. In the last two rows, every episode starts at least 3.5 m from the MAP, which falls in the long range.
Figure 8 : Navigation auxiliary predictions of UniWAM on MAP-Bench, overlaid over consecutive frames of each episode. Magenta boxes are the predicted target boxes, and cyan points are the predicted MAP positions projected into the image, both predicted jointly with the actions by the navigation stream of UniWAM at every step of the approach.
Figure 9 : Alignment between predicted 2D end-effector tracks and executed 3D actions on several tasks and two embodiments. Curves show the tracks of the two arms predicted in the previous action chunk, and circles mark the positions reached after that chunk.
Figure 10 : Failure case in the color-sorting task. From left to right, a pink and a blue cube stand next to each other, and the gripper descends toward the point between them and closes on the empty gap instead of a cube.
Institute for AI Industry Research (AIR), Tsinghua University · 2Shanghai Jiao Tong University · 3The Hong Kong University of Science and Technology (Guangzhou) +8