Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.
Figures & tables
Figure 1 : Clean-action versus velocity prediction under a common loss and a given arm–body allocation. (a) Clean chunks occupy a prescribed 13-dimensional subspace, whereas velocity and noise span the ambient space. (b) Variance of the endpoint estimate across noise draws with the clean chunk fixed. (c) Excess risk over the Bayes-optimal denoiser, as a v / x ratio. (d) Task-allocation error of sampled chunks against the number of Euler steps.
Figure 2 : MM-ABC architecture. Left: sparse VLM features condition MM-APT; a training-only future branch receives geometric supervision from frozen VGGT-Omega features. Middle: joint transformer blocks (MM-JiT in the diagram) combine stream-specific transformations with masked attention and read-only perceptual context. Right: action streams communicate within each segment, and far tokens additionally read near tokens. Future queries follow the same temporal ordering but remain isolated from actions.
Figure 3 : Overview of the MM-ABC multi-embodiment pretraining corpus. The corpus contains 5,166.2 cleaned hours from 12 datasets and 51 subsets, spanning 17 embodiments. Sector areas show training sampling probabilities aggregated by dataset; the surrounding panels illustrate environments and embodiments.
Figure 4 : Pretraining data and embodiment composition. (a) Dataset hours and duration shares. (b) Arm configuration, mobility, and control frequency.
Figure 5 : MM-30: self-collected mobile manipulation dataset. Representative tasks, object and scene diversity, and episode counts per task.
Figure 6 : MM-ABC data engine. Sources undergo semantic auditing, conversion with dataset-specific adapters, and four stages of quality filtering. Retained contiguous segments are encoded in the masked 80D interface, independently validated, and incorporated into the training mixture.
Figure 7 : Simulation benchmarks. Representative task scenes from LIBERO/LIBERO-Plus, RoboCasa365, EBench, and ManiSkill-HAB, covering fixed-base and mobile manipulation.
Figure 8 : EBench evaluation. Success rate (left) and stage-wise progress score (right) for MM-ABC and selected baselines [ 20 , 19 , 100 , 101 , 102 , 103 , 104 ] . Higher is better.
\mmabctablehead Method
Atomic-S
Composite-S
Composite-U
Average
GR00T-N1.5 [ 41 ]
60.6
35.0
33.3
43.7
Fast-WAM [ 103 ]
59.1
36.4
33.2
43.5
LingBot-VA [ 106 ]
63.5
37.3
32.1
45.1
ABot-M0.5 [ 6 ]
\secondbest 70.6
\secondbest 44.3
\secondbest 45.6
\secondbest 54.2
MM-ABC (Ours)
\best 78.4
\best 52.7
\best 50.3
\best 61.2
Table 1: RoboCasa365 target evaluation with full demonstrations . Success rate (%). S/U denote seen/unseen tasks; the average is weighted by the split sizes (18/16/16). Best: shaded; second best: bold.
\mmabctablehead
Pick
Pick
Place
Place
Open
Open
Close
\mmabctablehead Method
Apple
Bowl
Apple
Bowl
Fridge
Drawer
Drawer
Mean
DP3 ∗ [ 76 ]
0.0
20.0
31.0
32.0
0.0
0.0
68.0
21.6
ACT [ 108 ]
28.0
28.0
8.7
13.0
2.0
0.0
85.7
23.6
DP [ 18 ]
21.3
20.7
28.0
69.3
7.3
0.0
55.0
28.8
RDT-1B [ 77 ]
12.0
10.7
32.0
18.7
82.7
44.0
\best 100.0
42.9
AC-DiT ∗ [ 8 ]
33.3
36.0
33.3
17.3
90.7
81.3
\secondbest 97.3
55.6
Table 2: ManiSkill-HAB SetTable [ 9 ] . Skill and mean success rates (%). ∗ denotes depth input; – indicates an unavailable result. AnchorVLA’s mean covers six available skills. Best: shaded; second best: bold.
TidyHouse
PrepareGroceries
\mmabctablehead Method
Pick
Place
Mean
Pick
Place
Mean
ACT [ 108 ]
2.2
31.6
16.9
2.0
27.5
14.8
DP [ 18 ]
0.0
30.3
15.2
0.4
17.7
9.1
DP3 [ 76 ]
0.0
61.0
30.5
0.0
31.3
15.7
InCoM [ 1 ]
16.7
\best 78.9
47.8
15.0
\secondbest 65.9
\secondbest 40.5
GeoHAT [ 2 ]
\secondbest 30.3
\secondbest 73.3
\secondbest 51.8
\secondbest 19.7
60.7
40.2
Table 3: ManiSkill-HAB TidyHouse and PrepareGroceries [ 9 ] . Pick and Place average success rates (%) over nine object categories; Mean averages both groups. Best: shaded; second best: bold.
\mmabctablehead Method
Spatial
Object
Goal
Long
Average
OpenVLA [ 17 ]
84.7
88.4
79.2
53.7
76.5
WorldVLA [ 113 ]
87.6
96.2
83.4
60.0
81.8
π0 -FAST [ 114 ]
96.4
96.8
88.6
60.2
85.5
NORA [ 115 ]
92.2
95.4
89.4
74.6
87.9
π0 [ 19 ]
96.8
98.8
95.8
85.2
94.2
UniVLA [ 116 ]
96.5
96.8
95.6
92.0
95.2
Table 4: LIBERO. Success rate (%) with 50 trials per task. Baselines include suite-specific and shared-policy configurations. Best: shaded; second best: bold.
\mmabctablehead Method
Camera
Robot
Language
Light
Background
Noise
Layout
Total
Additional training on perturbed demonstrations
π0 + [ 19 ]
79.6
21.1
72.5
84.7
86.2
68.3
69.4
67.4
GR00T-N1.6+ [ 41 ]
\secondbest 92.6
33.5
80.1
93.6
95.4
\best 93.6
75.0
79.4
OpenVLA-OFT+ [ 112 ]
\best 92.8
30.3
85.8
94.9
93.9
89.3
77.6
79.6
Trained on original LIBERO only
OpenVLA [ 17 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
Table 5: LIBERO-Plus. Success rate (%) across seven perturbation categories [ 112 ] . Total is the success rate over all 10,030 evaluation episodes, following the official protocol. + denotes additional training on LIBERO-Plus perturbed demonstrations. Rankings span both groups. Best: shaded; second best: bold.
\mmabctablehead Task
Interleaved
Last layer
x -pred
v -pred
MM-ABC
DeliverStraw
\secondbest 0.0
\best 0.5
\best 0.5
\best 0.5
\secondbest 0.0
GetToastedBread
0.0
\best 1.0
0.0
\best 1.0
\secondbest 0.5
KettleBoiling
33.0
29.5
35.5
\secondbest 41.5
\best 49.0
LoadDishwasher
18.5
\secondbest 28.0
22.0
27.0
\best 33.0
PackIdentical
0.5
1.5
1.5
\best 7.0
\secondbest 6.0
PreSoakPan
29.5
50.5
48.5
\secondbest 52.0
\best 67.5
Table 6: Component ablations on RoboCasa365 composite-seen tasks. Success rate (%) without robot-data pretraining. Average is the unweighted mean across 16 tasks. Rankings are computed within each row, across the five variants. Best: shaded; second best: bold.
Figure 9 : Real-world task execution. Representative stages of the five tasks, from top to bottom: Birthday Party Setup, Office Folder Arrangement, Kitchen Work, Industrial Parts Organization, and Chemistry Lab Operation. Within each task, frames progress from left to right.
Figure 10 : Real-world task success. Success rate (%) over 20 trials per task for each method.
Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Anyverse Dynamics