Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.
Figures & tables
Figure 1 : Clean-action versus velocity prediction under a common loss and a given arm–body allocation. (a) Clean chunks occupy a prescribed 13-dimensional subspace, whereas velocity and noise span the ambient space. (b) Variance of the endpoint estimate across noise draws with the clean chunk fixed. (c) Excess risk over the Bayes-optimal denoiser, as a v / x ratio. (d) Task-allocation error of sampled chunks against the number of Euler steps.
Figure 2 : MM-ABC architecture. Left: sparse VLM features condition MM-APT; a training-only future branch receives geometric supervision from frozen VGGT-Omega features. Middle: joint transformer blocks (MM-JiT in the diagram) combine stream-specific transformations with masked attention and read-only perceptual context. Right: action streams communicate within each segment, and far tokens additionally read near tokens. Future queries follow the same temporal ordering but remain isolated from actions.
Figure 3 : Overview of the MM-ABC multi-embodiment pretraining corpus. The corpus contains 5,166.2 cleaned hours from 12 datasets and 51 subsets, spanning 17 embodiments. Sector areas show training sampling probabilities aggregated by dataset; the surrounding panels illustrate environments and embodiments.
Figure 4 : Pretraining data and embodiment composition. (a) Dataset hours and duration shares. (b) Arm configuration, mobility, and control frequency.
Figure 5 : MM-30: self-collected mobile manipulation dataset. Representative tasks, object and scene diversity, and episode counts per task.
Figure 6 : MM-ABC data engine. Sources undergo semantic auditing, conversion with dataset-specific adapters, and four stages of quality filtering. Retained contiguous segments are encoded in the masked 80D interface, independently validated, and incorporated into the training mixture.
Figure 7 : Simulation benchmarks. Representative task scenes from LIBERO/LIBERO-Plus, RoboCasa365, EBench, and ManiSkill-HAB, covering fixed-base and mobile manipulation.
Figure 8 : EBench evaluation. Success rate (left) and stage-wise progress score (right) for MM-ABC and selected baselines [ 20 , 19 , 100 , 101 , 102 , 103 , 104 ] . Higher is better.
\mmabctablehead Method
Atomic-S
Composite-S
Composite-U
Average
GR00T-N1.5 [ 41 ]
60.6
35.0
33.3
43.7
Fast-WAM [ 103 ]
59.1
36.4
33.2
43.5
LingBot-VA [ 106 ]
63.5
37.3
32.1
45.1
ABot-M0.5 [ 6 ]
\secondbest 70.6
\secondbest 44.3
\secondbest 45.6
\secondbest 54.2
MM-ABC (Ours)
\best 78.4
\best 52.7
\best 50.3
\best 61.2
Table 1: RoboCasa365 target evaluation with full demonstrations . Success rate (%). S/U denote seen/unseen tasks; the average is weighted by the split sizes (18/16/16). Best: shaded; second best: bold.
\mmabctablehead
Pick
Pick
Place
Place
Open
Open
Close
\mmabctablehead Method
Apple
Bowl
Apple
Bowl
Fridge
Drawer
Drawer
Mean
DP3 ∗ [ 76 ]
0.0
20.0
31.0
32.0
0.0
0.0
68.0
21.6
ACT [ 108 ]
28.0
28.0
8.7
13.0
2.0
0.0
85.7
23.6
DP [ 18 ]
21.3
20.7
28.0
69.3
7.3
0.0
55.0
28.8
RDT-1B [ 77 ]
12.0
10.7
32.0
18.7
82.7
44.0
\best 100.0
42.9
AC-DiT ∗ [ 8 ]
33.3
36.0
33.3
17.3
90.7
81.3
\secondbest 97.3
55.6
Table 2: ManiSkill-HAB SetTable [ 9 ] . Skill and mean success rates (%). ∗ denotes depth input; – indicates an unavailable result. AnchorVLA’s mean covers six available skills. Best: shaded; second best: bold.
TidyHouse
PrepareGroceries
\mmabctablehead Method
Pick
Place
Mean
Pick
Place
Mean
ACT [ 108 ]
2.2
31.6
16.9
2.0
27.5
14.8
DP [ 18 ]
0.0
30.3
15.2
0.4
17.7
9.1
DP3 [ 76 ]
0.0
61.0
30.5
0.0
31.3
15.7
InCoM [ 1 ]
16.7
\best 78.9
47.8
15.0
\secondbest 65.9
\secondbest 40.5
GeoHAT [ 2 ]
\secondbest 30.3
\secondbest 73.3
\secondbest 51.8
\secondbest 19.7
60.7
40.2
Table 3: ManiSkill-HAB TidyHouse and PrepareGroceries [ 9 ] . Pick and Place average success rates (%) over nine object categories; Mean averages both groups. Best: shaded; second best: bold.
\mmabctablehead Method
Spatial
Object
Goal
Long
Average
OpenVLA [ 17 ]
84.7
88.4
79.2
53.7
76.5
WorldVLA [ 113 ]
87.6
96.2
83.4
60.0
81.8
π0 -FAST [ 114 ]
96.4
96.8
88.6
60.2
85.5
NORA [ 115 ]
92.2
95.4
89.4
74.6
87.9
π0 [ 19 ]
96.8
98.8
95.8
85.2
94.2
UniVLA [ 116 ]
96.5
96.8
95.6
92.0
95.2
Table 4: LIBERO. Success rate (%) with 50 trials per task. Baselines include suite-specific and shared-policy configurations. Best: shaded; second best: bold.
\mmabctablehead Method
Camera
Robot
Language
Light
Background
Noise
Layout
Total
Additional training on perturbed demonstrations
π0 + [ 19 ]
79.6
21.1
72.5
84.7
86.2
68.3
69.4
67.4
GR00T-N1.6+ [ 41 ]
\secondbest 92.6
33.5
80.1
93.6
95.4
\best 93.6
75.0
79.4
OpenVLA-OFT+ [ 112 ]
\best 92.8
30.3
85.8
94.9
93.9
89.3
77.6
79.6
Trained on original LIBERO only
OpenVLA [ 17 ]
0.8
3.5
23.0
8.1
34.8
15.2
28.5
15.6
Table 5: LIBERO-Plus. Success rate (%) across seven perturbation categories [ 112 ] . Total is the success rate over all 10,030 evaluation episodes, following the official protocol. + denotes additional training on LIBERO-Plus perturbed demonstrations. Rankings span both groups. Best: shaded; second best: bold.
\mmabctablehead Task
Interleaved
Last layer
x -pred
v -pred
MM-ABC
DeliverStraw
\secondbest 0.0
\best 0.5
\best 0.5
\best 0.5
\secondbest 0.0
GetToastedBread
0.0
\best 1.0
0.0
\best 1.0
\secondbest 0.5
KettleBoiling
33.0
29.5
35.5
\secondbest 41.5
\best 49.0
LoadDishwasher
18.5
\secondbest 28.0
22.0
27.0
\best 33.0
PackIdentical
0.5
1.5
1.5
\best 7.0
\secondbest 6.0
PreSoakPan
29.5
50.5
48.5
\secondbest 52.0
\best 67.5
Table 6: Component ablations on RoboCasa365 composite-seen tasks. Success rate (%) without robot-data pretraining. Average is the unweighted mean across 16 tasks. Rankings are computed within each row, across the five variants. Best: shaded; second best: bold.
Figure 9 : Real-world task execution. Representative stages of the five tasks, from top to bottom: Birthday Party Setup, Office Folder Arrangement, Kitchen Work, Industrial Parts Organization, and Chemistry Lab Operation. Within each task, frames progress from left to right.
Figure 10 : Real-world task success. Success rate (%) over 20 trials per task for each method.
Mobile manipulation is a fundamental capability for general-purpose robotic agents, requiring both coordinated control of the mobile base and manipulator and robust perception under dynamically changing viewpoints. However, existing approaches face two key challenges: strong coupling between base and arm actions complicates control optimization, and perceptual attention is often poorly allocated as viewpoints shift during mobile manipulation. We propose InCoM, an intent-driven perception and structured coordination framework for mobile manipulation. InCoM infers latent motion intent to dynamically reweight multi-scale perceptual features, enabling stage-adaptive allocation of perceptual attention. To support robust cross-modal perception, InCoM further incorporates a geometric-semantic structured alignment mechanism that enhances multimodal correspondence. On the control side, we design a decoupled coordinated flow matching action decoder that explicitly models coordinated base-arm action generation, alleviating optimization difficulties caused by control coupling. Experimental results demonstrate that InCoM significantly outperforms state-of-the-art methods, achieving success rate gains of 28.2%, 26.1%, and 23.6% across three ManiSkill-HAB scenarios without privileged information. Furthermore, its effectiveness is consistently validated in real-world mobile manipulation tasks, where InCoM maintains a superior success rate over existing baselines.
Jiahao Liu, Cui Wenbo, Zhongpu Xia +3
Institute of Automation, Chinese Academy of Sciences · School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences · Anyverse Dynamics
Mobile manipulation requires perceptual evidence at different spatial scales for base motion and arm control, while the two action modalities remain kinematically coupled. Existing policies often employ specialized action generation for different subsystems but condition heterogeneous action branches on a shared perceptual representation, leaving subsystem-specific perception-action correspondence implicit. We present MoPA, a framework that aligns perceptual conditioning with mobility and manipulation while preserving coordination at the action level. Dual Perceptual Streams employ two mutually masked query banks to extract separate perceptual representations from a shared vision-language context. Perception2Action Adaptation jointly updates each query bank and its corresponding action stream at every layer of a structured Mixture-of-Transformers decoder, while enabling information exchange between the two action streams. Coupled conditional flow matching learns a joint vector field for coordinated generation of both action chunks. On the ManiSkill-HAB benchmark, MoPA achieves state-of-the-art performance across all three task suites. Across four real-world tasks, MoPA achieves a mean full-task success rate of 76.3%, outperforming the best baseline by 12.5 percentage points. Ablation studies and further analyses validate the effectiveness of the proposed design. Website is available at: https://mopa-policy.github.io/.
Mobile manipulation is a key capability for general-purpose robots, yet remains challenging for current embodied learning methods. VLA policies are typically reactive and lack explicit world modeling, while existing World Action Models (WAMs) are still poorly aligned with the structure of mobile manipulation: they operate on coarse video chunks, model entangled navigation-manipulation actions, and train inverse dynamics under supervision that does not match autoregressive inference. As a result, they often miss fine-grained contact dynamics, suffer from action-distribution conflicts, and accumulate errors over long-horizon rollouts. We propose ABot-M0.5, a new WAM built on the insight that mobile manipulation requires alignment at three levels: temporal granularity, action space, and train-test consistency. To align temporal granularity, we introduce intermediate latent actions that capture local visual state transitions and serve as an bridging action space between video latents and embodiment-specific controls. To align action space, we design a dual-level Mixture-of-Transformers architecture that disentangles both modality representations and heterogeneous action subspaces such as base movement and arm manipulation. To align inference conditions, we propose the dream-forcing training strategy that progressively trains inverse dynamics on model-predicted videos, improving train-test alignment and robustness during autoregressive prediction. Experiments on challenging mobile and fine-grained manipulation benchmarks demonstrate that ABot-M0.5 achieves state-of-the-art performance in both long-horizon task success and finegrained control accuracy. These results highlight the critical importance of granularity-aligned, action-disentangled, and inference-consistent world-action modeling.