cs.ROSep 30, 2026

UniWAM Technical Report: Unified Mobile Manipulation via Mixed-Stream World-Action Modeling and Manipulation Anchor Pose Supervision

Authors: Wei Xue, Keliang Liu, Mingzhang Cui, Jinhua Xie, Jinjie Wei, Jianan Hou, Jingcheng Lu, Lintao Wang, +9 more

Organizations: Physical Superintelligence Lab, Fysics AI · College of Intelligent Robotics and Advanced Manufacturing, Fudan University

Abstract

Mobile manipulation requires precise navigation to a manipulation-ready pose followed by reliable object interaction. These two stages differ in action spaces and visual requirements, which complicates unified policy learning. In addition, collecting diverse real-world navigation data with explicit manipulation-ready pose supervision remains costly and difficult to scale. We introduce UniWAM, a unified mixed-stream world-action model with separate action encoders and output heads for navigation and manipulation, sharing a common backbone. This design supports joint representation learning on independently sampled navigation and manipulation data. UniWAM supports independent inference for either stream and batch-parallel inference for both. We further introduce Manipulation Anchor Pose (MAP) supervision for where to stop and how to orient for manipulation. An automated pipeline constructs MAP-Data from large-scale 3D scenes, yielding over 1.5 million episodes and 7,500 hours. MAP-Data provides per-frame target-object bounding boxes and image-plane MAP coordinates as auxiliary navigation supervision. Together with projected end-effector trajectories for manipulation, these prediction targets provide stream-specific image-plane supervision for action learning from egocentric observations. With large-scale MAP-Data, UniWAM outperforms the strongest external baselines on our MAP-Bench by 30.1% in position error and 44.0% in heading error. Across 24 real-robot tasks, UniWAM achieves leading results in MAP navigation and mobile manipulation, with competitive manipulation performance. We have released code, data, and benchmark.

Figures & tables

Explore similar work

CardsList
  1. MobileWAM: Bridging World Action Models to Mobile Manipulation with Chain-of-Foresight

    Aug 5, 2026Zehua Fan, Junjie He, Wenxuan Song +14World ModelsKinematic Priors

  2. ABot-M0.5: Unified Mobility-and-Manipulation World Action Model

    Jul 1, 2026Ronghan Chen, Yandan Yang, Zuojin Tang +18Efficient World-Action ModelRecent World-Action Models

  3. UMR: Universal Manipulation Representation

    Sep 28, 2026Song Liu, Linyi Li, Yanshun Zhao +11Cross-Embodiment TransferUnified Representation