This paper presents Adaptive Whole-body Loco-Manipulation, AdaptManip, a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery. Unlike prior imitation learning-based approaches that rely on human demonstrations and are often brittle to disturbances, AdaptManip aims to train a robust loco-manipulation policy via reinforcement learning without human demonstrations or teleoperation data. The proposed framework consists of three coupled components: (1) a recurrent object state estimator that tracks the manipulated object in real time under limited field-of-view and occlusions; (2) a whole-body base policy for robust locomotion with residual manipulation control for stable object lifting and delivery; and (3) a LiDAR-based robot global position estimator that provides drift-robust localization. All components are trained in simulation using reinforcement learning and deployed on real hardware in a zero-shot manner. Experimental results show that AdaptManip significantly outperforms baseline methods, including imitation learning-based approaches, in adaptability and overall success rate, while the learned estimator keeps tracking the object when visual observations are intermittent. We further demonstrate fully autonomous real-world navigation, object lifting, and delivery on a humanoid robot.
Figures & tables
Fig. 1 : Fully autonomous humanoid loco-manipulation using online recurrent state estimation. (1) navigating toward the object, (2) lifting the object through coordinated whole-body motion, and (3) delivering the object to the target location. Our method relies solely on onboard sensing and does not require teleoperation data or an external mocap system.
Method
Onbd
NoHumRef
LocoMan
NoFutRef
NoTeleOp
TWIST [ 7 ]
✗
✗
✓
✓
✓
ResMimic [ 4 ]
✗
✗
✓
✓
✓
VisualMimic [ 3 ]
✓
✗
✓
✓
✗
HDMI [ 10 ]
✗
✗
✓
✓
✓
PhysHSI [ 11 ]
✗
✗
✓
✓
✓
OmniRetarget [ 5 ]
✗
✗
✓
✓
✓
TABLE I : Comparison of representative methods across key aspects: onboard-only sensing (Onbd), absence of human demonstrations (NoHumRef), locomotion–manipulation capability (LocoMan), absence of future references (NoFutRef), and no teleoperation (NoTeleOp).
Fig. 2 : Three-stage AdaptManip experiment plan and deployment. Stage 1: LiDAR odometry and proprioception enable autonomous navigation. Stage 2: Recurrent multimodal object-pose estimation supports coordinated lifting. Stage 3: Image-based refinement and residual policies ensure stable delivery. All stages operate using only onboard sensing.
Fig. 3 : Overview of the training and deployment pipeline. (1) A base whole-body control policy πwbc is trained in IsaacLab to generate base whole-body behavior such as walking. (2) A manipulation residual policy πres is trained on top of the base policy, taking proprioception and the estimated object state X^box to produce residual actions Δat . The residual action aims to adaptively lift a 3D object. (3) A recurrent online object state estimator fuses vision and proprioceptive cues using a V-LSTM and MLP to infer X^box , and is trained jointly with the residual manipulation policy. During real-world deployment, the robot uses onboard estimators and LiDAR odometry and executes the combined policies πwbc and πres to complete the whole-body loco-manipulation task.
Parameter
Range
Operation
Base Mass [kg]
[-2.5, 2.5]
Add
kp
[ 0.8, 1.2]
Scale
kd
[ 0.8, 1.2]
Scale
Ground Static Friction
[ 0.3, 1.5]
Absolute
Ground Dynamic Friction
[ 0.3, 0.9]
Absolute
Base Force Disturbance [N]
[-4.0, 4.0]
Absolute
TABLE II : Domain randomization parameters for policy training.
Method
Whole
Stage1
Stage2
Drops ( ↓ )
Regrasps ( ↑ )
IsaacLab
Pure RL
0.62 (±0.48)
0.98 (±0.14)
0.88 (±0.32)
1.79 (±2.39)
5.74 (±3.91)
Pure RL + FK
0.88 (±0.32)
0.93 (±0.25)
0.92 (±0.27)
0.29 (±0.89)
2.94 (±1.88)
Motion Imitation
0.42 (±0.46)
0.96 (±0.19)
0.93 (±0.18)
3.81 (±4.77)
2.94 (±1.99)
AdaptManip (Ours)
0.85 (±0.35)
0.97 (±0.17)
0.92 (±0.26)
0.49 (±1.14)
2.12 (±1.95)
Pure Vision
0.63 (±0.42)
0.95 (±0.22)
0.91 (±0.20)
12.44 (±9.36)
23.55 (±15.32)
TABLE III : Performance comparison between AdaptManip and baselines across 135 trials in MuJoCo and IsaacLab.
Fig. 4 : State estimation error of our method. Shows mean ± 1 standard deviation across 50 episodes. The green region shows the area where vision is available, and the purple region shows the area where there is contact between the robot and box.
Fig. 5 : Hardware demonstration of the three-stage whole-body loco-manipulation task.
Fig. 6 : Hardware demonstration of whole-body loco-manipulation. The yellow region indicates the grasp formation phase where the robot carefully coordinates its arms for a secure hold, while the red region highlights the robot’s ability to recover from a transient loss of stability through corrective arm motions.
Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.
Zejie Tian, Ruibing Hou, Bingpeng Ma +2
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China.
Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.
Humanoid loco-manipulation requires stable whole-body control under varying object masses and pickup/placement heights. This becomes particularly challenging in sim-to-real transfer, where object-induced load variation and robot-side dynamics mismatch interact during physical contact. Existing history-based adapters often compress these factors into a single latent representation, which can weaken robustness under heavy-load manipulation. We propose \textbf{SplitAdapter: Load-Aware Humanoid Loco-Manipulation via Factorized Adaptation}, which freezes a pretrained box manipulation policy and extends it with object/load and dynamics-aware context encoders trained with split world-model objectives, GRL-based cross-adversarial regularization, and hierarchical Feature-wise Linear Modulation (FiLM). In sim-to-sim experiments and real-world deployment, SplitAdapter improves Full-task success over the base policy and world-model FiLM baselines across object masses of 2, 4, and 6 kg and pickup/placement heights of 0, 30, and 60 cm, with the largest improvements under heavy-load conditions.