This paper presents Adaptive Whole-body Loco-Manipulation, AdaptManip, a fully autonomous framework for humanoid robots to perform integrated navigation, object lifting, and delivery. Unlike prior imitation learning-based approaches that rely on human demonstrations and are often brittle to disturbances, AdaptManip aims to train a robust loco-manipulation policy via reinforcement learning without human demonstrations or teleoperation data. The proposed framework consists of three coupled components: (1) a recurrent object state estimator that tracks the manipulated object in real time under limited field-of-view and occlusions; (2) a whole-body base policy for robust locomotion with residual manipulation control for stable object lifting and delivery; and (3) a LiDAR-based robot global position estimator that provides drift-robust localization. All components are trained in simulation using reinforcement learning and deployed on real hardware in a zero-shot manner. Experimental results show that AdaptManip significantly outperforms baseline methods, including imitation learning-based approaches, in adaptability and overall success rate, while the learned estimator keeps tracking the object when visual observations are intermittent. We further demonstrate fully autonomous real-world navigation, object lifting, and delivery on a humanoid robot.
Figures & tables
Fig. 1 : Fully autonomous humanoid loco-manipulation using online recurrent state estimation. (1) navigating toward the object, (2) lifting the object through coordinated whole-body motion, and (3) delivering the object to the target location. Our method relies solely on onboard sensing and does not require teleoperation data or an external mocap system.
Method
Onbd
NoHumRef
LocoMan
NoFutRef
NoTeleOp
TWIST [ 7 ]
✗
✗
✓
✓
✓
ResMimic [ 4 ]
✗
✗
✓
✓
✓
VisualMimic [ 3 ]
✓
✗
✓
✓
✗
HDMI [ 10 ]
✗
✗
✓
✓
✓
PhysHSI [ 11 ]
✗
✗
✓
✓
✓
OmniRetarget [ 5 ]
✗
✗
✓
✓
✓
TABLE I : Comparison of representative methods across key aspects: onboard-only sensing (Onbd), absence of human demonstrations (NoHumRef), locomotion–manipulation capability (LocoMan), absence of future references (NoFutRef), and no teleoperation (NoTeleOp).
Fig. 2 : Three-stage AdaptManip experiment plan and deployment. Stage 1: LiDAR odometry and proprioception enable autonomous navigation. Stage 2: Recurrent multimodal object-pose estimation supports coordinated lifting. Stage 3: Image-based refinement and residual policies ensure stable delivery. All stages operate using only onboard sensing.
Fig. 3 : Overview of the training and deployment pipeline. (1) A base whole-body control policy πwbc is trained in IsaacLab to generate base whole-body behavior such as walking. (2) A manipulation residual policy πres is trained on top of the base policy, taking proprioception and the estimated object state X^box to produce residual actions Δat . The residual action aims to adaptively lift a 3D object. (3) A recurrent online object state estimator fuses vision and proprioceptive cues using a V-LSTM and MLP to infer X^box , and is trained jointly with the residual manipulation policy. During real-world deployment, the robot uses onboard estimators and LiDAR odometry and executes the combined policies πwbc and πres to complete the whole-body loco-manipulation task.
Parameter
Range
Operation
Base Mass [kg]
[-2.5, 2.5]
Add
kp
[ 0.8, 1.2]
Scale
kd
[ 0.8, 1.2]
Scale
Ground Static Friction
[ 0.3, 1.5]
Absolute
Ground Dynamic Friction
[ 0.3, 0.9]
Absolute
Base Force Disturbance [N]
[-4.0, 4.0]
Absolute
TABLE II : Domain randomization parameters for policy training.
Method
Whole
Stage1
Stage2
Drops ( ↓ )
Regrasps ( ↑ )
IsaacLab
Pure RL
0.62 (±0.48)
0.98 (±0.14)
0.88 (±0.32)
1.79 (±2.39)
5.74 (±3.91)
Pure RL + FK
0.88 (±0.32)
0.93 (±0.25)
0.92 (±0.27)
0.29 (±0.89)
2.94 (±1.88)
Motion Imitation
0.42 (±0.46)
0.96 (±0.19)
0.93 (±0.18)
3.81 (±4.77)
2.94 (±1.99)
AdaptManip (Ours)
0.85 (±0.35)
0.97 (±0.17)
0.92 (±0.26)
0.49 (±1.14)
2.12 (±1.95)
Pure Vision
0.63 (±0.42)
0.95 (±0.22)
0.91 (±0.20)
12.44 (±9.36)
23.55 (±15.32)
TABLE III : Performance comparison between AdaptManip and baselines across 135 trials in MuJoCo and IsaacLab.
Fig. 4 : State estimation error of our method. Shows mean ± 1 standard deviation across 50 episodes. The green region shows the area where vision is available, and the purple region shows the area where there is contact between the robot and box.
Fig. 5 : Hardware demonstration of the three-stage whole-body loco-manipulation task.
Fig. 6 : Hardware demonstration of whole-body loco-manipulation. The yellow region indicates the grasp formation phase where the robot carefully coordinates its arms for a secure hold, while the red region highlights the robot’s ability to recover from a transient loss of stability through corrective arm motions.