Most humanoid loco-manipulation controllers require human motion data to learn whole-body coordination and posture, leaving policies reliant on external sources to provide this data. We present OCLO (Online-posture Compliant LOco-manipulation), a humanoid loco-manipulation system trained without human motion data and commanded only through two end-effector targets. Because these targets do not uniquely determine whole-body posture, OCLO generates pelvis height and torso orientation online using an analytic reachability prior, further refined through policy-in-the-loop sampling with a task-agnostic cost. OCLO also learns whole-body compliance by displacing end-effector references according to measured forces through a spring-damper model, encouraging the legs, waist, and pelvis to yield to external loads. In simulation, using the reachability prior leads to a 77.8% success rate in acquiring the commanded reference, a vast improvement over the 37.8% success rate accomplished without the prior. Further, refinement reduces end-effector orientation error across all evaluated tasks. The same posture module improves a pretrained SONIC controller on four of five tasks. Without compliance training, policies tend to lose balance under disturbances rather than sacrifice tracking. On a Unitree G1, OCLO maintains balance under end-effector disturbances that cause its ablations to fail and performs seven loco-manipulation tasks, including crouched walking and picking up a box from a low surface. Project website: https://oclo-humanoid.github.io/
Figures & tables
System
No human data
Compliance
Posture online
Humanoid hardware
OmniH2O [ 22 ]
–
–
–
✓
Mobile-TeleVision [ 2 ]
∘
–
–
✓
SONIC [ 1 ]
–
–
–
✓
HOMIE [ 3 ]
✓
–
–
✓
AMO [ 19 ]
∘
–
–
✓
HugWBC [ 4 ]
✓
–
–
✓
TABLE I: Positioning against prior systems. ✓ yes, ∘ partial, – no. No human data: no motion capture, retargeted motion, or demonstrations; Compliance: force-yielding or impedance tracking; Posture online: robot-selected pelvis height or torso orientation; Humanoid hardware: physical humanoid loco-manipulation.
Fig. 2: Training and deployment of OCLO. (a) The lower-body policy is trained through a three-stage curriculum: locomotion, EE reaching with interaction loads, and finally posture adaptation through crouching and leaning. All training commands are procedurally generated, and the policy is frozen after training. (b) At deployment, two EE targets define the manipulation intent. The spring-damper model produces compliant EE references for differential IK, while the online posture module generates pelvis height and torso orientation. For real-time hardware operation, the analytic prior is applied directly (OCLO (Prior)). When higher pose accuracy is desired, it can be refined by a Cross-Entropy Method (CEM) optimizer using frozen-policy rollouts from a snapshot of the live robot state (OCLO (Prior + CEM)). The resulting posture command is executed by the lower-body policy together with the IK-controlled arms.
Term
Error ei
σi
wi
Stage
End-effector tracking
EE position (per arm)
∥pee−pref∥
0.2 m
+0.5
2–3
EE orientation (per arm)
quat. geodesic
0.5
+0.5
2–3
Locomotion
Velocity command (lin. / yaw)
∥vbxy−vcmdxy∥ / ∣ωbz−ωcmdz∣
0.35 / 0.25
+4.0 / +2.5
1–3
Posture & base
TABLE II: Principal reward terms. Tracking terms use ri=exp(−ei2/σi2) , and Stage 3 swaps the fixed-stance penalties (highlighted rows replace them) for posture tracking. Standard regularization terms (action rate, joint acceleration, velocity, torque, limits, contacts) are omitted.
Track
Base Push
EE Push
Sustained Load
Crouched Push †
Method
Succ. ↑
Ori. ↓
Succ. ↑
Ori. ↓
Succ. ↑
Ori. ↓
Succ. ↑
Ori. ↓
Succ. ↑
Ori. ↓
B1, no targets, no forces
56.7
14.7
13.3
43.0
32.2
17.4
41.1
16.6
4.4
68.4
B2, targets, no forces
56.7
23.0
23.3
45.9
40.0
24.5
70.0
25.4
16.7
72.4
OCLO (CEM)
37.8
10.3
27.8
17.1
25.6
11.6
42.2
10.7
40.0
52.1
OCLO (Prior)
82.2
17.6
61.1
22.6
62.2
17.9
64.4
16.6
45.6
51.3
OCLO (Prior + CEM)
77.8
11.5
62.2
17.1
63.3
12.5
72.2
11.7
48.9
48.2
TABLE III: Simulation evaluation. Success (%) and EE orientation error ( ∘ ) over 90 paired episodes per cell, with identical episode draws across methods. The upper block compares training ablations B1 and B2 with OCLO under three posture-generation variants; bold indicates the best result within this block. The lower block evaluates pretrained SONIC [ 1 ] using the same posture interface. † For Crouched Push, success is determined by falls only.
Fig. 3: Success rate under increasing disturbance force. (a) Base Push at 40, 60, and 80 N and (b) Crouched Push at 30, 60, and 90 N, 20 episodes per point. † Crouched Push is scored on falls alone.
Fig. 4: Evaluation under external force. The robot maintains balance and recovers its posture under (a) forward pulling, (b) pushing, and (c) lateral pulling from the left. Each column shows representative frames over time for the corresponding disturbance.
Method
Torso pull
Torso push
EE pull
EE push
Lateral pull
40–60 N
20–40 N
50–70 N
40–60 N
B1
8/10
7/10
1/10
—
0/10
0/10
B2
8/10
8/10
5/10
—
3/10
2/10
OCLO (Prior)
10/10
9/10
10/10
10/10
9/10
8/10
OCLO (Prior + CEM)
9/10
9/10
9/10
—
8/10
8/10
TABLE IV: Real-world disturbance rejection. Successful trials out of ten per cell. Pulls are applied through an inline force gauge at the stated range. Pushes are applied manually, the EE push through a PVC pipe, and are not instrumented. B1, B2, and OCLO (Prior) hold the same commanded posture throughout, and OCLO (Prior + CEM) refines it online. A dash marks an untested condition.
Humanoid loco-manipulation requires adaptive whole-body coordination to seamlessly integrate locomotion and physical interaction. Despite recent advances, learning autonomous loco-manipulation remains challenging due to the scarcity of diverse, physically executable robot-object interaction data and the difficulty of learning unified whole-body control directly from onboard observations. We present ViLoMan, a scalable framework for autonomous humanoid loco-manipulation. ViLoMan first transforms partial kinematic demonstrations of human-object interactions into complete, physically executable robot trajectories. It then leverages these trajectories within a teacher-student distillation framework to learn a unified policy that maps egocentric depth observations and proprioceptive measurements directly to joint-level whole-body actions. During deployment, the policy requires neither reference motions nor intermediate commands. We evaluate ViLoMan on door-closing tasks across diverse door configurations and robot initial conditions in both simulation and the real world. Experimental results demonstrate that a single policy enables a Unitree G1 humanoid to complete the full task using only onboard depth sensing and proprioception, while generalizing robustly across task variations and transferring effectively from simulation to reality. Project page: viloman-anonymous.pages.dev.
Zejie Tian, Ruibing Hou, Bingpeng Ma +2
State Key Laboratory of AI Safety, Institute of Computing Technology, CAS, China.
Whole-body humanoid loco-manipulation requires coordinating the robot's entire kinematic chain. However, most existing systems typically decouple the upper and lower bodies into separate controllers, limiting such coordination and yielding behaviors similar to those of a wheeled dual-arm platform. In this paper, we ask what it takes to build a whole-body native vision-language-action (VLA) model that maps language and pixels directly to all of the humanoid's degrees of freedom. We conduct a systematic empirical study organized as a roadmap of one-variable-at-a-time experiments across three phases: whole-body teleoperation, VLA model design, and heterogeneous co-training. Our study yields several intriguing findings: a joint-based whole-body teleoperation interface outperforms alternatives that only partially expose the humanoid's degrees of freedom; a VLA pretrained on static and wheeled dual-arm platforms transfers surprisingly well to a humanoid's full action space; and co-training with HuMI, the humanoid analog of UMI, extends the policy to new objects and instructions without additional whole-body teleoperation on those targets. Following this roadmap yields OpenHLM, an open-source recipe for whole-body humanoid loco-manipulation. In a challenging long-horizon task that spans a wide vertical range of the humanoid, OpenHLM outperforms two state-of-the-art humanoid VLA baselines (GR00T N1.6 and Ψ0) using less than half the total demonstration time. Our code, training data, and model checkpoints are available at [https://openhlm-project.github.io/].
Yingdong Hu, Haodong Zhu, Boyuan Zheng +6
Tsinghua University · Shanghai Qi Zhi Institute · Spirit AI
Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.