Humanoids are increasingly capable of tracking complex whole-body motions, but physical interaction introduces a different challenge. When a robot makes contact with a person or the environment, it needs to respond to external forces while preserving the motion needed for the task. This response can vary across directions in the end-effectors and on the body. For example, an end effector may need to accommodate contact force in one direction while maintaining motion accuracy in another, while the robot body may resist an external force or move with it. We present a compliance framework for humanoid loco-manipulation that combines directional and tunable end-effector (EE) compliance with selectable root compliance for external force rejection or force following. A hierarchical reinforcement learning controller modulates a fixed whole-body tracking policy through high-level EE and root commands, while interaction forces are estimated from proprioceptive history. Our simulation and real-world experiments on a humanoid demonstrate directional stiffness control, online stiffness adjustment, distinct root compliance, compliant manipulation, and collaborative carrying.
Figures & tables
Method
Direction
Angular
Tunable
Gen. Track.
Scope
Loco.
Gentle [ 5 ]
×
×
100–300 †
✓
Upper body
×
CHIP [ 7 ]
×
×
1/k:0 – 0.05‡
✓
EE
×
SoftMimic [ 4 ]
×
×
40–1000
×
EE
×
CEER [ 6 ]
×
×
×
✓
EE
×
LAC [ 11 ]
×
✓
10–500
✓
Arms + torso
×
Ours
✓
×
100–600
✓
EE + root
✓
TABLE I: Comparison of humanoid compliance methods.
Fig. 2: System overview. Stage 1 trains a stiff low-level whole-body tracking policy with an EE–root command interface. In Stage 2, the low-level policy is held fixed while high-level residual policies learn EE or root command corrections that produce the desired response to external forces under commanded stiffness or a selected root mode. At deployment, an analytical compliance-space gate composes 10 EE compliance experts according to the commanded Cartesian stiffness, while a root compliance policy is selected through mode switching. The low-level policy remains active throughout, enabling online adjustment of EE stiffness and root mode switching.
Policy
Nom. Err. (m) ↓
Median (N/m)
P95 (N/m)
3-kp stiff (30 N)
0.0456±0.0140
621.4
1793.0
CEER [ 6 ]
0.1385±0.1657
103.3
285.0
SONIC [ 1 ]
0.0955±0.0482
192.3
342.4
3-kp stiff (40 N)
0.0524±0.0148
544.7
961.8
3-kp stiff (50 N)
0.0528±0.0134
672.1
3426.0
3-kp stiff (60 N)
0.0830±0.0269
406.5
1088.8
TABLE II: Low-level tracking-policy and training-force characterization. Apparent stiffness values describe natural behavior and are not errors relative to a target.
Fig. 3: Natural apparent stiffness of representative low-level policies. Positive and negative directions are combined for each axis. Thin lines show P5–P95, thick lines show the interquartile range, and dots show the median. Dotted lines indicate 200 and 600N/m .
Policy
EE Comp. Err. (m) ↓
Nom. EE Err. (m) ↓
App. Stiffness (N/m)
Analytical MoE
0.023±0.016
0.049±0.045
219.6±49.2
Fixed-stiffness HL
0.028±0.014
0.050±0.040
241.5±55.7
Fixed-stiffness E2E
0.082±0.034
0.100±0.046
316.6±512.0
TABLE III: Fixed-stiffness compliance evaluation at 200N/m .
Fig. 4: Commanded versus apparent stiffness over the 16-configuration 100–600 N/m evaluation, shown separately for the x , y , and z axes. Curves show the median response, shaded regions show the interquartile range, and the dashed line denotes ideal tracking. For readability, the unusually broad stiffness-conditioned E2E IQR band is omitted only in the x panel.
Accuracy
Trend
Method
Compliance Matrix Err. ↓
EE Comp. Err. (m) ↓
Stiffness MAPE ↓
Slope →1
R2↑
Analytical MoE
0.278
0.0349
0.249
0.764
0.779
Stiff.-cond. HL
0.677
0.0566
1.325
0.136
0.040
Stiff.-cond. E2E
1.064
0.0884
5.208
3.430
0.003
Oracle-force LL
0.304
0.0459
0.256
0.485
0.792
TABLE IV: Full-range directional EE compliance over 16 stiffness configurations (ideal trend slope and R2 : 1 ).
Mode
Root drift (m) ↓
Velocity error (m/s) ↓
Apparent B median / target
3-kp LL only
0.3627
0.0955
–
Resistance
0.3103
0.0987
–
B=60
1.7806
0.1143
57.0 / 60
B=200
0.7163
0.1031
201.4 / 200
TABLE V: Root-compliance evaluation using the 3-kp root-residual action space without yaw control.
Mode
3-kp
3-kp+yaw
5-kp
5-kp+yaw
Resistance
0.3103
0.2032
0.4964
0.1574
B=60
0.1143
0.1217
0.1094
0.1073
B=200
0.1031
0.0678
0.0904
0.0700
TABLE VI: Root-policy action-space ablation. Resistance is evaluated by planar root drift (m), and damping by velocity-response error (m/s). All variants include root-command residuals; 5-kp adds foot-position residuals, and +yaw adds a root-yaw residual.
Fig. 5: Directional EE compliance tasks: writing “8” (a–c), straight-line writing under disturbance (d–f), and writing on a tilted surface (g–l). Tunable EE-stiffness demonstrations: force-gauge pulling (m–o) and yoga-ball grasping (p–q).
Humanoid robots have achieved impressive locomotion performance, yet contact-rich and long-horizon manipulation remains a major bottleneck. Manipulation is inherently contact-rich and demands compliant whole-body control for stable interaction, while its diversity and long-horizon nature favor modular, planner-compatible interfaces over joint-space tracking. We propose CEER, a compliant end-effector-root (EE-root) control abstraction for modular humanoid loco-manipulation within a hierarchical planning framework. CEER enables compliance-aware whole-body control in an interpretable task space defined by root motion commands and end-effector pose targets, and supports plug-and-play integration with heterogeneous high-level planners. A teacher-student framework is adopted to distill a general motion-tracking controller into a low-level policy that consumes only EE-root commands. We further construct a hierarchical system that integrates heterogeneous planners and task modules through the EE-root interface, enabling diverse manipulation tasks without retraining the underlying whole-body policy. Experiments in simulation and on hardware demonstrate 3.3 cm end-effector tracking accuracy with substantially reduced jerk compared to baselines, stable contact-rich manipulation under teleoperation, and up to 70% success in simulated single-object loco-manipulation tasks within a room-scale environment. These results indicate that compliant EE-root control provides a practical abstraction for humanoid loco-manipulation, enabling modular and scalable integration of diverse skills.
Xinyuan Luo, Xingrui Chen, Xunjian Yin +6
Department of Mechanical Engineering and Materials Science, Duke University, Durham, NC 27708, USA
Whole-body compliant control is essential for deploying heavy humanoids under high payload in human-centric environments. Most prior force-aware learning-based pipelines focus on end-effector resistance, per-link upper-body springs, or end-effector stiffness modulation, leaving arbitrary-site perturbations on heavy platforms with lower-body engagement largely unaddressed. We close this gap with CompliantWBC comprising: (1) A base policy trained with RL to maximize compliance-fidelity reward, guided by a multi-site whole-body impedance reference controller, extending classical Cartesian impedance to any controlled link; (2) A bounded residual policy that edits the per-link impedance equilibrium over a frozen base, correcting the coarse but structured wrench estimate supplied by a force encoder co-trained behind a gradient barrier; (3) A Phong-weighted force-origin sampler with an axis-decoupled pelvis anchor induces lower-body-inclusive compliance curriculum training via two interpretable parameters. We evaluate CompliantWBC in simulation against both compliant and stiff baselines, achieving best compliant fidelity of 2.58cm deviation from analytical solutions, and demonstrate it on a real heavy humanoid across static/dynamic force reaction, board wiping, squat under payload, and cooperative payload transport. Project website: https://dotandung.github.io/compliantwbc/
Tan-Dzung Do, Cuc T. Trinh, Tuan Dat Phuong +4
VinRobotics · National University of Singapore · VinUniversity +1
Bipedal loco-manipulation enables robots to interact with objects beyond the nominal workspace of their arms by coordinating locomotion and manipulation. Realizing this capability requires a low-level whole-body controller that translates task-level manipulation goals into coordinated arm and leg motions while maintaining balance. We present a unified whole-body controller trained with reinforcement learning that directly maps 6-DoF end-effector targets to coordinated actions for the bipedal base and robotic arm. Given only an end-effector target, the learned controller autonomously coordinates reaching, postural adaptation, and stepping without explicit base-velocity or footstep commands. A reward-gating strategy regulates the trade-offs among end-effector tracking, locomotion, and balance during training, while a temporal context estimator combines windowed Transformer encoding, recurrent GRU memory, and auxiliary dynamics prediction to extract dynamics-relevant information from observation history. Real-robot experiments demonstrate that the same controller supports reaching, postural adaptation, and stepping under commands from VR teleoperation, a learned diffusion policy, and scripted trajectories, providing a common end-effector interface for diverse manipulation tasks.