Humanoids are increasingly capable of tracking complex whole-body motions, but physical interaction introduces a different challenge. When a robot makes contact with a person or the environment, it needs to respond to external forces while preserving the motion needed for the task. This response can vary across directions in the end-effectors and on the body. For example, an end effector may need to accommodate contact force in one direction while maintaining motion accuracy in another, while the robot body may resist an external force or move with it. We present a compliance framework for humanoid loco-manipulation that combines directional and tunable end-effector (EE) compliance with selectable root compliance for external force rejection or force following. A hierarchical reinforcement learning controller modulates a fixed whole-body tracking policy through high-level EE and root commands, while interaction forces are estimated from proprioceptive history. Our simulation and real-world experiments on a humanoid demonstrate directional stiffness control, online stiffness adjustment, distinct root compliance, compliant manipulation, and collaborative carrying.
Figures & tables
Method
Direction
Angular
Tunable
Gen. Track.
Scope
Loco.
Gentle [ 5 ]
×
×
100–300 †
✓
Upper body
×
CHIP [ 7 ]
×
×
1/k:0 – 0.05‡
✓
EE
×
SoftMimic [ 4 ]
×
×
40–1000
×
EE
×
CEER [ 6 ]
×
×
×
✓
EE
×
LAC [ 11 ]
×
✓
10–500
✓
Arms + torso
×
Ours
✓
×
100–600
✓
EE + root
✓
TABLE I: Comparison of humanoid compliance methods.
Fig. 2: System overview. Stage 1 trains a stiff low-level whole-body tracking policy with an EE–root command interface. In Stage 2, the low-level policy is held fixed while high-level residual policies learn EE or root command corrections that produce the desired response to external forces under commanded stiffness or a selected root mode. At deployment, an analytical compliance-space gate composes 10 EE compliance experts according to the commanded Cartesian stiffness, while a root compliance policy is selected through mode switching. The low-level policy remains active throughout, enabling online adjustment of EE stiffness and root mode switching.
Policy
Nom. Err. (m) ↓
Median (N/m)
P95 (N/m)
3-kp stiff (30 N)
0.0456±0.0140
621.4
1793.0
CEER [ 6 ]
0.1385±0.1657
103.3
285.0
SONIC [ 1 ]
0.0955±0.0482
192.3
342.4
3-kp stiff (40 N)
0.0524±0.0148
544.7
961.8
3-kp stiff (50 N)
0.0528±0.0134
672.1
3426.0
3-kp stiff (60 N)
0.0830±0.0269
406.5
1088.8
TABLE II: Low-level tracking-policy and training-force characterization. Apparent stiffness values describe natural behavior and are not errors relative to a target.
Fig. 3: Natural apparent stiffness of representative low-level policies. Positive and negative directions are combined for each axis. Thin lines show P5–P95, thick lines show the interquartile range, and dots show the median. Dotted lines indicate 200 and 600N/m .
Policy
EE Comp. Err. (m) ↓
Nom. EE Err. (m) ↓
App. Stiffness (N/m)
Analytical MoE
0.023±0.016
0.049±0.045
219.6±49.2
Fixed-stiffness HL
0.028±0.014
0.050±0.040
241.5±55.7
Fixed-stiffness E2E
0.082±0.034
0.100±0.046
316.6±512.0
TABLE III: Fixed-stiffness compliance evaluation at 200N/m .
Fig. 4: Commanded versus apparent stiffness over the 16-configuration 100–600 N/m evaluation, shown separately for the x , y , and z axes. Curves show the median response, shaded regions show the interquartile range, and the dashed line denotes ideal tracking. For readability, the unusually broad stiffness-conditioned E2E IQR band is omitted only in the x panel.
Accuracy
Trend
Method
Compliance Matrix Err. ↓
EE Comp. Err. (m) ↓
Stiffness MAPE ↓
Slope →1
R2↑
Analytical MoE
0.278
0.0349
0.249
0.764
0.779
Stiff.-cond. HL
0.677
0.0566
1.325
0.136
0.040
Stiff.-cond. E2E
1.064
0.0884
5.208
3.430
0.003
Oracle-force LL
0.304
0.0459
0.256
0.485
0.792
TABLE IV: Full-range directional EE compliance over 16 stiffness configurations (ideal trend slope and R2 : 1 ).
Mode
Root drift (m) ↓
Velocity error (m/s) ↓
Apparent B median / target
3-kp LL only
0.3627
0.0955
–
Resistance
0.3103
0.0987
–
B=60
1.7806
0.1143
57.0 / 60
B=200
0.7163
0.1031
201.4 / 200
TABLE V: Root-compliance evaluation using the 3-kp root-residual action space without yaw control.
Mode
3-kp
3-kp+yaw
5-kp
5-kp+yaw
Resistance
0.3103
0.2032
0.4964
0.1574
B=60
0.1143
0.1217
0.1094
0.1073
B=200
0.1031
0.0678
0.0904
0.0700
TABLE VI: Root-policy action-space ablation. Resistance is evaluated by planar root drift (m), and damping by velocity-response error (m/s). All variants include root-command residuals; 5-kp adds foot-position residuals, and +yaw adds a root-yaw residual.
Fig. 5: Directional EE compliance tasks: writing “8” (a–c), straight-line writing under disturbance (d–f), and writing on a tilted surface (g–l). Tunable EE-stiffness demonstrations: force-gauge pulling (m–o) and yoga-ball grasping (p–q).