FAME: Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid
Organizations: Department of Computer Science, University of Colorado Boulder, Boulder, CO 80309.
Abstract
Maintaining balance under external hand forces is critical for humanoid bimanual manipulation, where interaction forces propagate through the kinematic chain and constrain the feasible manipulation envelope. We propose FAME, a force-adaptive reinforcement learning framework that conditions a standing policy on a learned latent context encoding upper-body joint configuration and bimanual interaction forces jointly, since the base moment a load induces depends on the arm configuration through which it acts. Training applies isotropically sampled 3D forces at each hand under an upper-body pose curriculum, exposing the policy to manipulation-induced perturbations across continuously varying arm configurations. At deployment the interaction force is not measured but reconstructed online from joint torques and states through rigid-body inverse dynamics, requiring no wrist force/torque sensing. We evaluate over upper-body configurations under swept hand forces, scoring each trial by a task-level criterion that requires the robot both to remain upright and to hold its hands near where the task placed them; all such results run with the estimated force in the loop. At a ,mm tolerance FAME reaches task success, against for a policy given the same force without encoding, for a pose-conditioned curriculum policy, and for an adversarially trained locomotion policy, which stays upright but recovers by stepping and so relocates the hands. We further demonstrate transfer to task-generated interaction forces in a MuJoCo kitchen environment, and to asymmetric and bimanual loading on a full-scale Unitree H1-2. Code and videos are available on the https://correlllab.github.io/fame_website.
Figures & tables
| Term | Expression | Weight | Remarks |
| Base Stability & Uprightness | |||
| Base height tracking | |||
| Vertical velocity | Penalize bouncing | ||
| Angular velocity (xy) | Suppress roll/pitch motion | ||
| Orientation | Projected gravity error | ||
| Stand still (zero command) | Encourage static stability | ||
| Information available | Task success (%) | Deviation | |||||
| Variant | Pose curr. | Force to policy | Arm cfg. in encoder | Latent encoding | mm | mm | EE (mm) |
| FAME-pose (HOMIE-style [ 11 ] ) | ✓ | — | ✓ | ✓ | 4.3 | 5.6 | 242.4 |
| +Curr+F raw | ✓ | ✓ | — | — | 16.6 | 21.7 | 220.3 |
| ALMI [ 15 ] | adversarial motion imitation; no force input | 24.7 | 39.3 | 202.3 | |||
| FAME | ✓ | ✓ | ✓ | ✓ | 38.9 | 48.5 | 179.6 |
| Actor observation term (one-step) | Symbol | Dim |
| Command (scaled) | 3 | |
| Commanded base height | 1 | |
| IMU angular velocity (body frame) | 3 | |
| Projected gravity (body frame) | 3 | |
| Joint position error (scaled) | ||
| Joint velocity (scaled) |
| Critic observation term | Symbol | Dim |
| All actor one-step terms | 76 | |
| Latent context (from encoder) | 8 | |
| Base linear velocity (privileged, scaled) | 3 | |
| One-step critic obs: | ||
| History stacking: | ||
| Term | Value |
| External Disturbances | |
| Push robot | , |
| Upper-body disturbance curriculum | initially; increased with standing quality, interval |
| Actuation & Observation Noise | |
| Joint injection noise | , |
| Actuation offset | , |
| Hyperparameter | Value | Hyperparameter | Value |
| Optimizer | Adam | Discount factor | 0.99 |
| Learning rate | GAE | 0.95 | |
| Schedule | adaptive | Clip param | 0.2 |
| Learning epochs | 5 | Value loss coef | 1.0 |
| Mini-batches | 4 | Clipped value loss | True |
| Max grad norm | 1.0 | Entropy coef | 0.01 |