Partnered human-humanoid interaction couples locomotion with continuous physical contact. A humanoid needs to coordinate with a person's motion while responding to interaction forces and maintaining stable and natural movement. We present CoDance, a framework for learning reactive and compliant human-humanoid interaction from video. We study partnered dancing as a challenging instantiation, where a humanoid coordinates its footsteps with a moving partner and maintains continuous two-hand contact. Given a single video of two human dancers, CoDance retargets their motions into a robot reference and a moving partner. We introduce a multi-link compliance augmentation that transforms the kinematic demonstration into force-aware training data by adapting the robot reference under structured forces at both hands. Policies trained on this data follow the observed partner while preserving the demonstrated locomotion style and responding compliantly to physical interaction. In simulation, the policies adapt their footsteps to changes in the partner and reproduce approximately 80% of the wrist displacement encoded by the augmented demonstrations. On a physical humanoid, CoDance enables sustained two-hand dancing with a human partner including repeated transitions between forward and backward motions.
Figures & tables
Fig. 2 : Overview of CoDance . A paired-dancing video is converted into force-aware training data through motion retargeting and multi-link compliance augmentation. The policy learns to follow the observed partner, maintain the demonstrated locomotion style, and remain compliant at both hands. At deployment, the simulated partner is replaced by a person. Training is illustrated with the decoupled policy.
Parameter
Value
Robot stiffness K
log-uniform, 101000Nm−1
Environment stiffness
same range, independent draw
Force ceiling
140N
Displacement scale dmax
0.15m
per-event draw
U(0.03,dmax) m
Balance budget (CoM shift)
0.10m
TABLE I : Augmentation parameters.
Fig. 3 : Evaluation results on compliance interaction. Under force, both compliant policies move the wrist most of the distance encoded in the adapted reference, while the stiff tracker moves it far less. Error bars show the standard deviation over evaluation seeds.
Fig. 4 : Visualization of the free and adapted references (top) and of the decoupled policy with and without the adversarial term, at the middle of four force holds. The partner is in teal, and arrows mark the force at the left wrist (blue) and right wrist (pink), with the applied force solid and the scheduled force faded. Without the adversarial term, the pelvis turns away from the dancer’s heading, the waist and hips twist back toward the direction of travel, and the stance widens into a crouch. Top row also shows that separate retargeting does not preserve relative position in the paired motion clip.
Fig. 5 : Visualization of one real-robot dance with a human partner. The robot keeps close to the following distance for most of the dance. Top: the trajectories of the robot and the partner. Bottom: the partner’s root in the robot’s pelvis frame.
Fig. 6 : Visualization of a standing policy under force on one or both hands. The red overlay is the reference standing pose. Left: one hand loaded by a hanging weight. Middle: the two hands pulled in different directions. Right: both hands pulled outward. The arms move with the pull and the feet stay still.
TABLE V : Reward terms of the two trained policies. Tracking terms are exponential kernels exp(−e2/σ2) of the error e with the listed σ ; weights multiply the per-step term. The two policies share every term except the locomotion objective: the decoupled policy draws its locomotion style from the adversarial term, the whole-body tracker (in the style of SoftMimic) from reference tracking of the lower body of the adapted reference.
Term
Per step
History
Actor noise
Actor
Critic
Base linear velocity (IMU)
3
5
✓
Base angular velocity (IMU)
3
5
±0.2
✓
✓
Joint positions
29
5
±0.01
✓
✓
Joint velocities
29
5
±0.5
✓
✓
Previous actions
29
5
✓
✓
Projected gravity
3
5
±0.05
✓
✓
TABLE VI : Observations of the two trained policies, identical for both. Per-step sizes are stacked over the listed history. Noise is uniform in the listed range and applied to the actor only; the actor’s joint positions also carry the per-joint encoder bias of Table VII . Both networks are MLPs with hidden layers 512, 256, 128 and ELU activations. The AMP discriminator of the decoupled policy reads 2 frames of 63 features (positions and 6D orientations of 7 lower-body and torso links in the pelvis frame), which the policy itself does not see.
Setting
Value
Randomization (at episode start)
Torso centre of mass offset
uniform, x±0.025m , y±0.05m , z±0.05m
Foot friction coefficient
uniform, 0.31.2 , one draw for all foot geoms
Joint encoder bias
uniform, ±0.01rad per joint, added to the actor’s joint positions
Actor observation noise
uniform, per observation term, at the ranges the policy trained with
Reset root pose
x , y±0.15m , yaw ±0.3rad about the clip pose
TABLE VII : Randomization, force replay and episode settings, identical for both policies.