Track-and-Complete: Learning Humanoid Skills from a Single Failed Human Video
Organizations: School of Mechanical Engineering, Yonsei University, Seoul 03722, Korea
Abstract
Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until the failure occurs, after which the task-completion reward guides the policy to learn the intended task goal without requiring a successful task trajectory. We evaluate our method on six in-the-wild failed human tasks from the Oops! dataset. Our experimental results demonstrate the effectiveness of the proposed approach for learning from failed attempts when no successful demonstration is available. Thus, these findings establish failed human videos as a viable source of supervision for humanoid skill learning.
Figures & tables
| Method | Embodiment | Demonstration | Single video | Use of failure |
| VideoMimic [ 1 ] | Humanoid | Target video | Multi | – |
| MeshMimic [ 2 ] | Humanoid | Target video | – | |
| HDMI [ 3 ] | Humanoid | Target video | – | |
| OKAMI [ 4 ] | Humanoid | Target video | – | |
| ORION [ 5 ] | Arm | Target video | – | |
| LUCID [ 6 ] | Arm | Video corpus | Multi | – |
| Group | Component | Dim. | After |
| Proprioception | Root position, orientation, velocity | 13 | Observed |
| Joint position | 29 | Observed | |
| Joint velocity | 29 | Observed | |
| Previous action | 29 | Observed | |
| Task state | Task object, goal, and scene state | Observed | |
| Reference | Activity, phase, release time | 3 | Masked-out |
| Task | Selected reward terms and weights |
|---|---|
| Kick target | Pad approach ( ); latched strike ( ); after the strike: two-foot contact ( ), pelvis height ( ), and stability and uprightness ( ); not fallen ( ). |
| Football | Foot–ball approach ( ); ball contact ( ); ball speed toward the goal ( ); ball–goal distance ( ); uprightness ( ); not fallen ( ); hovering without contact penalty ( ). |
| Backflip | Landing-region approach ( ); left and right foot contact with load ( each); landing pelvis height ( ); low linear ( ) and angular ( ) velocity; uprightness ( ); latched inversion bonus ( ). |
| Box jump | Feet near the box top ( ); loaded top contact ( ); standing height ( ); pelvis over the box ( ); controlled standing ( ); uprightness ( ); settling ( ); non-top ( ) and invalid-contact ( ) penalties; reward set to on success. |
| Handstand | Hand-patch approach ( ); left and right palm contact ( each); bilateral palm contact ( ); after bilateral contact: inversion progress ( ) and feet above the pelvis ( ); pelvis height ( ); settling ( ). |
| Log walk | Progress toward the far end of the log ( ); foot contact on the log top ( ); pelvis height above the log ( ); uprightness ( ); hovering without contact penalty ( ). |
| Task | Success criterion |
|---|---|
| Kick target | Foot strikes the target at m/s, upright , settled, held s. |
| Football | Ball enters the goal region, upright , settled, held s. |
| Backflip | Root becomes fully inverted, then both feet in the landing region and in contact, upright , settled, held s. |
| Box jump | Both feet on the box top, upright , settled, held s. |
| Handstand | Both hands in the support region and in contact, feet at least m above the pelvis, held s. |
| Log walk | Net displacement along the log m without falling. |
| Policy | Kick target | Football | Backflip | Box jump | Handstand | Log walk | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Survival (%) | Success (%) | Survival (%) | Success (%) | Survival (%) | Success (%) | Survival (%) | Success (%) | Survival (%) | Success (%) | Survival (%) | Success (%) | |
| Task reward only | 100.0 | 0.0 | 100.0 | 0.0 | 100.0 | 0.0 | 94.6 | 0.0 | 100.0 | 0.0 | 100.0 | 0.0 |
| Unified policy (ours) | 100.0 | 99.3 | 100.0 | 100.0 | 99.7 | 41.7 | 98.6 | 75.8 | 100.0 | 56.3 | 100.0 | 100.0 |