MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking
Organizations: The Hong Kong University of Science and Technology (Guangzhou)
Abstract
Human videos provide rich motion targets for humanoid learning, yet visually plausible references can still produce persistent failures under physics-based execution. These failures reveal where training supervision should change. We present MimicX, a policy-in-the-loop framework that uses execution feedback to refine video-driven humanoid motion tracking. Starting from reconstructed and retargeted motion, MimicX localizes difficult transitions and affected body regions, then jointly adapts tracking objectives and the reset curriculum for policy continuation. Repeated rollout verification selects execution-priority improvements subject to tracking guards. Across four core video tasks, MimicX consistently improves tracking accuracy and Robust Execution Horizon relative to the Fixed Reference baseline. Task-averaged results show a 25.7% reduction in body-tracking error and a 255.6% increase in execution horizon. Additional video, motion-reference, and collision-scene studies evaluate the method beyond the core tasks, while MimicX-HLoop accelerates feedback through heterogeneous execution. Overall, MimicX turns policy failure into actionable supervision for deciding what to refine and which refinement to retain.
Figures & tables
| Task | Body mean (m) | Body p95 (m) | Horizon (steps) | Success (%) | ||||||
| Fixed | MimicX | Red. (%) | Fixed | MimicX | Fixed | MimicX | Gain (%) | Fixed | MimicX | |
| Tennis | 0.241 | 0.157 | 35.0 | 0.518 | 0.278 | 322.0 | 801.0 | +148.8 | 11.1 | 100.0 |
| Football | 0.286 | 0.244 | 14.7 | 0.612 | 0.504 | 53.7 | 427.3 | +696.3 | 0.0 | 0.0 |
| Dance | 0.270 | 0.244 | 9.9 | 0.502 | 0.467 | 193.3 | 341.3 | +76.6 | 0.0 | 0.0 |
| Kung Fu | 0.249 | 0.141 | 43.3 | 0.634 | 0.296 | 97.7 | 196.0 | +100.7 | 0.0 | 0.0 |
| Method | Policy setup | Joint RMSE (rad, ) | Reduction (%) | Local body error (m, ) | Reduction (%) |
| Fixed Reference | Fixed objective | 0.389 0.166 | +0.0 | 0.151 0.072 | +0.0 |
| BeyondMimic (MjLab) | Fixed objective | 0.389 0.027 | +0.1 | 0.137 0.021 | +9.8 |
| SONIC (released) | Pretrained | 0.857 0.018 | -120.2 | 0.236 0.001 | -55.8 |
| MimicX | Closed loop | 0.201 0.002 | +48.3 | 0.055 0.002 | +63.4 |
| Method | Supervision | Execution | Tracking | Dynamics | |||||
| C | T | V | Success (%, ) | Horizon (%, ) | Body (m, ) | Anchor (m, ) | Linear vel. (m/s, ) | Angular vel. (rad/s, ) | |
| Fixed Reference | – | – | – | 2.8 | 21.5 | 0.262 | 0.219 | 0.691 | 2.955 |
| Failure Curriculum | – | – | 25.0 | 62.7 | 0.215 | 0.135 | 0.451 | 1.935 | |
| Task-Aware Refinement | – | 22.2 | 70.2 | 0.222 | 0.111 | 0.485 | 2.393 | ||
| MimicX | 25.0 | 64.0 | 0.196 | 0.147 | 0.481 | 2.238 | |||
| Completion | Valid steps / | p95 reduction | ||||
| Task | Fixed | MimicX | Fixed | MimicX | Body | Root |
| Track Running | 3/3 | 3/3 | 97/97 | 97/97 | 53.80% | 24.95% |
| Stair Ascent | 3/3 | 3/3 | 449/449 | 449/449 | 11.86% | 7.88% |
| Forest Traversal | 0/3 | 3/3 | 39/404 | 404/404 | 79.15% | 89.04% |
| Platform Jump | 3/3 | 3/3 | 58/58 | 58/58 | 5.78% | 50.85% |
| Parkour | 0/3 | 0/3 | 220/255 | 220/255 | 2.04% | 0.59% |
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
| Study | Recorded evidence | Tasks | Methods | Train seeds | Eval. repeats |
| Core tracking | 48 policies; 144 rollouts | 4 | 4 | 3 | 3 |
| Repeated verification | 12 decisions; 6 accept, 6 protect | 4 | – | 3 | 3 |
| Same-motion trackers | 42 eligible recordings | 4 | 3–4 | 3 a | 1 |
| Collision scenes | 30 final rollouts | 5 | 2 | 1 | 3 |
| Additional videos | 4 paired task aggregates | 4 | 2 | – | – |
| Supplied motions | 11 LaFAN1; 3 ACCAD references | 14 | 2 | – | – |
| Method | Horizon (steps, ) | Reward ( ) | Body mean (m, ) | Body P95 (m, ) | Anchor error (m, ) | Linear vel. (m/s, ) | Angular vel. (rad/s, ) |
| Tennis Swing ( steps) | |||||||
| Fixed Reference | 322.0 302.2 | 0.0657 0.0023 | 0.241 0.020 | 0.518 0.125 | 0.301 0.023 | 0.566 0.053 | 2.287 0.031 |
| Failure Curriculum | 801.0 0.0 | 0.0816 0.0008 | 0.184 0.002 | 0.282 0.007 | 0.208 0.014 | 0.426 0.004 | 1.798 0.012 |
| Task-Aware Refinement | 768.0 57.2 | 0.0788 0.0014 | 0.178 0.011 | 0.293 0.012 | 0.177 0.023 | 0.454 0.003 | 2.074 0.027 |
| MimicX | 801.0 0.0 | 0.0750 0.0012 | 0.157 0.008 | 0.278 0.009 | 0.276 0.006 | 0.439 0.017 | 1.979 0.064 |
| Football Juggling ( steps) | |||||||
| Metric | Task macro (%, ) | Paired mean (%, ) | Wins | Paired median (%, ) | Minimum (%) | Maximum (%) |
| Reward | +39.44 | +41.96 | 12/12 | +39.97 | +10.58 | +91.90 |
| Robust execution horizon | +255.57 | +320.70 | 12/12 | +255.78 | +18.07 | +904.65 |
| Mean aligned worst-body error | +25.72 | +25.65 | 12/12 | +24.26 | +7.35 | +44.58 |
| 95th-percentile aligned worst-body error | +31.13 | +29.25 | 10/12 | +28.54 | -5.51 | +62.52 |
| Root-and-anchor error | +33.28 | +31.38 | 10/12 | +30.66 | -11.79 | +80.65 |
| Body linear-velocity error | +32.07 | +31.81 | 12/12 | +28.49 | +10.17 | +62.60 |
| Method | Training reward | Body error (m) | Anchor error (m) | ||||||
| Initial | Final | AUC | Initial | Final | AUC | Initial | Final | AUC | |
| Fixed Reference | 0.29 0.03 | 19.30 1.29 | 15.96 0.55 | 1.032 0.021 | 0.189 0.008 | 0.205 0.002 | 0.080 0.015 | 0.351 0.015 | 0.388 0.004 |
| Failure Curriculum | 0.23 0.01 | 23.70 1.91 | 18.45 0.18 | 1.159 0.355 | 0.152 0.013 | 0.188 0.001 | 0.095 0.004 | 0.261 0.023 | 0.377 0.002 |
| Task-Aware Refinement | 1.12 0.03 | 56.33 1.07 | 45.89 0.58 | 1.159 0.355 | 0.163 0.007 | 0.197 0.005 | 0.095 0.004 | 0.332 0.036 | 0.436 0.011 |
| MimicX | 4.36 0.35 | 126.25 1.59 | 103.74 0.88 | 0.800 0.741 | 0.143 0.018 | 0.181 0.019 | 0.056 0.009 | 0.359 0.035 | 0.457 0.023 |
| Task | Method | AUC | 20% | 40% | 60% | 80% | 100% | Body (m) |
| Tennis | Fixed Reference | 0.580 | 0.463 | 0.674 | 0.725 | 0.462 | 0.460 | 0.239 |
| Failure Curriculum | 0.963 | 0.705 | 1.000 | 1.000 | 1.000 | 1.000 | 0.183 | |
| Task-Aware Refinement | 0.941 | 0.845 | 0.861 | 1.000 | 1.000 | 0.959 | 0.183 | |
| MimicX | 0.921 | 1.000 | 1.000 | 0.842 | 0.842 | 1.000 | 0.156 | |
| Football | Fixed Reference | 0.170 | 0.566 | 0.113 | 0.097 | 0.130 | 0.118 | 0.286 |
| Failure Curriculum | 0.574 | 0.566 | 0.570 | 0.576 | 0.578 | 0.581 | 0.237 |
| Task | Method | Runs | Joint RMSE (rad) | Local body error (m) |
| Tennis | Fixed Reference | 3 | 0.389 0.166 | 0.151 0.072 |
| BeyondMimic (MjLab) | 3 | 0.389 0.027 | 0.137 0.021 | |
| SONIC (released) | 3 | 0.857 0.018 | 0.236 0.001 | |
| MimicX | 3 | 0.201 0.002 | 0.055 0.002 | |
| Football Juggling | Fixed Reference | 3 | 0.507 0.050 | 0.190 0.023 |
| BeyondMimic (MjLab) | 3 | 0.446 0.032 | 0.166 0.007 |
| Video task | Fixed Reference | MimicX | Error reduction (%, ) | success (pp, ) |
| Basketball | 0.204 | 0.217 | -6.22 | +0.0 |
| Soccer | 0.279 | 0.248 | +11.03 | +33.3 |
| Michael Jackson Dance | 0.413 | 0.326 | +21.04 | +0.0 |
| Uniandes Dance | 0.289 | 0.183 | +36.68 | +0.0 |
| LaFAN1: 11 evaluated references | |||||||
| Motion reference | Fixed | MimicX | Red. (%) | Motion reference | Fixed | MimicX | Red. (%) |
| Dance 1 | 0.2866 | 0.2499 | +12.81 | Jumps 1 | 0.2871 | 0.1606 | +44.06 |
| Dance 2 | 0.2961 | 0.2885 | +2.55 | Run 1 | 0.2523 | 0.1985 | +21.30 |
| Fall and Get Up 1 | 0.3746 | 0.3972 | -6.03 | Sprint 1 | 0.2861 | 0.2171 | +24.12 |
| Fall and Get Up 2 | 0.3540 | 0.3544 | -0.13 | Walk 1 | 0.2462 | 0.1419 | +42.36 |
| Fight 1 | 0.2988 | 0.2385 | +20.18 | Walk 4 | 0.2341 | 0.3008 | -28.51 |
| Case | Measurement | Before | After | Change |
| Tennis initialization | Terminations in 800 steps | 7 | 0 | Stable rollout |
| Tennis articulation | Mean max-body error (m) | 0.1411 | 0.1202 | 14.81% lower |
| Tennis articulation | Right-wrist error (m) | 0.0977 | 0.0497 | 49.13% lower |
| Tennis articulation | Left-wrist error (m) | 0.0932 | 0.0453 | 51.39% lower |
| Football execution | Terminations in 455 steps | 1 | 0 | Complete interval |
| Dance execution | Terminations in 784 steps | 1 | 0 | Complete interval |