Human videos provide rich motion targets for humanoid learning, yet visually plausible references can still produce persistent failures under physics-based execution. These failures reveal where training supervision should change. We present MimicX, a policy-in-the-loop framework that uses execution feedback to refine video-driven humanoid motion tracking. Starting from reconstructed and retargeted motion, MimicX localizes difficult transitions and affected body regions, then jointly adapts tracking objectives and the reset curriculum for policy continuation. Repeated rollout verification selects execution-priority improvements subject to tracking guards. Across four core video tasks, MimicX consistently improves tracking accuracy and Robust Execution Horizon relative to the Fixed Reference baseline. Task-averaged results show a 25.7% reduction in body-tracking error and a 255.6% increase in execution horizon. Additional video, motion-reference, and collision-scene studies evaluate the method beyond the core tasks, while MimicX-HLoop accelerates feedback through heterogeneous execution. Overall, MimicX turns policy failure into actionable supervision for deciding what to refine and which refinement to retain.
Figures & tables
Figure 2: Execution feedback coordinates supervision. Rollout errors localize failure windows, affected bodies, and dominant local, global, or dynamic channels. These signals condition three proposals that coordinate tracking rewards and reset curricula. Each candidate continues PPO from the same checkpoint with a matched budget and seed. Repeated verification compares candidate execution with the current policy to select the retained update using execution-first criteria over repeated rollouts.
Figure 3: Verified updates with heterogeneous execution. Repeated rollouts rank candidates by completion, reliable motion length, and tracking quality. A body-error guard protects the current policy when no candidate qualifies. The same verification workload can run sequentially or through MimicX-HLoop. Its scheduler dispatches ready rollouts to GPUs and completed trajectories to CPU diagnosis workers, overlapping independent work while respecting dependencies. Reports are joined before selection; inputs, seeds, budgets, and the selection rule remain fixed. Execution order changes feedback latency while preserving the complete evidence supplied to the gate and its policy decision across all repeated evaluation seeds.
Figure 4: From video observations to humanoid execution. Tennis shows the input, recovered human motion, simulated policy rollout, and the corresponding policy pose in the reconstructed scene at a recovered source phase. Forest shows the input, recovered human motion, retargeted robot reference, and policy execution in the task-equivalent collision scene, sampled at each clip’s midpoint.
Figure 5: Controlled evidence across four video-derived skills. (a) All twelve task-matched trials across seven reward and tracking channels. (b) Task-wise reductions in mean and p95 tracking error, in percent. (c) Full-horizon completion across repeated evaluations. (d) Worst-case execution horizon. Black marks in (a) denote median and interquartile range. Fixed, Curriculum, and Task-aware abbreviate Fixed Reference, Failure Curriculum, and Task-Aware Refinement; gray, blue, green, and coral identify the four methods (Curric. and Task in (c)). Tail denotes body-error p95 and Ankle denotes ankle-height error.
Task
Body mean (m)
Body p95 (m)
Horizon (steps)
Success (%)
Fixed
MimicX
Red. (%)
Fixed
MimicX
Fixed
MimicX
Gain (%)
Fixed
MimicX
Tennis
0.241
0.157
35.0
0.518
0.278
322.0
801.0
+148.8
11.1
100.0
Football
0.286
0.244
14.7
0.612
0.504
53.7
427.3
+696.3
0.0
0.0
Dance
0.270
0.244
9.9
0.502
0.467
193.3
341.3
+76.6
0.0
0.0
Kung Fu
0.249
0.141
43.3
0.634
0.296
97.7
196.0
+100.7
0.0
0.0
Table 1: MimicX improves tracking and execution horizon on all four tasks. Means are over three continuation seeds, each evaluated with three fixed rollout seeds. Body mean and p95 summarize aligned worst-body error over time; horizon is the earliest failure across repeats, with T+1 assigned on success. Reduction and gain are relative to Fixed Reference, computed before rounding. Bold marks improvements.
Figure 6: More accurate policies for the same demonstrated motion. Common-reference evaluation of Fixed Reference, BeyondMimic’s MjLab reproduction (BM), released SONIC, and MimicX. (a) Tennis joint-error mean and range over time; (b) Football control-step body-error distribution; (c) Tennis body error in cm for each recording; (d) mean Football joint errors with every recording. Lower errors are better. All methods are evaluated over the same 518-step Tennis and 454-step Football intervals without failure resets; body error uses root-local G1 forward kinematics. Full protocols and task outcomes appear in Table 16 .
Method
Policy setup
Joint RMSE (rad, ↓ )
Reduction (%)
Local body error (m, ↓ )
Reduction (%)
Fixed Reference
Fixed objective
0.389 ± 0.166
+0.0
0.151 ± 0.072
+0.0
BeyondMimic (MjLab)
Fixed objective
0.389 ± 0.027
+0.1
0.137 ± 0.021
+9.8
SONIC (released)
Pretrained
0.857 ± 0.018
-120.2
0.236 ± 0.001
-55.8
MimicX
Closed loop
0.201 ± 0.002
+48.3
0.055 ± 0.002
+63.4
Table 2: Direct policy comparison on the same Tennis Swing reference. MimicX improves both common articulation metrics across all three seeds relative to BeyondMimic’s MjLab reproduction. Bold marks minima. Reductions use Fixed Reference as the baseline.
Figure 7: Execution differences on the same motion timeline. Each case compares BeyondMimic’s MjLab reproduction (blue) and MimicX (coral) at 20/50/80% of the source interval. Recorded policies use continuation seed 202 and evaluation seed 1001, with matched camera and crop settings within each task.
Figure 8: Kung Fu from video to recorded execution. Seven input frames show the selected motion interval. The lower row progresses from recovered human motion to recorded MimicX policy states. Source and rollout timestamps identify the sampled interval; the complete-task evaluation is reported in Table 1 .
Method
Supervision
Execution
Tracking
Dynamics
C
T
V
Success (%, ↑ )
Horizon (%, ↑ )
Body (m, ↓ )
Anchor (m, ↓ )
Linear vel. (m/s, ↓ )
Angular vel. (rad/s, ↓ )
Fixed Reference
–
–
–
2.8
21.5
0.262
0.219
0.691
2.955
Failure Curriculum
✓
–
–
25.0
62.7
0.215
0.135
0.451
1.935
Task-Aware Refinement
✓
✓
–
22.2
70.2
0.222
0.111
0.485
2.393
MimicX
✓
✓
✓
25.0
64.0
0.196
0.147
0.481
2.238
Table 3: Supervision components improve complementary tracking channels. Values are unweighted means across four tasks after averaging three continuation seeds within each task. Bold marks column bests.
Table 11
Figure 9: Coordinated supervision changes the learning trajectory. Reward, body error, and episode length across twelve 250-update Tennis continuations. Thin traces retain raw values; emphasized curves use an 11-update moving mean, with sample standard deviation across seeds. A single legend applies to all three panels.
Completion
Valid steps / T
p95 reduction
Task
Fixed
MimicX
Fixed
MimicX
Body
Root
Track Running
3/3
3/3
97/97
97/97
53.80%
24.95%
Stair Ascent
3/3
3/3
449/449
449/449
11.86%
7.88%
Forest Traversal
0/3
3/3
39/404
404/404
79.15%
89.04%
Platform Jump
3/3
3/3
58/58
58/58
5.78%
50.85%
Parkour
0/3
0/3
220/255
220/255
2.04%
0.59%
Table 6: Policy-in-the-loop refinement on five collision-scene video tasks. Equal-budget continuations share the same initial policy, scene, and reference. Completion rises from 9/15 to 12/15 full executions; both p95 errors improve on all five tasks. Steps report the minimum valid prefix divided by target length; reductions compare seed-mean errors over matched prefixes before failure.
Figure 10: Video-driven skills in collision scenes. Stair Ascent compares Fixed Reference and refinement at 20/50/80% of the 449-step schedule, using evaluation seed 2002 and time-specific close crops shared by both methods. Both complete all three evaluations; refinement reduces body and root prefix p95 errors by 11.86% and 7.88%. Platform Jump shows each clip’s midpoint: input, human recovery, robot reference, and policy execution.
Table 15
Figure 11: Reliable refinement and faster feedback across motion inputs. (a) Execution verification rejects reward-improving regressions. (b) Five paired wall-time measurements: connectors link sequential execution and HLoop; triangles show bulk-synchronous execution. (c) Per-job durations: rollout and diagnosis histograms, five selection measurements, and median markers; each stage has its own linear axis and units. (d) BeyondMimic minus MimicX body error over the complete 454-step Football interval, averaged over three recordings; positive values indicate improved tracking.
Appendix figures & tables25 assets
Supplementary material from the paper’s appendix.
Appendix
Study
Recorded evidence
Tasks
Methods
Train seeds
Eval. repeats
Core tracking
48 policies; 144 rollouts
4
4
3
3
Repeated verification
12 decisions; 6 accept, 6 protect
4
–
3
3
Same-motion trackers
42 eligible recordings
4
3–4
3 a
1
Collision scenes
30 final rollouts
5
2
1
3
Additional videos
4 paired task aggregates
4
2
–
–
Supplied motions
11 LaFAN1; 3 ACCAD references
14
2
–
–
Appendix
Table 9: Experimental coverage and independent sampling units. Training seeds, repeated evaluations and task aggregates are distinguished explicitly.
Method
Horizon (steps, ↑ )
Reward ( ↑ )
Body mean (m, ↓ )
Body P95 (m, ↓ )
Anchor error (m, ↓ )
Linear vel. (m/s, ↓ )
Angular vel. (rad/s, ↓ )
Tennis Swing ( H=800 steps)
Fixed Reference
322.0 ± 302.2
0.0657 ± 0.0023
0.241 ± 0.020
0.518 ± 0.125
0.301 ± 0.023
0.566 ± 0.053
2.287 ± 0.031
Failure Curriculum
801.0 ± 0.0
0.0816 ± 0.0008
0.184 ± 0.002
0.282 ± 0.007
0.208 ± 0.014
0.426 ± 0.004
1.798 ± 0.012
Task-Aware Refinement
768.0 ± 57.2
0.0788 ± 0.0014
0.178 ± 0.011
0.293 ± 0.012
0.177 ± 0.023
0.454 ± 0.003
2.074 ± 0.027
MimicX
801.0 ± 0.0
0.0750 ± 0.0012
0.157 ± 0.008
0.278 ± 0.009
0.276 ± 0.006
0.439 ± 0.017
1.979 ± 0.064
Football Juggling ( H=455 steps)
Appendix
Table 10: Four-task continuation comparison. Entries show mean ± sample SD over three continuation-seed trials, each aggregating three evaluation rollouts. Bold marks each task–metric best mean, including ties.
Metric
Task macro (%, ↑ )
Paired mean (%, ↑ )
Wins
Paired median (%, ↑ )
Minimum (%)
Maximum (%)
Reward
+39.44
+41.96
12/12
+39.97
+10.58
+91.90
Robust execution horizon
+255.57
+320.70
12/12
+255.78
+18.07
+904.65
Mean aligned worst-body error
+25.72
+25.65
12/12
+24.26
+7.35
+44.58
95th-percentile aligned worst-body error
+31.13
+29.25
10/12
+28.54
-5.51
+62.52
Root-and-anchor error
+33.28
+31.38
10/12
+30.66
-11.79
+80.65
Body linear-velocity error
+32.07
+31.81
12/12
+28.49
+10.17
+62.60
Appendix
Table 11: Paired improvements of MimicX over Fixed Reference. All eight recorded metrics are retained. Relative changes are percentages, with positive values denoting reward/horizon increases or error reductions; wins count strictly positive task–continuation-seed pairs.
Table 20
Method
Training reward
Body error (m)
Anchor error (m)
Initial
Final
AUC
Initial
Final
AUC
Initial
Final
AUC
Fixed Reference
0.29 ± 0.03
19.30 ± 1.29
15.96 ± 0.55
1.032 ± 0.021
0.189 ± 0.008
0.205 ± 0.002
0.080 ± 0.015
0.351 ± 0.015
0.388 ± 0.004
Failure Curriculum
0.23 ± 0.01
23.70 ± 1.91
18.45 ± 0.18
1.159 ± 0.355
0.152 ± 0.013
0.188 ± 0.001
0.095 ± 0.004
0.261 ± 0.023
0.377 ± 0.002
Task-Aware Refinement
1.12 ± 0.03
56.33 ± 1.07
45.89 ± 0.58
1.159 ± 0.355
0.163 ± 0.007
0.197 ± 0.005
0.095 ± 0.004
0.332 ± 0.036
0.436 ± 0.011
MimicX
4.36 ± 0.35
126.25 ± 1.59
103.74 ± 0.88
0.800 ± 0.741
0.143 ± 0.018
0.181 ± 0.019
0.056 ± 0.009
0.359 ± 0.035
0.457 ± 0.023
Appendix
Table 14: Dense Tennis learning-log summaries from all 3,000 records: four methods, three continuation seeds, and 250 recorded updates per trajectory. Entries show mean ± sample SD, computed across seeds after obtaining each trajectory’s endpoint or integral. Training reward retains its logged scale and differs from evaluation reward.
Task
Method
AUC
20%
40%
60%
80%
100%
Body (m)
Tennis
Fixed Reference
0.580
0.463
0.674
0.725
0.462
0.460
0.239
Failure Curriculum
0.963
0.705
1.000
1.000
1.000
1.000
0.183
Task-Aware Refinement
0.941
0.845
0.861
1.000
1.000
0.959
0.183
MimicX
0.921
1.000
1.000
0.842
0.842
1.000
0.156
Football
Fixed Reference
0.170
0.566
0.113
0.097
0.130
0.118
0.286
Failure Curriculum
0.574
0.566
0.570
0.576
0.578
0.581
0.237
Appendix
Table 15: Execution quality throughout policy continuation. Every cell averages three continuation seeds. The normalized horizon AUC integrates five checkpoints from 20% to 100% of the recorded continuation; the checkpoint columns report horizon fractions, followed by final body error. The four-task, four-method, three-seed design gives 48 curves and 240 evaluations.
Figure 12: Aligned motion-state traces. Orientation, acceleration and angular velocity share a clock across 800 recorded Tennis states. Acceleration uses position differences; angular velocity uses quaternion increments. Derivative stencils crossing resets are excluded. Raw values and 11-sample means distinguish short transients from the motion trend.
Task
Method
Runs
Joint RMSE (rad)
Local body error (m)
Tennis
Fixed Reference
3
0.389 ± 0.166
0.151 ± 0.072
BeyondMimic (MjLab)
3
0.389 ± 0.027
0.137 ± 0.021
SONIC (released)
3
0.857 ± 0.018
0.236 ± 0.001
MimicX
3
0.201 ± 0.002
0.055 ± 0.002
Football Juggling
Fixed Reference
3
0.507 ± 0.050
0.190 ± 0.023
BeyondMimic (MjLab)
3
0.446 ± 0.032
0.166 ± 0.007
Appendix
Table 16: Same-reference policy evaluation across motions. Mean ± sample SD; lower is better. Bold marks the lowest task-wise mean.
Video task
Fixed Reference
MimicX
Error reduction (%, ↑ )
Δ success (pp, ↑ )
Basketball
0.204
0.217
-6.22
+0.0
Soccer
0.279
0.248
+11.03
+33.3
Michael Jackson Dance
0.413
0.326
+21.04
+0.0
Uniandes Dance
0.289
0.183
+36.68
+0.0
Appendix
Table 17: Additional video-task outcomes. Body error is reported in meters ( ↓ ); bold marks the lower error within each task. Values are recorded task aggregates. Success change is in percentage points (pp).
Figure 13: Complementary tracking diagnostics across recordings and tasks. (a) Tennis joint-error means for all three recordings and their range. (b) Distribution of all 3×518 Tennis body-error measurements per method. (c) Football joint-error means at each of 454 control steps. (d) Body-error reductions on 18 additional video and supplied-motion cases, grouped by input source; the full distribution includes negative changes. BM denotes the standalone BeyondMimic MjLab tracker; ACCAD includes the Form 1 sequence. Panels (a–c) use the common-reference comparison, while (d) compares MimicX with Fixed Reference on additional tasks.
LaFAN1: 11 evaluated references
Motion reference
Fixed
MimicX
Red. (%)
Motion reference
Fixed
MimicX
Red. (%)
Dance 1
0.2866
0.2499
+12.81
Jumps 1
0.2871
0.1606
+44.06
Dance 2
0.2961
0.2885
+2.55
Run 1
0.2523
0.1985
+21.30
Fall and Get Up 1
0.3746
0.3972
-6.03
Sprint 1
0.2861
0.2171
+24.12
Fall and Get Up 2
0.3540
0.3544
-0.13
Walk 1
0.2462
0.1419
+42.36
Fight 1
0.2988
0.2385
+20.18
Walk 4
0.2341
0.3008
-28.51
Appendix
Table 18: Motion-reference breadth: fourteen evaluated references grouped by source dataset. Body error is reported in meters ( ↓ ); bold compares methods within a row. Cohort reductions average per-reference relative reductions rather than taking the ratio of cohort-mean errors.
Figure 14: Execution on additional task-equivalent collision scenes. Fixed Reference and MimicX at matched 20/50/80% source phases, evaluation seed 2002. Crops are identical between methods at each time. Track Running completes all repeated evaluations with lower body and root errors after refinement. Parkour panels depict the shared pre-failure interval; its full-horizon outcome is reported in Table 6 .
Case
Measurement
Before
After
Change
Tennis initialization
Terminations in 800 steps
7
0
Stable rollout
Tennis articulation
Mean max-body error (m)
0.1411
0.1202
14.81% lower
Tennis articulation
Right-wrist error (m)
0.0977
0.0497
49.13% lower
Tennis articulation
Left-wrist error (m)
0.0932
0.0453
51.39% lower
Football execution
Terminations in 455 steps
1
0
Complete interval
Dance execution
Terminations in 784 steps
1
0
Complete interval
Appendix
Table 19: Execution-guided refinement in developmental case studies. Source-verified before/after transitions show complementary improvements in stability, task-critical articulation and motion completion. These are selected single-rollout mechanism cases, separate from the controlled multi-seed results.
Figure 15: Tennis across time and embodiment. Seven original input frames accompany three reconstructed human states and four policy states. Source phases associate the two rows.
Figure 16: Football Juggling across time and embodiment. Original-video frames provide the scene and action context for the reconstructed human motion and policy sequence below.
Figure 17: Dance across time and embodiment. Seven original-video frames show the selected interval; the lower row follows human reconstruction and policy execution.
Figure 18: Tennis task-critical execution. Fixed Reference and MimicX at four recorded event times. The neutral robot is the policy and the translucent method-colored robot is the reference overlay. Gray identifies Fixed Reference and coral identifies MimicX.
Figure 19: Football Juggling task-critical execution. Four recorded event times compare policy alignment with the reference motion. Alternating support and raised legs reveal whole-body juggling coordination.
Figure 20: Task-critical execution on articulated whole-body motions. Fixed Reference and MimicX are compared at matched events. Solid robots show execution; translucent robots show method-colored references.
Figure 21: Source-view correspondence in Track Running and Stair Ascent. Columns show input video, reconstructed human, and G1 reference; rows follow five synchronized instants.
Figure 22: Video-driven motion in task scenes. Tennis (top) and Football Juggling (bottom), each shown as seven recorded robot states. Spatial offsets separate the poses and increasing opacity encodes chronology. The selected Tennis sequence is MimicX; the Football sequence is Fixed Reference. Scene objects provide presentation context for the recorded motion.
Figure 23: Additional Soccer video task. Fixed Reference and MimicX at four recorded events, with method-colored reference overlays. This task is separate from the core Football Juggling video.
Figure 24: Uniandes Dance video task. Four task-critical events show policy–reference alignment.
Figure 25: AMASS/ACCAD Form 1: provided retargeting. Fixed Reference and MimicX on the provided robot-motion reference, with recorded policy states and method-colored reference overlays. Source attribution follows the motion audit in Table 18 . This pathway starts from an existing motion reference.
Figure 26: AMASS/ACCAD Form 1: new retargeting. A second retargeting of the same human source motion, evaluated through the shared policy-refinement interface. The two plates compare executions from different robot references derived from this common human motion.
Learning humanoid skills from videos typically requires a successful human demonstration, which often demands custom data collection. Although failures have traditionally been treated only as negative examples in robot learning, they can still reveal a usable trajectory prefix before the task fails, as well as the intended outcome. To leverage this information from a failed-attempt video, we propose TRACC, a pipeline that imitates the useful portion of the motion trajectory and then completes the task based on the inferred task outcome. The usable motion prefix serves as prior knowledge until the failure occurs, after which the task-completion reward guides the policy to learn the intended task goal without requiring a successful task trajectory. We evaluate our method on six in-the-wild failed human tasks from the Oops! dataset. Our experimental results demonstrate the effectiveness of the proposed approach for learning from failed attempts when no successful demonstration is available. Thus, these findings establish failed human videos as a viable source of supervision for humanoid skill learning.
Sarmad Idrees, Jongeun Choi
School of Mechanical Engineering, Yonsei University, Seoul 03722, Korea
Humanoid motion tracking is central to teleoperation and whole-body imitation, yet evaluation often disagrees with what people perceive in videos. Kinematic errors average per-frame pose differences but miss the physical artifacts that matter most, particularly unstable support and incorrect contacts such as foot skating and mistimed touch-downs. Meanwhile, widely used test suites are small and lack the diversity needed to stress contact-rich, long-horizon behaviors. We introduce HumanTracker to make humanoid tracking evaluation both perceptually aligned and scalable. The HumanTracker benchmark contains approximately 153 hours of optical motion trajectories from multiple professional performers, organized into four motion families with text labels for fine-grained diagnosis. We further propose HumanScore, a preference-aligned metric trained on 12K motion pairs containing 24K motions. Across representative state-of-the-art trackers, HumanScore better predicts human preferences and reveals contact and stability failures that kinematic metrics often miss.
Dairu Liu, Zekun Qi, Jiayu Zeng +11
1Nankai University · 3Galbot · 2Tsinghua University +3
Humanoid motion trackers perform reliably within learned tracking distributions, but falls can move the robot into low-height, contact-rich states from which an advancing command is temporarily unreachable. Tracking-only policies may chase infeasible references, producing rapid, large-amplitude limb corrections that increase risk to the robot and its surroundings. We present StableMimic, a unified tracker trained beyond the nominal tracking distribution. Perturbed resets around multiple human get-up references expose prone, supine, off-balance, and intermediate ground-contact states, shaping structured recovery that returns the robot to the trackable region. Because tracking and recovery occupy markedly different state--action distributions, StableMimic uses dedicated experts for each regime and a proprioceptive gate that continuously blends their actions. A hidden successor-state objective teaches human-reference-shaped recovery without exposing reference identity or phase to the deployed Actor; deployment requires no get-up reference, recovery command, trajectory retrieval, or external policy switch. On the complete retargeted LAFAN1 dance subset, StableMimic achieves the lowest errors on all four tracking metrics among five methods. Across 100 matched push-to-fall trials per method, it recovers in 100/100 and attains the lowest values on six of seven post-fall motion and load measures, supporting improved interaction safety under this protocol. Real Unitree G1 dance and standing-reference deployments qualitatively demonstrate bounded limb motion, autonomous recovery, and command resumption.
Weihao Wu, Ming Huang, Ruofei Liu +3
Southern University of Science and Technology (SUSTech), Shenzhen, China.