Video is an abundant, inexpensive source of human motion data that is rich in extreme athletic behaviors. Making it usable for humanoid robots, however, is not a matter of simply retargeting a reconstructed trajectory: video-derived motion is physically inconsistent, devoid of actuation information, and says nothing about failure or recovery. We present KungfuAthleteBot (KAB), a framework that treats learning high-dynamic motion from video as the central problem and resolves each of these three failure modes in turn. (C1) We build the KungfuAthlete dataset from videos of national-level martial artists and introduce a physics-guided parabolic trajectory correction that removes height floating, ground penetration, and high-frequency jitter from reconstructed aerial and landing phases. (C2) Because video carries no force information, strict tracking of a reconstructed trajectory is dynamically infeasible, and error-driven initialization keeps re-launching the policy from infeasible aerial poses. We introduce physics-driven pseudo-low-kinetic-energy (LKE) sampling, our central mechanism for making such references learnable: it biases initialization towards dynamically feasible states, letting the policy discover feasible actuation patterns instead of imitating infeasible ones. (C3) Finally, we introduce a direct training paradigm in which disturbance rejection and fall recovery are learned inside the same policy that tracks the video motion, requiring no recovery reference data and no manual mode switching. On a humanoid robot, KAB learns dynamic skills from video and recovers from arbitrary falls in about 0.7 s, the fastest reported recovery for a unified policy. Ablations on the unified policy confirm the necessity of its components, supporting the view that repairing and compensating video data, rather than only collecting more of it, is what unlocks high-dynamic humanoid skills.
Figures & tables
Figure 1: Learning high-dynamic motion from video. (a) A dataset extracted from videos of national-level martial artists. (b) Video-derived motion carries no actuation information, so direct mimic-based tracking (e.g., BeyondMimic) fails on such references. Our curriculum (Sec. 4.3 ) and low-kinetic-energy sampling (Sec. 4.2 ) let the policy explore feasible actuation patterns. (c) A single policy performs fall recovery and motion tracking, and (d) rejects external disturbances.
Dataset
Body Lin Vel ( w )
Body Ang Vel ( w )
Frame Range
AMASS
0.188±0.165
0.510±0.309
[2,8788]
PHUMA
0.193±0.130
0.585±0.368
[49,222]
LAFAN1
0.790±0.251
1.962±0.553
[3065,9171]
KungfuAthlete (Ground)
0.496±0.324
1.537±0.959
[24,15792]
KungfuAthlete (Jump)
0.948±0.473
2.832±1.431
[25,2842]
Table 1: Kinematic statistics of motion datasets. Entries are means over the whole body, denoted by w in the column headings; frame range is given in frames.
Figure 2: Existing methods exhibit significant height floating when reconstructing jumping motions; prior work (e.g. PHUMA) therefore discards such data, wasting jump-related samples (top row). We instead enforce a physically consistent parabolic correction to refine the falling trajectories of jumping motions.
Figure 3: Overview of the training pipeline. Top-left: low-kinetic-energy (LKE) sampling draws video motion from dynamically feasible frames after a tracking failure. Bottom-left: the initial state q0 is drawn from a Bernoulli mixture over a reference frame and a gravity-based randomized fall state from DGRSI . Centre: reference-conditioned tracking rewards with recovery penalties (feet height, knee jerk, feet slip, root rotation). Right: the distributional actor–critic loop (FastSAC, Sec. 6 ).
Table 3: Configurations for noise and domain randomization.
Method
Video ref.
Aerial repair
Energy init.
KungfuBot Xie et al. (2025)
✓
✗
✗
OmniXtreme Wang et al. (2026)
✗
–
✗
KAB (ours)
✓
✓
✓
Table 4: Positioning on the video-to-robot pathway. Prior work assumes the processed reference is either learnable or safely discardable; we target the regime where it is neither. ✓ explicit support, ✗ absent, ‘–’ not applicable.
Motion ID
Method
Empjpe
Empjve
succ
78
TWIST
85.2 ± 23.8
18.9 ± 9.7
0/6
GMT
68.3 ± 17.3
11.7 ± 4.8
0/6
SONIC
71.9 ± 17.2
4.9 ± 4.6
0/6
KAB
24.7 ± 12.1
5.4 ± 4.0
6/6
213
TWIST
87.4 ± 21.0
21.1 ± 9.8
0/6
GMT
52.9 ± 12.8
9.7 ± 5.3
0/6
Table 5: Is the dataset challenging, and can KAB track it? Empjpe ( 10−3 rad), Empjve ( 10−3 rad/frame) and success rate over six trials. Motion IDs index clips in the released dataset. The same metrics are reported for KAB, to answer directly whether our own method can track these motions.
Figure 4: (a) Reward curves of the three-stage training and the single-stage baseline on Motion 468. (b) Episode length under LKE and adaptive sampling in Stage II on Motion 278. LKE reaches usable tracking behaviour markedly earlier than error-driven adaptive sampling.
Method
Reference- free
Tracking & Recovery
No Get-Up Cmd.
Recovery Time (s)
HumanUP He et al. (2025b)
✓
✗
✗
6.708 ± 0.987
HoST Huang et al. (2025)
✓
✗
✗
2.055 ± 0.310
FIRM Xu et al. (2025)
✗
✗
✓
1.729 ± 0.424
BFM-Zero Li et al. (2025b)
✗
✓
✓
1.566 ± 0.065
Heracles Tao et al. (2026)
✗
✓
✓
2.471 ± 0.139
StableMimic Wu et al. (2026)
✗
✓
✓
n/r
Table 6: Comparison of fall recovery performance and capability unification. Within the methods listed, only KAB combines a reference-free recovery objective with unified tracking and recovery, and KAB attains the shortest recovery time under the protocol above.
Figure 5: Fall recovery from diverse fallen configurations. Top four rows: MuJoCo recovery sequences covering prone, supine and two twisted joint configurations with irregular limb postures. Bottom two rows: the corresponding real-world recovery on the Unitree G1.
Figure 6: Training curves on Motion 1307. (a,b) Adding the recovery rewards does not change the episode length or the tracking error relative to the tracking-only baseline. (c,d) The full unified configuration matches the baseline’s episode length and tracking accuracy, at the cost of slower convergence ( ∼ 16k vs. ∼ 2.5k iterations). (e) Recovery components without mixed training: with no termination tolerance the episode length stays at 1, and with a 3 s tolerance it converges to only ∼ 150 of the full 500 steps, so the policy survives but never stands up and tracks. (f) Removing the height reward leaves partial tracking (300–400 steps) without recovery, while removing the root-orientation reward causes early collapse at ∼ 5k iterations.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Name
Expression
Weight
Stage-agnostic tracking reward (Stages I, II, III)
Body Dof Match
mincos(θdes−θ)
3.0
Penalty Action Rate
j∑∥aj−ajprev∥2
-0.01
Limits DOF Position
j∑[(clip(pj−upperj,0))+(clip(lowerj−pj,0))]
-10.0
Penalty Self Collision
1self-collision
-10.0
Stage II additional tracking reward
Appendix
Table 7: Physics-stability motion-tracking reward rmt . Rows under “Stage-agnostic” are active in all stages; rows under “Stage II additional” are added on top of them in Stage II. The recovery terms are active whenever the recovery indicator is triggered.
Name
Expression
Weight
Motion Relative Body Position Error
exp(−σ21B1b∑∥pbrel−p^b∥2)
1.0
Motion Relative Body Orientation Error
exp(−σ21B1b∑∥quat_error(qbrel,q^b)∥2)
1.0
Motion Global Body Position Error
exp(−σ21B1b∑∥pbglo−p^b∥2)
0.5
Motion Global Body Orientation Error
exp(−σ21B1b∑∥quat_error(qbglo,q^b)∥2)
1.0
Motion Global Body Angular Velocity
exp(−σ21B1b∑∥ωb−ω^b∥2)
0.4
Motion Global Body Linear Velocity
exp(−σ21B1b∑∥lb−l^b∥2)
1.0
Appendix
Table 8: Motion tracking reward borrowed from BeyondMimic, rmt .
Group
Term
Actor & critic
motion command (reference joint targets)
reference body orientation (anchor frame)
base angular velocity
joint positions
joint velocities
previous action
Appendix
Table 9: Observation terms. Actor and critic share the reference-conditioned and proprioceptive terms; the critic additionally receives privileged terms that are unavailable at deployment. Noise is injected into actor observations only.
Figure 7: Motion 78: a 720° spin in the air, followed by a 360° kick.
Figure 8: Motion 1080: a kick followed by a roundhouse kick.
Humanoid robots are promising to acquire various skills by imitating human behaviors. However, existing algorithms are only capable of tracking smooth, low-speed human motions, even with delicate reward and curriculum design. This paper presents a physics-based humanoid control framework, aiming to master highly-dynamic human behaviors such as Kungfu and dancing through multi-steps motion processing and adaptive motion tracking. For motion processing, we design a pipeline to extract, filter out, correct, and retarget motions, while ensuring compliance with physical constraints to the maximum extent. For motion imitation, we formulate a bi-level optimization problem to dynamically adjust the tracking accuracy tolerance based on the current tracking error, creating an adaptive curriculum mechanism. We further construct an asymmetric actor-critic framework for policy training. In experiments, we train whole-body control policies to imitate a set of highly-dynamic motions. Our method achieves significantly lower tracking errors than existing approaches and is successfully deployed on the Unitree G1 robot, demonstrating stable and expressive behaviors. The project page is https://kungfubot.github.io.
Weiji Xie, Jinrui Han, Jiakun Zheng +6
Institute of Artificial Intelligence (TeleAI), China Telecom · Shanghai Jiao Tong University · East China University of Science and Technology +2
Recent reinforcement learning approaches have shown great promise in improving humanoid motion tracking performance and achieving fall recovery under disturbances. However, most existing works treat motion tracking and fall recovery as different tasks and require multi-stage training with specialized recovery rewards and/or separate recovery policies. Moreover, existing reinforcement learning-based methods often terminate training episodes immediately after severe tracking failures, limiting recovery-oriented exploration in unstable or fallen states. To address the above issues, we propose Stubborn, a streamlined and unified reinforcement learning framework to achieve robust humanoid motion tracking and fall recovery. Specifically, Stubborn uses an asymmetric Actor-Critic architecture and consists of three major components. First, a yaw-aligned tracking representation is adopted to reduce sensitivity to global drift and heading disturbances while preserving gravity-related balance information. Second, we introduce a Bernoulli-based probabilistic termination mechanism that enables the policy to encourage exploration of fall-recovery behaviors under varying failure modes. Third, we propose a probabilistic termination and tracking-error-driven strategy that dynamically reshapes the sampling distribution based on tracking performance, increasing the training efficiency for difficult motion segments and unstable states. Extensive comparisons with SOTA methods and ablation studies show that Stubborn achieved competitive performance, and the proposed probabilistic termination mechanism and adaptive sampling strategy contributed to the performance and robustness gains. For real-world demonstrations, please refer to https://aislab-sustech.github.io/Stubborn/.
Xiao Ren, Yuhui Yang, Zongbiao Weng +2
Southern University of Science and Technology, Shenzhen 518055, China
Humanoid robots operating in unstructured environments must recover from unexpected disturbances-a capability that remains challenging for end-to-end control policies. We present RECOVERFORMER, a fully end-to-end humanoid recovery policy that learns when and how to switch among recovery behaviors-including compensatory stepping, hand-environment contact, and center-of-mass reshaping-while maintaining robust performance under model mismatch. The architecture combines a causal transformer over a 50-step observation history with two novel heads: a latent recovery mode that enables smooth transitions among distinct recovery strategies, and a contact affordance head that predicts which environmental surfaces (walls, railings, table edges) are beneficial for stabilization. We evaluate RECOVERFORMER on the Unitree G1 humanoid in MuJoCo. Trained only on open floor, RECOVERFORMER transfers zero shot to walled environments, achieving 100% recovery success across 100-300 N pushes and across wall distances from 0.25-1.4m. Under zero-shot dynamics mismatch, RECOVERFORMER reaches 75.5% at plus +25% mass, 89% under 30 ms latency, 91.5% at low friction, and 99% under compound friction, latency and mass perturbation. The learned latent modes specialize across force regimes without mode-level supervision, validated by t-SNE analysis of 300 episodes. Taken together, these results show that a single end-to-end policy can deliver multi-modal, contact aware humanoid recovery that generalizes across perturbation magnitude, contact geometry, and dynamics shift.