Motion capture data is rarely directly usable, as they typically exhibit missing segments, jitter, drift and contact artifacts. Traditionally, corrupted motions are cleaned by animators through the manual identification of keyframes from noisy motion, subsequent keyframe correction, and interpolation between corrected keyframes to reconstruct coherent motion. While the rise of generative motion models has made automatic cleanup feasible, most approaches operate as black box denoisers with limited controllability, making it difficult to preserve reliable segments or enforce specific user intents. Inspired by animation workflows, we present CleanMDM, a unified multimodal motion cleanup framework that formulates cleanup as masked conditional generation with plug-and-play conditions. This single model supports arbitrary combinations of noisy 3D motion, sparse 2D keyframes, sparse 3D keyframes, and text. This design enables both automatic cleanup without additional user annotation and controllable cleanup under multimodal guidance. To further improve motion realism, we incorporate the Latent Motion Quality Discriminator (LMQD) to better match kinematic distributions and reduce skating, jitter, and interpenetration artifacts, and we apply Mesh-Aware Contact Projection as a test-time optimization step to enhance contact and physical consistency. Experiments across multiple datasets demonstrate that CleanMDM consistently outperforms prior cleanup and generation baselines, and that low cost conditions (text and 2D keyframes) provide reliable controllability gains in multimodal cleanup scenarios.
Figures & tables
Figure 1: Overview of motion cleanup. Top: the input motion with a localized corrupted interval, represented as masked noisy motion. Middle: the same clip augmented with optional, plug-and-play conditions: corrupted frames are highlighted in red for visualization, a reference image provides sparse 2D keyframes, white silhouettes indicate sparse 3D keyframes, and a text prompt specifies high-level intent. Bottom: CleanMDM generates a clean and temporally consistent motion sequence.
Figure 2: CleanMDM framework and architecture. Left: CleanMDM formulates motion cleanup as masked conditional generation in the autoencoder latent space, using a noisy 3D draft with frame-wise confidence and optional plug-and-play cues (sparse 2D/3D keyframes and text). Solid-bordered inputs are required; dashed-bordered inputs are optional. Visualization: semi-transparent skeletons indicate masked-out (unobserved) regions, and red skeletons indicate corrupted frames in the noisy draft. LMQD is used only during training for latent adversarial regularization, and an optional Mesh-Aware Contact Projection post-process is applied at inference to reduce contact artifacts. Right: CleanMDM block with frame weighted geometric fusion injected via AdaLN+gated residuals into self-attention (RoPE) and FFN; text is injected through cross-attention.
Setting
Method
MPJPE(mm) ↓
Skate ↓
Acc Err ↓
Pen.F(%) ↓
Pen.D(mm) ↓
Jitter ↓
KF Err(mm) ↓
FID ↓
Div →
M2M Score ↑
M2M R@3 ↑
Auto cleanup
GT
0.00
0.13
0.00
0.00
0.00
3.73
0.00
0.00
9.40
1.00
100%
StableMotion
59.36
0.20
2.24
0.00
1.25
7.90
✗
0.12
7.09
0.95
90.72%
CleanMDM
18.25(8.84)
0.18
2.04
0.03
0.29
6.28
✗
0.12
9.17
0.99
99.92%
CleanMDM + MACP
45.15(19.91)
0.14
1.88
0.00
0.00
3.74
✗
0.24
9.16
0.96
98.73%
Inbetween
GT
0.00
0.12
0.00
0.00
0.00
3.15
0.00
0.00
9.44
1.00
100%
ACMDM
54.4
0.21
4.13
0.00
0.54
10.71
21.9
0.39
9.65
0.95
95.11%
Table 1: Comparison to prior work under three evaluation settings. We report results for auto cleanup (auto cleanup from a noisy draft, no interval annotation), inbetween (mask guided interval completion), and multimodal cleanup (interval completion with additional 2D/3D keyframes and text). MPJPE(mm) reports global error, with local MPJPE in parentheses (root effects removed). Skate denotes foot skating ratio; Acc and Jitter measure second- and third-order temporal errors; Pen.F(%) and Pen.D(mm) denote ground penetration frequency and distance.
Method
MPJPE ↓
Skate ↓
Acc ↓
Pen.F ↓
Pen.D ↓
Jitter ↓
KF Err ↓
FID ↓
Div →
GT
0.00
0.13
0.00
0.00
0.00
3.73
0.00
0.00
9.40
C0
35.45
0.19
2.12
0.07
0.95
5.73
12.7
0.09
8.04
C1
34.33
0.21
2.27
0.08
0.84
6.26
13.2
0.07
7.95
C2
27.45
0.20
2.17
0.07
0.89
6.12
13.3
0.07
8.05
C3
27.32
0.20
2.19
0.06
0.78
6.17
13.0
0.08
8.06
C4
28.87
0.19
2.07
0.05
0.72
5.74
12.2
0.13
9.17
Table 2: Main ablation on multimodal conditions and regularizers. C0–C6 differ only in the enabled conditions and post-processing (see Table 4 ).
Table 5
Appendix figures & tables13 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 3: Qualitative comparisons with task-matched baselines. Top: multimodal cleanup compared with GenMo. The input consists of an unmasked noisy draft (yellow), sparse 3D keyframes (gray bodies), sparse 2D keyframes (blue skeleton), and a text prompt. Bottom-left: auto cleanup compared with StableMotion; CleanMDM corrects foot sliding (circled) while preserving the remaining motion. Bottom-right: inbetween compared with CondMDI; CleanMDM better satisfies the 3D keyframe constraints (yellow translucent overlays) and yields smoother transitions. Rows show GT , Input , Baseline , and Ours .
Source Group
Dataset
Used
Motion LLaMA
HumanSC3D
✓
Motion LLaMA
FineDance
✓
Motion LLaMA
Fit3D
✓
Motion LLaMA
Hi4D
✓
Motion LLaMA
InterHuman
✗
Motion LLaMA
InterX
✓
Appendix
Table 5: MotionLLaMA-related subsets shown in the data manager.
AIST++ (128)
animation_ref (20)
Method
Skate ↓
Jitter ↓
Pen.F ↓
Pen.D ↓
Succ. ↑
Skate ↓
Jitter ↓
Pen.F ↓
Pen.D ↓
Succ. ↑
GVHMR (CCD-IK)
0.13
23.1513
14.20
48.46
1.00
0.2594
48.0547
17.53
40.41
1.00
GVHMR (CCD-IK) + PHC
0.0913
30.7293
0.00
0.95
0.867
0.0725
34.2244
0.00
1.31
0.35
GVHMR + MACP (ours)
0.056
27.3287
0.49
1.63
1.00
0.0583
63.5704
0.00
0.78
1.00
Appendix
Table 6: Effect of Mesh-Aware Contact Projection on video-mocap outputs. We compare different post-processing strategies applied to GVHMR results on AIST++ and animation_ref. Our Mesh-Aware Contact Projection improves contact quality (lower skating and penetration) and maintains a high success ratio, indicating stable optimization behavior across datasets.
Loss Category
Loss Term
Weight
Reconstruction
Rotation preservation
1
Reconstruction
Horizontal velocity preservation (non-contact)
10−4
Reconstruction
Vertical velocity preservation (non-contact)
10−3
Foot contact
Sole mesh height consistency
10−1
Foot contact
Sole mesh non-penetration
10−1
Foot contact
Zero velocity for contacting sole vertices
10−1
Appendix
Table 7: Loss weights used in Mesh-Aware Contact Projection. The table expands the implementation weights for Eq. 8 .
Figure 4: Effect of test-time Mesh-Aware Contact Projection on video mocap outputs. Top: raw GVHMR Shen et al. [2024] results exhibit contact artifacts, including foot sliding and occasional floating (highlighted). Bottom: applying our Mesh-Aware Contact Projection as a post-process enforces more consistent ground contact, reducing sliding and correcting floating while preserving the overall motion.
Method
MPJPE ↓
Skate ↓
Acc Err ↓
Pen.F ↓
Pen.D ↓
Jitter ↓
FID ↓
Div →
GT
0.00
0.13
0.00
0.00
0.00
3.73
0.00
9.40
C0
13.8
0.20
1.87
0.07
1.09
5.77
0.10
6.53
C5
12.8
0.19
1.78
0.03
0.79
5.49
0.06
8.14
Appendix
Table 8: Multimodal training improves noisy-only denoising. C0 and C5 are evaluated on the same auto repair setup with identical noisy inputs and no mask guidance or extra modalities at test time. Lower is better ( ↓ ) except Div ( ↑ ).
Method
MPJPE ↓
Skate ↓
Acc Err ↓
Pen.F ↓
Pen.D ↓
Jitter ↓
FID ↓
Div →
GT
0.00
0.13
0.00
0.00
0.00
3.73
0.00
9.40
C2 (2D MLP; ours)
27.45
0.20
2.17
0.07
0.89
6.12
0.07
8.05
C2 (MotionBERT 2D)
31.45
0.21
2.24
0.10
0.89
6.25
0.12
9.20
Appendix
Table 9: Ablation on the 2D condition encoder. All settings are identical except for the 2D branch.
Method
MPJPE ↓
Skate ↓
Acc Err ↓
Pen.F ↓
Pen.D ↓
Jitter ↓
KF Err ↓
FID ↓
Div →
GT
0.00
0.13
0.00
0.00
0.00
3.73
0.00
0.00
9.40
C5 (MLP)
35.59
0.10
14.94
0.08
0.73
71.02
30.44
2.99
8.91
C5 (CleanMDM)
21.95
0.18
1.98
0.06
0.76
5.54
11.4
0.11
9.23
Appendix
Table 10: Latent-space vs. joint-space 3D conditioning (multimodal cleanup). C5 (MLP) replaces the autoencoder-based 3D keyframe encoder with an MLP in joint space. All other components are identical to C5.
Method
MPJPE ↓
Skate ↓
Acc Err ↓
Pen.F ↓
Pen.D ↓
Jitter ↓
C0 (diffusion only)
13.8
0.2008
1.87
0.07
1.09
6.5286
C0+LMQD (latent seq, ours)
11.8
0.1780
1.96
0.02
0.46
6.3115
C0+LMQD (power seq)
13.3
0.1813
1.96
0.05
0.91
6.4056
C0+LMQD (joint seq)
13.3
0.1843
2.01
0.04
0.65
6.5149
C0+LMQD (AMP seq)
13.3
0.1847
2.01
0.04
0.69
6.5257
C0+LMQD (accel seq)
13.3
0.1848
2.01
0.04
0.66
6.5167
Appendix
Table 11: Ablation on LMQD feature representations (denoising). All variants use the same denoising setup; only the discriminator feature is changed.
Dataset
Method
MPJPE (mm) ↓
Jitter ↓
Slide R ↓
Slide A ↓
IDEA400
Input
0.00
5.43
0.30
0.06
StableMotion
40.10
5.10
0.45
0.14
CleanMDM
15.61
4.94
0.25
0.05
Animation
Input
0.00
17.60
0.54
0.15
StableMotion
138.61
11.27
0.60
0.11
CleanMDM
69.40
15.95
0.50
0.11
Appendix
Table 12: Generalization on public datasets. We report MPJPE to the real-world input motion (lower indicates less drift), together with temporal smoothness (Jitter) and contact quality (Slide R/A). These public datasets do not provide ground-truth cleaned motions. CleanMDM consistently reduces sliding and jitter while maintaining fidelity to the real-world input motion, indicating a transferable motion prior that improves cleanup without over-smoothing.
Figure 5: (a) Real-world cleanup on KungFu Tornado Kick (Motion-X) Lin et al. [2023] . Top to bottom: real-world input, StableMotion Mu et al. [2025] , and CleanMDM (ours) . Orange bodies mark low-quality frames in the real-world input, while purple bodies mark the corresponding repaired frames in StableMotion and CleanMDM. StableMotion shows visible support-foot floating and dampened rotation; CleanMDM closely tracks the input motion. See the preceding paragraph for the MPJPE/Jitter trade-off, and Fig. 6 for a second example on the same dataset.
Figure 6: (b) Additional KungFu example: Play basketball . Top to bottom: real-world input, StableMotion Mu et al. [2025] , and CleanMDM (ours) . Orange bodies mark low-quality frames in the real-world input, while purple bodies mark the corresponding repaired frames in StableMotion and CleanMDM. StableMotion lifts the subject off the ground in one of these frames, breaking the grounded action semantics; CleanMDM preserves the input contact behavior. Companion to Fig. 5 .
Setting
Method
Artifacts ↑
Naturalness ↑
Faithfulness ↑
Auto cleanup
StableMotion
3.45
3.61
2.33
CleanMDM
4.13
4.41
4.63
Inbetween
CondMDI
2.91
3.08
2.29
CleanMDM
3.93
4.07
4.33
Multimodal cleanup
GenMo
3.20
3.50
3.66
CleanMDM
3.93
4.12
4.22
Appendix
Table 13: Complete user-study results. Mean ratings from 10 professional animators. Higher is better.
Text-driven human motion editing aims to realize a requested change while preserving compatible source content. Existing diffusion editors rely largely on learned conditioning for preservation of the unedited part, yet their outputs can lose temporal detail as denoising proceeds. We propose the \textbf{Source-Preserving Denoising framework (SuperMotion)}, which explicitly reuses the source at each reverse step for source preservation. We first align the source motion to the output timeline and predict a preservation gate that controls reuse across frames and feature dimensions. A clean-space source anchor then utilizes the learned preservation gate to blend the predicted clean motion with the aligned source and passes the corrected estimate directly to the sampling posterior. Because the aligned source is a realized motion rather than a regression output, the anchor injects sample-level temporal detail that a reconstruction-trained denoiser tends to smooth away. To learn effective source reuse, we supervise the anchored estimate against the editing target and match its second temporal differences through a temporal high-frequency loss. These objectives require no explicit edit masks. Extensive experiments show that SuperMotion improves editing accuracy, reaching 33.20% full-pool R@1 on MotionFix, while reducing temporal-detail attenuation and preserving motion dynamics as it realizes the requested changes. Ablations confirm that the learned preservation gate is responsible for the gain and that it reuses the source to retain the unedited content properly.
Fa-Ting Hong, Peter Wonka
King Abdullah University of Science and Technology
Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two major limitations: a representational gap caused by pre-trained text encoders that lack motion-specific information, and error propagation during the iterative denoising process. This paper introduces Reconstruction-Anchored Diffusion Model (RAM) to address these challenges. First, RAM leverages a motion latent space as intermediate supervision for text-to-motion generation. To this end, RAM co-trains a motion reconstruction branch with two key objective functions: self-regularization to enhance the discrimination of the motion space and motion-centric latent alignment to enable accurate mapping from text to the motion latent space. Second, we propose Reconstructive Error Guidance (REG), a testing-stage guidance mechanism that exploits the motion diffusion model's inherent self-correction ability to mitigate error propagation. At each denoising step, REG uses the motion reconstruction branch to reconstruct the previous estimate, reproducing the prior error patterns. By amplifying the residual between the current prediction and the reconstructed estimate, REG highlights the improvements in the current prediction. Extensive experiments demonstrate that RAM achieves significant improvements and state-of-the-art performance. Our code will be released.
Yifei Liu, Changxing Ding, Ling Guo +2
South China University of Technology · Joy Future Academy
We introduce MoLingo, a text-to-motion (T2M) model that generates realistic, lifelike human motion by denoising in a continuous latent space. Recent works perform latent space diffusion, either on the whole latent at once or auto-regressively over multiple latents. In this paper, we study how to make diffusion on continuous motion latents work best. We focus on two questions: (1) how to build a semantically aligned latent space so diffusion becomes more effective, and (2) how to best inject text conditioning so the motion follows the description closely. We propose a semantic-aligned motion encoder trained with frame-level text labels so that latents with similar text meaning stay close, which makes the latent space more diffusion-friendly. We also compare single-token conditioning with a multi-token cross-attention scheme and find that cross-attention gives better motion realism and text-motion alignment. With semantically aligned latents, auto-regressive generation, and cross-attention text conditioning, our model sets a new state of the art in human motion generation on standard metrics and in a user study. We will release our code and models for further research and downstream usage.
Yannan He, Garvita Tiwari, Xiaohan Zhang +4
University of Tübingen, Germany · Tübingen AI Center, Germany · Max Planck Institute for Informatics, Saarland Informatics Campus, Germany +2