Recent motion generative models have demonstrated strong capabilities in synthesizing physically plausible character motion, but often overlook established animation principles used by professional animators to ground and design their animation work. Understanding and incorporating these principles into motion generative pipelines is essential for producing motions that serve not only physically grounded applications but also the needs of the character animation community. This enables the creation of characters that not only move in physically plausible ways but also feel alive, expressive, and engaging. To close this gap, we focus on the Exaggeration principle of animation and investigate how it can be incorporated into modern motion generative pipelines to produce more expressive character motions. To this end, we introduce a framework that operates at two stages of existing motion generative pipelines. The first stage introduces exaggeration during training, where we perform supervised fine-tuning of pre-trained text-to-motion models on our curated exaggeration dataset. The second stage operates at inference time, where we: (i) introduce a mathematical formulation of exaggeration based on dynamic movement primitives (DMPs); and (ii) leverage this formulation as an exaggeration guidance signal to guide existing diffusion and flow-matching text-to-motion generation models toward exaggerated motion without additional training. Through qualitative and quantitative evaluations against three strong motion generation models, we show that our methods generate more exaggerated and expressive motions while preserving neutral reference motion intent and physical plausibility.
Figures & tables
Figure 1: Our approach generates controllable and plausible motion exaggerations ( {k=3,k=6,k=10} in green) from a neutral reference motion (dashed-line box in gray). Increasing k progressively amplifies the dance motion while preserving its underlying action intent.
Figure 2: An overview of the controllable motion exaggeration framework, consisting of two stages: (1) Training-Time Exaggeration , which constructs a graded motion-capture dataset with 26 actions and five exaggeration levels, augment and retarget the data, and perform supervised fine-tuning with LoRA and prior preservation to learn exaggeration as a controllable, text-conditioned attribute (Sec. 3.2 ); (2) Inference-Time Exaggeration , where a frozen pre-trained motion generation model first produces a neutral motion Qneut , which is transformed by a DMP-based exaggeration operator EK into an exaggerated target QK . Then, iterative guidance steers the sampling process toward QK , while the frozen generator regularizes the motion through its learned physical prior (Sec. 3.3 ). The bottom examples show progressively increasing levels of exaggeration.
Algorithm 1 Controllable Motion Exaggeration
Figure 3: Qualitative comparison of motion exaggeration. For each action, the top row shows the motion generated by the pre-trained model, while the bottom row shows the corresponding motion generated with our exaggeration method. Across MDM (green), KIMODO (orange), and HY-Motion (purple), our method (blue) consistently produces more expressive and exaggerated movements while preserving the underlying action.
Exaggeration
Action Preservation
Quality
Model
Method
ELRA@1↑
ρclf↑
ρphys↑
Δphys↑
Nov. ↑
R@3 ↑
FID ↓
Div ↑
MMDist ↓
Skate ↓
MDM
Baseline
0.200
0.000
0.026
0.175
0.117
0.7031
0.8388
9.5722
3.6133
0.070
+ DMP
0.200
0.000
0.358
0.953
0.130
0.5625
20.9116
6.8088
5.2530
0.134
+ TF
0.200
0.000
0.252
0.552
0.023
0.7070
5.7547
8.3935
3.9870
0.091
+ SFT
0.408
0.635
0.399
1.173
0.168
0.5547
1.6349
8.8127
4.6498
0.073
+ SFT+TF
0.408
0.643
0.545
1.610
0.276
0.5273
5.8575
7.7830
5.0234
0.085
Table 1: Quantitative evaluation of exaggeration and action preservation. We report results for HY-Motion (1B), MDM, and KIMODO under the baseline, DMP, TF, SFT, and combined SFT+TF settings. The Exaggeration metrics measure exaggeration control and physical intensity, while Action Preservation metrics evaluate text-motion alignment and preservation of the learned motion manifold. Quality reports foot-skate artifacts. ↑ and ↓ indicate higher and lower values are better, respectively. Bold values highlight the best result within each model and metric, while underlined values indicate the second-best result.
(a) Ablation on r
Backbone
Exaggeration
Action Preservation
KIMODO
86.7%
82.0%
HY-Motion
78.9%
50.0%
MDM
70.0%
–
Overall
81.8%
76.7%
Table 2: User-study results. Exaggeration reports how often participants selected the highest requested level as the most exaggerated, while Action Preservation reports the percentage of positive preservation ratings.
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
MDM
HY-Motion
KIMODO
(HumanML3D)
(SMPL-H)
(SMPL-X)
Feature dimension
263
201
273
Body joints
22
22
22
Frame rate
20 FPS
30 FPS
30 FPS
Prediction target
clean motion ( x0 )
velocity ( v )
clean motion ( x0 )
Text encoder
CLIP
CLIP-L + Qwen3-8B
LLM2Vec (Llama-3-8B)
Appendix
Table 3: Native motion representations and generative parameterizations used by the three backbones including MDM, HY-Motion, and KIMODO. Motion feature dimensions are reported per frame.
Setting
MDM
HY-Motion
KIMODO
Objective
masked x0 MSE
masked velocity MSE
masked x0 smooth- L1
Training steps
16 k
16 k
16 k
Batch size / GPU
64
16
8
Learning rate
1×10−4
1×10−4
1×10−4
Maximum frames
196
294
300
LoRA rank r
{64,128}
{16,64}
16
Appendix
Table 4: Default supervised fine-tuning hyperparameters and configurations.
Backbone
Exag. actions
Highest intensity
Neutral reference
Preserv. actions
Action preserved
KIMODO
3
86.7% (52/60)
80.0% (48/60)
5
82.0% (82/100)
HY-Motion
1
78.9% (15/19)
10.0% (2/20)
1
50.0% (10/20)
MDM
1
70.0% (14/20)
25.0% (5/20)
0
–
Overall
5
81.8% (81/99)
55.0% (55/100)
6
76.7% (92/120)
Appendix
Table 5: User-study results for controllable exaggeration and action preservation. We report how often participants identified the highest requested intensity as the most exaggerated and the neutral reference motion as the most neutral. Action preservation is the percentage of ratings marked very well preserved or mostly preserved . Percentages are computed over valid responses, with counts shown in parentheses; “–” indicates that no example was evaluated.
Exaggeration
Action Preservation
Quality
Setting
ELRA@1↑
ρclf↑
ρphys↑
Δphys↑
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MMDist ↓
Skate ↓
Reference rows
ExMoCap
0.983
0.995
0.564
1.494
–
–
–
–
–
–
0.050
HumanML3D
–
–
–
–
0.4530
0.6597
0.7675
0.0010
9.2461
3.2417
0.044
Base MDM (no FT)
0.199
0.014
0.099
0.273
0.4019
0.6062
0.7211
0.4444
9.5169
3.5280
0.076
(a) Fine-tuning strategy
Appendix
Table 6: Ablation study of training-time (SFT) exaggeration approach on MDM. We evaluate the effect of (a) fine-tuning strategy, (b) LoRA rank, (c) prior-preservation weight, and (d) prior-preservation data source across metrics measuring exaggeration control, prompt adherence, motion-manifold fidelity, and motion quality. The goal is to identify effective hyperparameters and configurations for supervised fine-tuning of the MDM model. Section (a) compares two fine-tuning strategies: (i) LoRA (ours) , which trains LoRA adapters with the proposed prior-preservation objective, and (ii) Full FT , which updates all model parameters using the same objective. Full fine-tuning produces stronger exaggeration, but substantially degrades prompt adherence and motion-manifold fidelity. In contrast, LoRA with prior preservation provides a better trade-off between exaggeration control and preservation of the base model’s action semantics and motion manifold. In Sections (b) and (c), we ablate this LoRA+prior-preservation configuration by training identical models with different LoRA ranks r and prior-preservation weights λ , respectively. Increasing the LoRA rank consistently improves exaggeration control but progressively reduces prompt adherence and manifold fidelity. Increasing the prior-preservation weight exhibits the opposite trend, improving prompt adherence and manifold fidelity while reducing the strength of exaggeration. Therefore, we select r∈{64,128} and λ∈{0.1,0.25} , providing a practical trade-off between exaggeration and preservation. Finally, Section (d) compares different sources for the prior-preservation data: real HumanML3D motions, motions generated by the pre-trained MDM, and a mixture of the two. Using real HumanML3D motions generally provides stronger preservation of prompt adherence and motion-manifold fidelity, while generated or mixed priors provide comparable but slightly weaker results.
Exaggeration
Action Preservation
Quality
Setting
ELRA@1↑
ρclf↑
ρphys↑
Δphys↑
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MMDist ↓
Skate ↓
Reference rows
ExMoCap
0.995
0.999
0.542
1.534
–
–
–
–
–
–
0.071
HumanML3D
–
–
–
–
0.4516
0.6596
0.7685
0.0010
9.0560
3.2341
–
Base HY-Motion (no FT)
0.335
0.453
0.335
0.970
0.3004
0.4617
0.5827
3.5259
8.1049
4.5667
0.046
(a) Fine-tuning strategy
Appendix
Table 7: Ablation study of Training-time (SFT) exaggeration approach on HY-Motion (0.4B). We evaluate the effect of (a) fine-tuning strategy, (b) LoRA rank, and (c) prior-preservation weight across metrics measuring exaggeration control, prompt adherence, motion-manifold fidelity, and motion quality. The goal is to identify effective hyperparameters and configurations for supervised fine-tuning of the HY-Motion (0.4B) model. Section (a) compares two fine-tuning strategies: (i) LoRA (ours) , which trains LoRA adapters with the proposed prior-preservation objective, and (ii) Full FT , which updates all model parameters without regularization. Full fine-tuning produces substantially stronger exaggeration but severely degrades prompt adherence and motion-manifold fidelity. In contrast, LoRA with prior preservation significantly improves exaggeration control while retaining the pretrained model’s action semantics and motion manifold. In Section (b), we ablate the LoRA rank r while keeping the prior-preservation weight fixed. Increasing r improves exaggeration control from r=1 up to approximately r=16 , after which the gains largely saturate. Meanwhile, prompt-adherence and manifold metrics remain relatively stable across ranks, with no consistent degradation as observed for MDM. We therefore select a relatively small rank, r=16 , for HY-Motion (0.4B), which provides strong exaggeration control. In Section (c), we ablate the prior-preservation weight λ . Removing prior preservation ( λ=0 ) yields the strongest classifier-based exaggeration, but substantially worsens prompt adherence and manifold fidelity. Increasing λ generally improves prompt adherence, with the strongest R-precision obtained around λ∈{0.25,0.5} , while larger values eventually weaken exaggeration control. We therefore use λ=0.25 for the selected configuration, providing a strong balance between exaggeration and preservation of the pretrained motion prior.
Exaggeration
Action Preservation
Quality
Setting
ELRA@1↑
ρclf↑
ρphys↑
Δphys↑
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MMDist ↓
Skate ↓
Reference rows
ExMoCap
0.995
0.999
0.542
1.534
–
–
–
–
–
–
0.071
HumanML3D
–
–
–
–
0.4527
0.6619
0.7719
0.0010
9.4287
3.2242
–
Base HY-Motion 1B (no FT)
0.291
0.346
0.300
0.720
0.3266
0.5262
0.6351
2.9705
8.2249
4.3312
0.044
(a) Fine-tuning strategy
Appendix
Table 8: Ablation study of training-time (SFT) exaggeration approach on HY-Motion (1B). We evaluate the effect of (a) fine-tuning strategy, (b) LoRA rank, and (c) prior-preservation weight across metrics measuring exaggeration control, prompt adherence, motion-manifold fidelity, and motion quality. The goal is to identify effective hyperparameters and configurations for supervised fine-tuning of the HY-Motion (1B) model. Section (a) compares two fine-tuning strategies: (i) LoRA (ours) , which trains LoRA adapters with the proposed prior-preservation objective, and (ii) Full FT , which updates all model parameters without regularization. Full fine-tuning produces stronger exaggeration, but significantly degrades prompt adherence and motion-manifold fidelity. In contrast, LoRA with prior preservation maintains much stronger prompt adherence and manifold fidelity while still providing substantial exaggeration control. In Section (b), we ablate the LoRA rank r with a fixed prior-preservation weight. Increasing r generally improves classifier-based exaggeration control up to r=128 . Prompt adherence and manifold fidelity also remain relatively stable across the tested ranks, with r=256 providing the strongest R-precision and MMDist among the rank ablations. We therefore select r=64 for the HY-Motion (1B) configuration, providing strong exaggeration control while retaining competitive prompt adherence and manifold fidelity without the additional capacity of larger adapters. In Section (c), we ablate the prior-preservation weight λ . Removing prior preservation ( λ=0 ) gives the strongest classifier-based exaggeration, but significantly reduces prompt adherence and increases FID and MMDist. Introducing a small prior-preservation weight improves prompt adherence and manifold fidelity, with λ=0.25 providing particularly strong R-precision together with a high physical-intensity gap, while larger weights progressively weaken classifier-based exaggeration. We therefore use λ=0.25 for the selected configuration, providing a practical balance between exaggeration control and preservation of the pretrained motion prior.
Exaggeration
Action Preservation
Quality
Setting
ELRA@1↑
ρclf↑
ρphys↑
Δphys↑
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div →
MMDist ↓
Skate ↓
Reference rows
ExMoCap
0.990
0.998
0.540
1.527
–
–
–
–
–
–
0.071
HumanML3D
–
–
–
–
0.4537
0.6611
0.7748
0.0010
9.4376
3.2301
–
Base KIMODO (no FT)
0.203
-0.024
0.174
0.448
0.3291
0.5373
0.6442
1.7365
8.9833
4.0029
0.026
(a) LoRA rank
Appendix
Table 9: Ablation study of supervised fine-tuning (SFT) exaggeration approach on KIMODO. We evaluate the effect of (a) LoRA rank and (b) prior-preservation weight across metrics measuring exaggeration control, prompt adherence, motion-manifold fidelity, and motion quality. The goal is to identify effective hyperparameters and configurations for supervised fine-tuning of the KIMODO model. In Section (a), we ablate the LoRA rank r with a fixed prior-preservation weight. Unlike MDM and HY-Motion, in KIMODO, the effect of LoRA rank on exaggeration control is not monotonic: ELRA@1 and ρclf improve from r=1 to moderate ranks but fluctuate thereafter, while the strongest physical-intensity correlation and neutral-to-extreme physical gap are obtained at the largest tested rank, r=512 . Prompt adherence and manifold fidelity also vary non-monotonically across ranks, with different ranks providing the strongest individual metrics; for example, r=256 gives the highest R@1, while r=32 gives the lowest FID. These results suggest that increasing LoRA capacity does not uniformly improve either exaggeration control or prompt adherence for KIMODO. In Section (b), we ablate the prior-preservation weight λ . Removing prior preservation ( λ=0 ) gives the strongest classifier-based exaggeration metrics, but substantially degrades prompt adherence and manifold fidelity. Introducing a small prior-preservation weight improves the text-motion alignment metrics, with λ∈{0.1,0.25} providing strong R-precision and MMDist, while larger values generally reduce classifier-based exaggeration. The physical-intensity metrics remain relatively stable across moderate values of λ , with λ=0.5 achieving the largest Δphys . Based on these results, we select a moderate LoRA rank and a small prior-preservation weight to balance exaggeration control with preservation of the pretrained motion prior.
Exaggeration control
Prompt adherence & manifold
Quality
Model
Method
ELRA@1↑
ρclf↑
ρphys↑
Δphys↑
Nov. ↑
R@1 ↑
R@2 ↑
R@3 ↑
FID ↓
Div ↑
MMDist ↓
Skate ↓
HY-Motion-0.4B
ExMoCap
0.995
0.999
0.542
1.534
–
–
–
–
–
–
–
–
HumanML3D
–
–
–
–
–
0.4481
0.6554
0.7640
0.0011
9.3215
3.2445
–
Baseline
0.331
0.267
0.315
0.961
0.069
0.3945
0.6016
0.7305
2.5235
8.3203
3.5885
0.051
+ DMP
0.231
0.055
0.471
1.257
0.185
0.1875
0.2969
0.3828
34.5822
4.6140
6.2699
0.080
+ TF
0.223
0.039
0.366
0.947
0.107
0.2383
0.3867
0.5273
22.0843
5.9314
5.1676
0.072
Appendix
Table 10: Contribution of the proposed motion-exaggeration approaches. We compare the original HY-Motion-0.4B, HY-Motion-1B, MDM, and KIMODO baselines with DMP, training-free guidance (TF), supervised fine-tuning (SFT), and their composition (SFT+TF). Each approach is evaluated using its corresponding exaggeration control: text-conditioned intensity levels for the Baseline and SFT, increasing DMP scale factors for DMP and TF, and both controls jointly for SFT+TF. The metrics evaluate exaggeration control, prompt adherence and motion-manifold fidelity, and motion quality, revealing the trade-off between stronger exaggeration and preservation of the underlying action. “ExMoCap” provides a reference for the exaggeration-control metrics, while “HumanML3D” provides a reference for the standard text-to-motion metrics; neither reference row is included when identifying the best method. ↑ and ↓ indicate that higher and lower values are better, respectively. The best result among the generated-motion methods for each backbone and metric is shown in bold.
Figure 5: Ablation study of LoRA rank and prior-preservation weight for MDM. We evaluate the trade-off between exaggeration control and prompt adherence by varying the LoRA rank r (left) and prior-preservation weight λ (right) for MDM ( Tevet et al., 2023 ) . Higher ELRA@1 and Δphys indicate stronger exaggeration control, while higher R-precision indicates better prompt adherence. The pre-trained MDM baseline is denoted by a star.
Figure 6: Qualitative evaluation of our SFT+TF approach compared to KIMODO. Selected frames from Kick and Plie-bend actions show that SFT+TF amplifies the characteristic poses and motion amplitudes of the original actions generated by KIMODO, producing more expressive exaggeration while preserving the underlying neutral action and its temporal progression.
Figure 7: Qualitative evaluation of our SFT approach compared to KIMODO. Selected frames from Reaching and Walk actions show that SFT amplifies the characteristic poses and motion amplitudes of the original actions generated by KIMODO, producing more expressive exaggeration while preserving the underlying neutral action and its temporal progression.
Figure 8: Qualitative evaluation of our TF approach compared to KIMODO. Selected frames from Dance and Throw actions show that TF amplifies the characteristic poses and motion amplitudes of the original actions generated by KIMODO, producing more expressive exaggeration while preserving the underlying neutral action and its temporal progression.
Figure 9: More qualitative results on MDM and HY-Motion. Selected frames from Slide and Jump actions compare the baseline models with our approach. Across both motion generation models, our method amplifies characteristic poses and motion amplitudes while preserving the underlying neutral action and temporal progression.
Figure 10: Controllable exaggeration using our proposed supervised fine-tuning (SFT) approach. The same kick motion is generated at progressively increasing exaggeration levels from left to right, with time progressing from top to bottom. Increasing exaggeration intensity through text prompt in the SFT approach produces progressively larger and more expressive poses while preserving the underlying neutral action.
Figure 11: Controllable exaggeration using our proposed training-free (TF) approach. The same dance motion is generated at progressively increasing exaggeration levels from left to right, with time progressing from top to bottom. Increasing exaggeration intensity through the parameter k in the TF approach produces progressively larger and more expressive poses while preserving the underlying neutral action.
Figure 12: Representative questions from our user study. Top: for the controllable-exaggeration evaluation, participants viewed anonymized motions at different exaggeration levels and selected the most exaggerated and most neutral motions. Bottom: for the action-preservation evaluation, participants compared a neutral source motion with its exaggerated counterpart and rated how well the counterpart preserved the source action using a four-level response scale.
Diffusion models have seen widespread adoption for text-driven human motion generation and related tasks due to their impressive generative capabilities and flexibility. However, current motion diffusion models face two major limitations: a representational gap caused by pre-trained text encoders that lack motion-specific information, and error propagation during the iterative denoising process. This paper introduces Reconstruction-Anchored Diffusion Model (RAM) to address these challenges. First, RAM leverages a motion latent space as intermediate supervision for text-to-motion generation. To this end, RAM co-trains a motion reconstruction branch with two key objective functions: self-regularization to enhance the discrimination of the motion space and motion-centric latent alignment to enable accurate mapping from text to the motion latent space. Second, we propose Reconstructive Error Guidance (REG), a testing-stage guidance mechanism that exploits the motion diffusion model's inherent self-correction ability to mitigate error propagation. At each denoising step, REG uses the motion reconstruction branch to reconstruct the previous estimate, reproducing the prior error patterns. By amplifying the residual between the current prediction and the reconstructed estimate, REG highlights the improvements in the current prediction. Extensive experiments demonstrate that RAM achieves significant improvements and state-of-the-art performance. Our code will be released.
Yifei Liu, Changxing Ding, Ling Guo +2
South China University of Technology · Joy Future Academy
Human motion generation models are fundamentally constrained by the limited diversity of motion capture datasets, which predominantly contain common, repetitive actions and fail to cover the long tail of complex human movements, resulting in a restricted motion vocabulary in learned latent representations and poor generalization to rare, compositional, and highly dynamic motions. In this work, we propose a framework for expanding the motion representation space by leveraging large-scale synthetic human motion, introducing a data generation pipeline that produces diverse, physically plausible motion sequences beyond the distribution of existing datasets and integrating it with a redesigned VQ-VAE tokenizer that adapts to this expanded motion space. Unlike conventional tokenizers trained on narrow data distributions, our approach jointly scales both the training distribution and the discrete codebook, enabling the model to capture a significantly richer set of motion primitives. We demonstrate that training with synthetic motion substantially improves the coverage and compositionality of the learned motion vocabulary, leading to consistent gains across motion generation tasks such as text-to-motion and motion continuation, while remaining fully compatible with existing frameworks including MotionGPT. Our results suggest that the primary bottleneck lies in the limited support of the learned motion representation, rather than model architecture alone. Scaling synthetic motion in tandem with representation learning offers a principled path toward more expressive, controllable, and generalizable human motion synthesis.
Recent advances in motion-aware large language models have shown remarkable promise for jointly learning motion understanding and generation knowledge. However, these models typically treat understanding and generation separately, limiting the mutual benefits that could arise from interactive feedback between tasks. In this work, we reveal that motion assessment and refinement tasks can act as crucial bridges to enable knowledge flow from motion understanding to generation. Specifically, we propose Interleaved Reasoning for Motion Generation (IRMoGen), a novel paradigm that tightly couples motion generation with assessment and refinement through iterative text-motion dialogue. To realize this, we introduce IRG-MotionLLM, the first model that seamlessly interleaves motion generation, assessment, and refinement to improve the alignment between generated motion and goal text. IRG-MotionLLM is developed progressively with a novel three-stage training scheme, initializing and subsequently enhancing native IRMoGen capabilities. To facilitate this development, we construct an automated data engine to synthesize interleaved reasoning annotations from existing text-motion datasets. Extensive experiments demonstrate the properties brought by IRMoGen training, and the advanced cross-benchmark and cross-evaluator performance of IRG-MotionLLM. Code and models are available at https://github.com/HumanMLLM/IRG-MotionLLM.
Yuan-Ming Li, Qize Yang, Nan Lei +5
Sun Yat-sen University · Tongyi Lab, Alibaba Group · Key Laboratory of Machine Intelligence and Advanced Computing, Ministry of Education, China +1