Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video, which arises from formulating the learning objective as a direct regression on the reference. To address this, we propose Control-based Motion Customization (CMC), a principled training framework that is structurally robust to content leakage. Our key idea is to steer generative dynamics toward desired motion while avoiding collapse toward the reference video, which we formalize using Stochastic Optimal Control (SOC). Under this formulation, customized videos acquire the target motion yet remain within the pre-trained model's prompt-conditional distribution, where appearance is determined by the text prompt rather than the reference video. Furthermore, to improve efficiency, we tailor the SOC formulation to motion customization by eliminating the need for an explicit reward and introducing a timestep-adaptive motion cost that focuses only on early generative stages, accelerating training by 2.5 times. Extensive experiments demonstrate that CMC effectively mitigates content leakage and achieves competitive motion fidelity while preserving the diversity of the base model across diverse scenarios.
Figures & tables
Figure 1: Overview of the proposed method. Our CMC mitigates content leakage, enabling diverse generation across varying prompts and samples.
Figure 3: Qualitative analysis on motion formation timestep. The y -axis indicates the switching timestep ( t∈[0,50] ) at which the initial prompt Pinit is switched to a new prompt Pmot or Papp . Left : Switching from Pinit to Pmot shows that motion becomes largely established after early timesteps ( t=10 ), with limited influence from subsequent prompts. Right : Switching to Papp shows that non-motion can still change at later timestep ( t=20 ).
Figure 4: Quantitative analysis on motion formation timestep.
Figure 5: Qualitative comparison with baselines. CMC achieves precise motion customization while significantly mitigating content leakage. The red arrow indicates the movement direction.
Method
Motion Fidelity ↓
Text Fidelity ↑
Temporal Consistency ↑
Diversity ↑
Backbone
0.2605
0.8520
0.9797
0.3139
Training-free methods
SMM ( Yatim et al., 2024 )
0.1123
0.8229
0.9650
0.2457
MOFT ( Xiao et al., 2024 )
0.1186
0.8272
0.9563
0.2292
MotionClone (MC) ( Ling et al., 2025 )
0.1303
0.7698
0.9253
0.2489
DiTFlow ( Pondaven et al., 2025 )
0.1189
0.8304
0.9814
0.2402
Table 1: Quantitative evaluation on CMC and baselines.
Figure 6: Ablation study. ts denotes the motion injection timestep ( ts=0 corresponds to the base T2V). Motion aligns with the reference from ts=10 . Removing w(t) leads to degenerated outputs.
Table S2: Motion formation timestep across model scales.
Running cost
Motion Fidelity ↓
Diversity ↑
w/o CMC
w/ CMC (Ours)
w/o CMC
w/ CMC (Ours)
Simple feature-matching (training-free)
0.145
0.140
0.197
0.206
Simple feature-matching (training-based)
0.134
-
0.171
-
MotionClone-style cost
0.188
0.212
0.254
0.282
Appendix
Table S3: Effect of CMC on top of existing methods.
Figure S1: Additional qualitative comparison with baselines. The red arrow indicates the camera movement direction.
Figure S2: Comparison with tuning-based methods. Tuning-based methods suffer from content leakage, producing near-identical samples that resemble the reference video, whereas CMC generates diverse samples. The red arrow indicates the movement direction.