Despite recent advances in video generation, motion customization remains challenging due to content leakage, where appearance attributes from the reference video unintentionally propagate into the generated output. We identify this issue as a consequence of the generative process collapsing toward the reference video, which arises from formulating the learning objective as a direct regression on the reference. To address this, we propose Control-based Motion Customization (CMC), a principled training framework that is structurally robust to content leakage. Our key idea is to steer generative dynamics toward desired motion while avoiding collapse toward the reference video, which we formalize using Stochastic Optimal Control (SOC). Under this formulation, customized videos acquire the target motion yet remain within the pre-trained model's prompt-conditional distribution, where appearance is determined by the text prompt rather than the reference video. Furthermore, to improve efficiency, we tailor the SOC formulation to motion customization by eliminating the need for an explicit reward and introducing a timestep-adaptive motion cost that focuses only on early generative stages, accelerating training by 2.5 times. Extensive experiments demonstrate that CMC effectively mitigates content leakage and achieves competitive motion fidelity while preserving the diversity of the base model across diverse scenarios.
Figures & tables
Figure 1: Overview of the proposed method. Our CMC mitigates content leakage, enabling diverse generation across varying prompts and samples.
Figure 3: Qualitative analysis on motion formation timestep. The y -axis indicates the switching timestep ( t∈[0,50] ) at which the initial prompt Pinit is switched to a new prompt Pmot or Papp . Left : Switching from Pinit to Pmot shows that motion becomes largely established after early timesteps ( t=10 ), with limited influence from subsequent prompts. Right : Switching to Papp shows that non-motion can still change at later timestep ( t=20 ).
Figure 4: Quantitative analysis on motion formation timestep.
Figure 5: Qualitative comparison with baselines. CMC achieves precise motion customization while significantly mitigating content leakage. The red arrow indicates the movement direction.
Method
Motion Fidelity ↓
Text Fidelity ↑
Temporal Consistency ↑
Diversity ↑
Backbone
0.2605
0.8520
0.9797
0.3139
Training-free methods
SMM ( Yatim et al., 2024 )
0.1123
0.8229
0.9650
0.2457
MOFT ( Xiao et al., 2024 )
0.1186
0.8272
0.9563
0.2292
MotionClone (MC) ( Ling et al., 2025 )
0.1303
0.7698
0.9253
0.2489
DiTFlow ( Pondaven et al., 2025 )
0.1189
0.8304
0.9814
0.2402
Table 1: Quantitative evaluation on CMC and baselines.
Figure 6: Ablation study. ts denotes the motion injection timestep ( ts=0 corresponds to the base T2V). Motion aligns with the reference from ts=10 . Removing w(t) leads to degenerated outputs.
Table S2: Motion formation timestep across model scales.
Running cost
Motion Fidelity ↓
Diversity ↑
w/o CMC
w/ CMC (Ours)
w/o CMC
w/ CMC (Ours)
Simple feature-matching (training-free)
0.145
0.140
0.197
0.206
Simple feature-matching (training-based)
0.134
-
0.171
-
MotionClone-style cost
0.188
0.212
0.254
0.282
Appendix
Table S3: Effect of CMC on top of existing methods.
Figure S1: Additional qualitative comparison with baselines. The red arrow indicates the camera movement direction.
Figure S2: Comparison with tuning-based methods. Tuning-based methods suffer from content leakage, producing near-identical samples that resemble the reference video, whereas CMC generates diverse samples. The red arrow indicates the movement direction.
Training-free motion customization imposes motion patterns from reference videos onto video generators through test-time computation. Most existing methods target full diffusion models, requiring many denoising steps and high computational cost. With the rise of efficient distilled models, a natural question arises: can test-time motion customization be applied directly to distilled generators with their accelerated sampling and efficiency gains? However, our analysis reveals that existing training-free techniques fail on distilled models. Distillation fundamentally alters the denoising dynamics that prior test-time guidance relies on, and the large denoising steps of distilled generators discard the dense intermediate states that score guidance requires, rendering existing motion control strategies incompatible with fast generation. To address this limitation, we propose MotionEcho, a novel training-free test-time distillation framework that enables motion customization for distilled video generators. The key idea is to correct the student model's sampling trajectory with restricted usage of a high-quality diffusion teacher at inference time. Teacher supervises the student's denoising by re-noising the student's endpoint onto its dense trajectory to form a motion-aligned clean endpoint, then interpolating it with the student's, while an adaptive scheduling mechanism determines when and how much teacher guidance is needed. As a result, MotionEcho restores generative trajectories for distilled video generators via lightweight, adaptive test-time teacher guidance, enabling accurate motion control without compromising generation efficiency. Extensive experiments on multiple distilled video generation models demonstrate that our method significantly improves motion fidelity and visual quality while retaining the efficiency advantages of distilled generation.
Jintao Rong, Xin Xie, Xinyi Yu +4
Zhejiang University of Technology · University of New South Wales (UNSW Sydney) · University of Auckland +1
Autoregressive video models can generate minute-long videos in real time, but they produce generic subjects from text rather than specific subjects from user-provided images. Existing customization methods either require costly per-subject optimization or use pretrained conditioning networks that jointly process all video frames with bidirectional attention. Neither approach is designed for causal streaming. We present Custom Forcing, a training-free method that stores reference-based anchor frames in the persistent KV cache of a frozen autoregressive video model. However, fixed anchors face two limitations: simple conditioning allows identity to drift, and the text prompt continues to favor a generic subject. To address these problems, drift-adaptive value amplification (DVA) scales reference influence with the degree of identity drift, while anchor contrast guidance (ACG) steers generation away from the generic class prior. Over two-minute rollouts, fixed anchors fall from 0.58 to 0.42 in DINO-I, while Custom Forcing keeps it between 0.58 and 0.62 without reducing motion. Custom Forcing also achieves higher subject similarity than bidirectional customization methods and better preserves identity over 30s than causal image-to-video and reference-to-video models, while generating each frame 9.5-28.5 times faster than these long-video baselines.
Yunseung Ok, Hyunsoo Kim, Minseo Kim +1
Kyung Hee University · The University of Texas at Austin
Video customization based on Text-to-Video (T2V) models aims to learn specific features from reference data to generate controllable videos. While significant strides have been made in image stylization and video motion customization, simultaneously controlling multiple concepts, such as content, style, and motion, remains a major challenge. In this work, we systematically define the task of multi-concept video customization, which requires the joint control of content, style, and motion. To facilitate research in this area, we construct a comprehensive benchmark and propose Disco-LoRA, a unified framework designed to tackle this problem by disentangling and flexibly recombining different concepts in two stages: (1) We decompose the objective into two sub-tasks: Content-Style and Content-Motion. Each sub-task is addressed using our Iterative Dual-LoRA Disentanglement Framework, which effectively disentangles distinct concepts within the data. (2) We identify layer-wise weight trends as crucial for LoRA identity, while weight magnitudes dictate composability. To harmonize these scales, we propose a Z-score-based statistical regularization that aligns weight distributions, preserving layer-wise trends while minimizing interference between different LoRAs. Extensive experiments show that Disco-LoRA excels in multi-concept video customization, effectively preserving appearance, style, and motion for controllable text-to-video generation.
Xuancheng Xu, Gengyun Jia, Bing-Kun Bao
Nanjing University of Posts and Telecommunications · Hefei University of Technology · Peng Cheng Laboratory