cs.CVJul 13, 2026

Controlling Motion Transfer in Diffusion Transformers via Attention Heads

Authors: Sunyoung Jung, Jiwoo Park, Yoonseok Choi, Kyobin Choo, Ming-Hsuan Yang, Seong Jae Hwang

Organizations: Yonsei University · LG Electronics · University of California, Merced

Abstract

Diffusion Transformers (DiTs) have advanced video generation with high-quality, temporally coherent results. However, extending them to motion transfer, which requires following reference motion while aligning with a target prompt, remains challenging due to limited understanding of motion and structure representations within DiTs. We analyze video DiTs at the attention-head level and identify distinct heads specialized for motion and spatial structure. Based on this insight, we propose a head-aware controllable motion transfer framework that requires no parameter updates. Our method refines motion cues from motion-specialized heads via semantic correspondence guidance and preserves structure through selective feature injection. This head-level control not only enables accurate motion transfer but also provides an interpretable foundation for controllable video generation with DiTs.

Explore similar work

CardsList
  1. Making Time Editable in Video Diffusion Transformers

    Jun 8, 2026Konstantin Kuklev, Viacheslav Vasilev, Alexander Kunitsyn +2Pre-Trained Video Diffusion ModelsTemporal Dynamics