cs.CVSep 29, 2026

BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling

Authors: Xiaojian Shen, Dahu Shi, Jianrong Zhang, Hai Li, Hongwei Zhao, Dawei Zhang, Yunzhi Zhuge, Zhiliang Wu, +2 more

Organizations: College of Software, Jilin University, Changchun, 130012, China · College of Computer Science and Technology, Zhejiang University, Hangzhou, 310007, China · ReLER, AAII, University of Technology Sydney, Sydney, Australia · College of Computer Science and Technology, Jilin University, Changchun, 130012, China · School of Computer Science and Technology, Zhejiang Normal University, Jinhua, 321004, China · School of Information and Communication Engineering, Dalian University of Technology, Dalian, 116024, China · School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, 518060, China · School of Computer Science and Informatics, Cardiff University, CF10 3AT Cardiff, U.K.

Abstract

Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.

Figures & tables

Explore similar work

CardsList
  1. FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation

    Nov 26, 2025Kaixing Yang, Xulong Tang, Ziqiao Peng +4Human Motion GenerationDuet

  2. Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment

    Jul 12, 2026Ryota Kimura, Sangheon Park, Natalia Polouliakh +1DuetDance-To-Music Generation

  3. Where the Body Keeps the Beat: Structured Motion Conditioning and Music Dynamics Supervision for Dance-to-Music Generation

    Sep 30, 2026Changchang Sun, Lu Cheng, Yan YanDuetDance-To-Music Generation