BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling
Authors: Xiaojian Shen, Dahu Shi, Jianrong Zhang, Hai Li, Hongwei Zhao, Dawei Zhang, Yunzhi Zhuge, Zhiliang Wu, +2 more
Organizations: College of Software, Jilin University, Changchun, 130012, China · College of Computer Science and Technology, Zhejiang University, Hangzhou, 310007, China · ReLER, AAII, University of Technology Sydney, Sydney, Australia · College of Computer Science and Technology, Jilin University, Changchun, 130012, China · School of Computer Science and Technology, Zhejiang Normal University, Jinhua, 321004, China · School of Information and Communication Engineering, Dalian University of Technology, Dalian, 116024, China · School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, 518060, China · School of Computer Science and Informatics, Cardiff University, CF10 3AT Cardiff, U.K.
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.
Figures & tables
Figure 1: Comparison with the existing state-of-the-art method. (a) POPDG [ 24 ] uses a cross-attention and an alignment module to associate music and dance. However, it may face challenges in achieving beat-level synchronization. (b) We propose Hierarchical Decoupled Attention (HDA) and a cycle-consistent learning strategy. HDA hierarchically decouples spatial and temporal information, while cycle-consistent learning leverages the inverse dance-to-music mapping to further enhance temporal alignment.
Figure 2: Overview of BeatDance. Hierarchical Decoupled Attention (HDA) employs a hierarchical design to capture both short-term and long-term music-dance dependencies, building upon a Decoupled Attention (DA) that splits the modeling into spatial and temporal branches. Additionally, our cycle-consistent learning mechanism uses an auxiliary Dance-to-Music (D2M) model to reconstruct the source music features from the generated dance, thereby enforcing a tighter music-dance alignment.
Dataset
Method
Motion Quality
Motion Diversity
Alignment
PFC ↓
PBC →
Div k →
Div g →
BAS ↑
PopDanceSet
Ground Truth
1.5824
8.7365
9.0219
7.2931
0.174
Bailando [ 31 ]
3.9751
4.8863
5.1835
5.4342
0.230
EDGE † [ 35 ]
3.8366
4.0348
6.1709
5.7568
0.224
POPDG † [ 24 ]
1.8253
5.9492
7.1342
5.8314
0.233
DanceEditor † [ 50 ]
1.8972
7.8429
5.6338
5.2459
0.238
Table 1: Comparison with the state-of-the-art methods on PopDanceSet and AIST++ test sets. Metrics are marked as: ↑ higher is better, ↓ lower is better, → closer to GT is better. Bold ( underlined ) indicates the best (second-best) results. † denotes diffusion-based methods.
Figure 3: Visual results on the PopDanceSet. We visually compare BeatDance with the state-of-the-art method POPDG [ 24 ] . Penetration artifacts (red) and Distorted motions (blue) are highlighted.
Figure 4: Visual comparison of beat consistency.
Comparison
BeatDance Win Rate
BeatDance vs. POPDG
91%
BeatDance vs. DanceEditor
94%
Table 2: Pairwise user-study results on the PopDanceSet test set.
Method
PFC
PBC
Div k
Div g
BAS
w/o DA
15.5267
6.3755
8.9975
6.1423
0.218
w/o HM
1.6046
5.7520
6.9480
6.3608
0.240
w/o CCL
1.5594
5.8300
6.6467
5.5107
0.252
Ours
1.5390
7.7493
6.7770
6.1801
0.273
Table 3: Ablation study of different components. DA, HM, and CCL denote the Decoupled Attention, Hierarchical Mechanism, and Cycle Consistency Learning, respectively.
Order
PFC
PBC
Div k
Div g
BAS
1
1.2401
7.2673
6.6585
5.8027
0.229
2
1.3925
6.9268
6.6322
5.4058
0.242
3
1.5390
7.7493
6.7770
6.1801
0.273
4
1.7447
7.3459
6.5358
5.6094
0.255
5
1.1339
6.4334
6.2685
5.1131
0.256
Table 4: Ablation study of the order of the Taylor series expansion.
τ
PFC
PBC
Div k
Div g
BAS
5
2.6125
8.5605
7.5657
5.8857
0.232
15
1.6489
6.9349
6.6073
5.5570
0.249
30
1.5390
7.7493
6.7770
6.1801
0.273
50
1.4113
6.7600
8.7246
6.0828
0.218
75
1.8742
5.3648
7.2374
5.5517
0.234
Table 5: Ablation study of the local window size τ .
Method
BCS ↑
BHS ↑
FAD ↓
BAS ↑
Ground Truth
100
100
0
0.174
GAN
86.5
52.8
7.63
0.191
VAE+GAN
100.2
54.6
7.52
0.193
Diffusion
110.8
62.1
7.60
0.193
VAE+Diffusion
113.8
64.8
6.86
0.195
Table 6: Ablation study of different generative architectures for the dance-to-music generation task.
Music-to-dance generation aims to translate auditory signals into expressive human motion, with broad applications in virtual reality, choreography, and digital entertainment. Despite promising progress, the limited generation efficiency of existing methods leaves insufficient computational headroom for high-fidelity 3D rendering, thereby constraining the expressiveness of 3D characters during real-world applications. Thus, we propose FlowerDance, which not only generates refined motion with physical plausibility and artistic expressiveness, but also achieves significant generation efficiency on inference speed and memory utilization. Specifically, FlowerDance combines MeanFlow with Physical Consistency Constraints, which enables high-quality motion generation with only a few sampling steps. Moreover, FlowerDance leverages a simple but efficient model architecture with BiMamba-based backbone and Channel-Level Cross-Modal Fusion, which generates dance with efficient non-autoregressive manner. Meanwhile, FlowerDance supports motion editing, enabling users to interactively refine dance sequences. Extensive experiments on AIST++ and FineDance show that FlowerDance achieves state-of-the-art results in both motion quality and generation efficiency. Code will be released upon acceptance.
Kaixing Yang, Xulong Tang, Ziqiao Peng +4
Renmin University of China · Malou Tech Inc · Wuhan University +1
Dance-to-music generation is a promising task for applications such as choreography support and automatic accompaniment, where temporal coordination between body movement and sound is essential. In particular, using human joint positions as the motion representation is attractive because they explicitly capture body dynamics while being lightweight, privacy-preserving, and easy to integrate with motion capture and pose-estimation pipelines. A central challenge in this setting, however, is the scarcity of high-quality paired dance-music data, since collecting accurately synchronized pairs is costly and often constrained by copyright and performance rights. This makes it difficult to train end-to-end models solely from paired data. To address this issue, we propose a dance-conditioned music generation framework that efficiently exploits both unpaired and paired data. Our method combines pretrained unimodal encoders for motion and music, beat-guided contrastive pretraining to align their feature spaces, and a ControlNet-style conditioning module on top of a pretrained text-to-audio diffusion model. Experiments on AIST++ demonstrate that the proposed techniques improve both dance-music alignment and audio quality, as confirmed by quantitative and qualitative evaluations. Compared to a state-of-the-art method, our approach achieves superior dance alignment performance and competitive audio quality. Code is available at https://github.com/kmraven/AudioLDM-ControlNet .
Ryota Kimura, Sangheon Park, Natalia Polouliakh +1
Sony Computer Science Laboratories · Keio University · Georgia Institute of Technology
Dance-to-music (D2M) generation aims to synthesize musically plausible soundtracks whose temporal structure aligns with a given dance performance. Representative D2M methods do not explicitly distinguish motion cues across body parts and frequency bands, while standard flow matching lacks a dedicated objective for supervising local music-latent dynamics. To address these issues, we propose Dyna2Music, a latent flow-matching framework that combines structured motion conditioning with explicit supervision of music-latent dynamics. An empirical analysis on AIST++ quantifies how spatial partitioning and frequency separation affect raw music-to-kinematic beat alignment and motion-reference density, informing the conditioning design. Accordingly, Dyna2Music decomposes joint velocities into slow and fast components and hierarchically fuses the resulting part-wise motion energy with pretrained joint features to condition music generation. To complement this representation, we introduce latent dynamics consistency (LDC), an auxiliary objective that matches adjacent-frame change magnitudes between a single-step clean-latent estimate and the paired reference. LDC makes local music-latent variation an explicit training target without adding trainable parameters or inference computation. Dyna2Music supports variable-length music generation, and experiments on AIST++ and TikTok demonstrate improved rhythmic alignment and audio quality over representative D2M baselines.
Changchang Sun, Lu Cheng, Yan Yan
University of Illinois Chicago · Penn State University