BeatDance: Generating Beat-Consistent 3D Dance with Hierarchical Spatial-Temporal Modeling
Authors: Xiaojian Shen, Dahu Shi, Jianrong Zhang, Hai Li, Hongwei Zhao, Dawei Zhang, Yunzhi Zhuge, Zhiliang Wu, +2 more
Organizations: College of Software, Jilin University, Changchun, 130012, China · College of Computer Science and Technology, Zhejiang University, Hangzhou, 310007, China · ReLER, AAII, University of Technology Sydney, Sydney, Australia · College of Computer Science and Technology, Jilin University, Changchun, 130012, China · School of Computer Science and Technology, Zhejiang Normal University, Jinhua, 321004, China · School of Information and Communication Engineering, Dalian University of Technology, Dalian, 116024, China · School of Biomedical Engineering, Shenzhen University Medical School, Shenzhen University, Shenzhen, 518060, China · School of Computer Science and Informatics, Cardiff University, CF10 3AT Cardiff, U.K.
Generating realistic 3D dance from music is a challenging task that requires accurate synchronization with musical rhythms while capturing the spatial complexity of human motion. Although existing methods can generate physically plausible dance motions, they often struggle to achieve precise alignment with music, such as the beat. To address this limitation, we propose a novel diffusion-based framework, BeatDance, with two components: 1) We present a Hierarchical Decoupled Attention (HDA) module, which first disentangles the learning of human pose and temporal dynamics. A hierarchical structure is then employed to capture both short-term and long-term dependencies, thereby enhancing spatial-temporal modeling. 2) We adopt cycle-consistent learning by introducing an auxiliary dance-to-music module. During training, discrepancies between the reconstructed and original music induce a stronger loss signal, effectively encouraging the consistency property between the music and dance motion. Extensive experimental results demonstrate that our proposed approach outperforms recent competitive methods on two benchmark datasets.
Figures & tables
Figure 1: Comparison with the existing state-of-the-art method. (a) POPDG [ 24 ] uses a cross-attention and an alignment module to associate music and dance. However, it may face challenges in achieving beat-level synchronization. (b) We propose Hierarchical Decoupled Attention (HDA) and a cycle-consistent learning strategy. HDA hierarchically decouples spatial and temporal information, while cycle-consistent learning leverages the inverse dance-to-music mapping to further enhance temporal alignment.
Figure 2: Overview of BeatDance. Hierarchical Decoupled Attention (HDA) employs a hierarchical design to capture both short-term and long-term music-dance dependencies, building upon a Decoupled Attention (DA) that splits the modeling into spatial and temporal branches. Additionally, our cycle-consistent learning mechanism uses an auxiliary Dance-to-Music (D2M) model to reconstruct the source music features from the generated dance, thereby enforcing a tighter music-dance alignment.
Dataset
Method
Motion Quality
Motion Diversity
Alignment
PFC ↓
PBC →
Div k →
Div g →
BAS ↑
PopDanceSet
Ground Truth
1.5824
8.7365
9.0219
7.2931
0.174
Bailando [ 31 ]
3.9751
4.8863
5.1835
5.4342
0.230
EDGE † [ 35 ]
3.8366
4.0348
6.1709
5.7568
0.224
POPDG † [ 24 ]
1.8253
5.9492
7.1342
5.8314
0.233
DanceEditor † [ 50 ]
1.8972
7.8429
5.6338
5.2459
0.238
Table 1: Comparison with the state-of-the-art methods on PopDanceSet and AIST++ test sets. Metrics are marked as: ↑ higher is better, ↓ lower is better, → closer to GT is better. Bold ( underlined ) indicates the best (second-best) results. † denotes diffusion-based methods.
Figure 3: Visual results on the PopDanceSet. We visually compare BeatDance with the state-of-the-art method POPDG [ 24 ] . Penetration artifacts (red) and Distorted motions (blue) are highlighted.
Figure 4: Visual comparison of beat consistency.
Comparison
BeatDance Win Rate
BeatDance vs. POPDG
91%
BeatDance vs. DanceEditor
94%
Table 2: Pairwise user-study results on the PopDanceSet test set.
Method
PFC
PBC
Div k
Div g
BAS
w/o DA
15.5267
6.3755
8.9975
6.1423
0.218
w/o HM
1.6046
5.7520
6.9480
6.3608
0.240
w/o CCL
1.5594
5.8300
6.6467
5.5107
0.252
Ours
1.5390
7.7493
6.7770
6.1801
0.273
Table 3: Ablation study of different components. DA, HM, and CCL denote the Decoupled Attention, Hierarchical Mechanism, and Cycle Consistency Learning, respectively.
Order
PFC
PBC
Div k
Div g
BAS
1
1.2401
7.2673
6.6585
5.8027
0.229
2
1.3925
6.9268
6.6322
5.4058
0.242
3
1.5390
7.7493
6.7770
6.1801
0.273
4
1.7447
7.3459
6.5358
5.6094
0.255
5
1.1339
6.4334
6.2685
5.1131
0.256
Table 4: Ablation study of the order of the Taylor series expansion.
τ
PFC
PBC
Div k
Div g
BAS
5
2.6125
8.5605
7.5657
5.8857
0.232
15
1.6489
6.9349
6.6073
5.5570
0.249
30
1.5390
7.7493
6.7770
6.1801
0.273
50
1.4113
6.7600
8.7246
6.0828
0.218
75
1.8742
5.3648
7.2374
5.5517
0.234
Table 5: Ablation study of the local window size τ .
Method
BCS ↑
BHS ↑
FAD ↓
BAS ↑
Ground Truth
100
100
0
0.174
GAN
86.5
52.8
7.63
0.191
VAE+GAN
100.2
54.6
7.52
0.193
Diffusion
110.8
62.1
7.60
0.193
VAE+Diffusion
113.8
64.8
6.86
0.195
Table 6: Ablation study of different generative architectures for the dance-to-music generation task.