Dance-to-music (D2M) generation aims to synthesize musically plausible soundtracks whose temporal structure aligns with a given dance performance. Representative D2M methods do not explicitly distinguish motion cues across body parts and frequency bands, while standard flow matching lacks a dedicated objective for supervising local music-latent dynamics. To address these issues, we propose Dyna2Music, a latent flow-matching framework that combines structured motion conditioning with explicit supervision of music-latent dynamics. An empirical analysis on AIST++ quantifies how spatial partitioning and frequency separation affect raw music-to-kinematic beat alignment and motion-reference density, informing the conditioning design. Accordingly, Dyna2Music decomposes joint velocities into slow and fast components and hierarchically fuses the resulting part-wise motion energy with pretrained joint features to condition music generation. To complement this representation, we introduce latent dynamics consistency (LDC), an auxiliary objective that matches adjacent-frame change magnitudes between a single-step clean-latent estimate and the paired reference. LDC makes local music-latent variation an explicit training target without adding trainable parameters or inference computation. Dyna2Music supports variable-length music generation, and experiments on AIST++ and TikTok demonstrate improved rhythmic alignment and audio quality over representative D2M baselines.
Figures & tables
Figure 1: Beat alignment score (BAS) and kinematic-beat density for different motion references on paired dance-music clips from the AIST++ training set. Filled markers and error bars (left axis) indicate mean music-to-kinematic BAS and 95% confidence intervals, respectively; the grey line (right axis) shows kinematic-beat density (beats/s). The spatial partitions comprise the whole body (one part), upper and lower body (two parts), or the torso, left and right arms, and left and right legs (five parts). One band denotes motion without frequency separation; two bands denote separate slow and fast components. Finer spatial partitions and slow-fast separation yield denser motion references, helping to explain the increase in raw BAS.
Figure 2: Overview of Dyna2Music. Left: SSMD extracts part-wise slow/fast motion energy and complementary features from frozen MotionBERT. Middle: The HSF-Encoder fuses these signals into conditioning tokens through attention and adaptive gating. Right: The tokens guide a latent flow-matching transformer trained with velocity regression and latent dynamics consistency (LDC), which supervises adjacent-frame change magnitudes. Generated latents are decoded into music by a frozen waveform VAE.
Category
Method
AIST++
TikTok
BCS ↑
CSD ↓
BHS ↑
HSD ↓
F1 ↑
BCS ↑
CSD ↓
BHS ↑
HSD ↓
F1 ↑
General video-to-music generation
CMT
32.77
23.43
10.51
8.74
14.80
20.20
18.43
12.01
11.54
13.64
Diff-Foley
31.02
23.10
12.85
9.51
16.80
23.21
20.73
7.86
9.22
9.72
M 2 UGen
32.10
20.81
43.19
25.70
33.24
19.38
11.91
40.64
25.83
24.71
VidMuse
33.41
18.41
49.72
17.41
38.06
19.33
10.12
42.64
15.79
25.37
GVMGen
28.18
27.77
12.10
16.35
12.49
18.54
21.56
10.60
11.78
10.82
Table 1: Onset-level synchrony (%) on AIST++ and TikTok. BCS/BHS denote precision/recall; CSD/HSD are their across-clip standard deviations. Best values are bold.
Category
Method
AIST++
TikTok
FAD v ↓
FAD p ↓
FAD c ↓
FAD v ↓
FAD p ↓
FAD c ↓
General video-to-music generation
CMT
10.672
68.863
0.958
18.556
80.538
1.368
Diff-Foley
31.843
184.737
1.606
15.147
96.495
1.318
M 2 UGen
4.081
41.447
0.808
11.493
60.614
1.051
VidMuse
4.778
28.645
0.975
9.930
52.895
1.008
GVMGen
12.694
56.269
1.132
10.474
54.344
1.147
Table 2: FAD (lower is better) on AIST++ and TikTok using VGGish ( v ), PANNs ( p ), and CLAP ( c ) features. Best and second-best values are bold and underlined, respectively.
Figure 3: Music-to-kinematic BAS ( σ=50 ms; higher is better) using the five-part, two-band motion reference on (a) AIST++ and (b) TikTok. GT denotes paired music; colors identify method groups.
Figure 4: Beat-relative rhythm on AIST++: (a) standardized onset envelopes aligned to reference-music beats; (b) autocorrelation with lag normalized by the reference track’s nominal beat period. Each column compares Dyna2Music with one baseline; GT is evaluated against its own timing. Shading denotes 95% bootstrap confidence intervals over 10 tracks.
Figure 5: Mean pairwise CLAP cosine similarity across different dances: (a) AIST++ ( 180 sequence pairs); (b) TikTok ( 1,358 clip pairs). The dashed line marks the corresponding similarity of GT music. Error bars show 95% bootstrap intervals over music tracks (AIST++) or source videos (TikTok). Higher values indicate more similar embeddings.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure A1: Matched-minus-mismatched rhythm statistics on AIST++: (1) onset strength within ±0.1 beats of reference beats; (2) autocorrelation within ±0.05 beats of one- and two-beat lags. Points and bars show track-averaged differences and 95% bootstrap confidence intervals; zero denotes equal matched and mismatched scores.
Figure A2: Music-to-kinematic BAS ( σ=50 ms) on AIST++ under four additional motion references. Panel labels specify part and band counts; colors follow Figure 3 . Higher values indicate closer alignment within each configuration.
Figure A3: Music-to-kinematic BAS ( σ=50 ms) on TikTok under four additional motion references. Panel labels specify part and band counts; colors follow Figure 3 . Higher values indicate closer alignment within each configuration.
Dataset
Parts × Bands
BCS ↑
CSD ↓
BHS ↑
HSD ↓
F1 ↑
AIST++
1×1
53.20
26.31
52.87
25.48
47.49
2×1
47.14
24.98
58.43
22.95
47.66
2×2
45.92
24.30
52.34
24.31
43.58
5×1
50.43
26.10
56.09
26.18
47.38
5×2 (full)
60.44
26.57
63.54
25.70
55.76
TikTok
1×1
19.71
11.29
26.66
11.79
20.89
Appendix
Table A1: Ablation of motion conditioning on AIST++ and TikTok using onset-level metrics (%). One band denotes merged slow/fast energy. Full-model results are reproduced from Table 1 ; best values within each dataset are bold.
Figure A4: Sensitivity to λLDC on (a) AIST++ and (b) TikTok. Blue: onset-level F1 on a 0 - 1 scale (left; higher is better). Orange: FAD c (right; lower is better). Horizontal positions are categorical; markers denote evaluated weights, with no 0.05 evaluation on TikTok.
Method
OVL ↑
REL ↑
AudioLDM-ControlNet
3.3
3.0
SONIQUE
3.1
3.4
Dyna2Music (Ours)
4.3
4.5
Appendix
Table A2: Subjective evaluation on AIST++: mean opinion scores for overall musical quality (OVL) and dance-music relevance (REL) on a 1 - 5 scale (higher is better). Best scores are bold.