Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension
Authors: Maral Ebrahimzadeh, Gilberto Bernardes, Sebastian Stober
Organizations: Otto von Guericke University Magdeburg, Artificial Intelligence Lab, Magdeburg, Germany · INESC TEC, Faculty of Engineering, University of Porto, Portugal
Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as skip-gram have been used to learn chord embeddings from symbolic corpora, but their ability to recover tonal structure and its relation to tonal tension remains underexplored. In this work, we investigate how skip-gram chord embeddings reflect tonal structure and whether they provide a useful basis for analyzing structural aspects of tonal tension. Using chord sequences with and without transposition-based augmentation, we evaluate the learned spaces from geometric, functional, and tension-related perspectives. We show that augmented embeddings exhibit strong transposition equivariance, recover a clear circle-of-fifths structure, and support interpretable shifts between key-related regions of the learned space. We then derive embedding-based measures from chord-to-key distance and contextual chord-distance relations, and show that they capture meaningful aspects of tonal tension structure through correspondence with matched tonal measures and moderate alignment with human tension profiles. Across analyses, transposition-based augmentation generally improves the stability, tonal coherence, and interpretability of the learned space.
Figures & tables
Dataset
Seqs.
Min
Max
Vocab
M/m cov.
BPS-FH
604
3
290
24
89.1%
Isophonics
515
3
128
24
94.9%
Table 1 : Summary statistics of the original skip-gram corpora, including the number of sequences, sequence-length range, vocabulary size, and the proportion of valid chord events retained after triad reduction and filtering.
Figure 1 : Two-dimensional MDS projections of the learned 24-chord spaces for representative augmented and non-augmented models.
Table 3 : Mean transposition-equivariance scores over the 11 non-identity pitch-class shifts.
Dataset
Model
LOO gap ↑
AUC ↑
BPS-FH
AUG
0.363 / 0.313
0.931 / 0.891
BPS-FH
ORIG
0.281 / 0.239
0.891 / 0.849
Isophonics
AUG
0.369 / 0.268
0.946 / 0.857
Isophonics
ORIG
0.269 / 0.197
0.922 / 0.825
Table 4 : Summary of emergent key structure. Values are reported as major / minor. The leave-one-out (LOO) gap is defined as mean(dout)−mean(din) , so higher values indicate stronger separation between out-of-key and in-key chords. Higher AUC indicates better discrimination of in-key versus out-of-key distances.
Dataset
Model
Cont. Δ↓
Target hit ↑
Source escape ↑
BPS-FH
AUG
-0.921
0.630
0.685
BPS-FH
ORIG
-0.724
0.377
0.410
Isophonics
AUG
-0.840
0.639
0.667
Isophonics
ORIG
-0.678
0.295
0.305
Table 5 : Tonal region steering results on source-only chords at α=2.0 .
Dataset
Model
Tkey
Tctx
Tcomb ( β=0.5 )
relL2 ↓
corr ↑
relL2 ↓
corr ↑
relL2 ↓
corr ↑
BPS-FH
AUG
0.063
0.883
0.074
0.995
0.065
0.976
BPS-FH
ORIG
0.173
0.321
0.246
0.925
0.180
0.770
Isophonics
AUG
0.144
0.804
0.200
0.968
0.155
0.914
Isophonics
ORIG
0.246
0.305
0.391
0.861
0.267
0.696
Table 6 : Transposition stability averaged over sequences and non-identity transpositions ( k=1,…,11 ). Lower relative ℓ2 is better, higher correlation is better.
Contextual chord distance
Key-relative chord distance
Dataset
Model
Pearson ↑
Spearman ↑
Pearson ↑
Spearman ↑
BPS-FH
AUG
0.733
0.701
0.624
0.698
BPS-FH
ORIG
0.707
0.687
0.513
0.511
Isophonics
AUG
0.637
0.628
0.526
0.592
Isophonics
ORIG
0.617
0.599
0.422
0.440
Table 7 : Length-weighted means of per-sequence correlations between embedding and matched TIV-based quantities. Higher values indicate stronger correspondence.
Figure 2 : Predicted and human-rated tension curves for three representative progressions, comparing the embedding-based measure with TIV, Lerdahl’s model, MorpheuS, and mean participant ratings.
Setting
Pearson r
Spearman ρ
RMSE
All-12 fit
0.473
0.531
0.189
LOOCV
0.406
0.456
0.199
Table 8 : Global performance of the embedding-based dynamic measure against mean human tension ratings on the 12 progressions.
Figure 3 : Score representations of the three representative examples discussed in Figure 2 : (a) Progression A, (b) Progression B, and (c) Progression C.