Towards cross-cultural study of folksong lyrics with machine translation
Authors: Anna Dvořáková, Anna Aljanaki, Danbinaerin Han, Peter van Kranenburg, Matěj Kratochvíl, Inna Lisniak, Zdeněk Vejvoda, Jan Hajič
Organizations: Charles University Czech Republic · University of Music and Performing Arts Graz Austria · Graduate School of Culture Technology KAIST Republic of Korea · Utrecht University Netherlands · Institute of Ethnology Czech Academy of Sciences Czech Republic · Estonian Literary Museum; M. T. Rylsky Institute of Art Studies, Folkloristics and Ethnology NAS of Ukraine
Music is universally present in human societies. Ethnomusicologists have long been documenting the diverse expressions of human musicality, and comparative musicology has recently brought several studies of folksong to a more global scale. Such cross-cultural research has not been conducted on lyrics: the language barrier has so far prevented work with multi-lingual data. However, Natural Language Processing (NLP) technologies have reached a stage where this language barrier may no longer be prohibitive. Combining folksong lyrics corpora across five languages, we machine-translate them to a pivot language with a pre-trained neural topic model, and we examine the relationship between content and social function within each language, and across languages for wedding songs. As expected, human evaluation of translation results shows that non-Indo-European languages suffer from overall worse translation quality. Experiments with topic models then indicate that the content of lyrics is at best partially related to the social function of folksongs across all languages. These experiments are just first steps into cross-cultural folk musics lyrics analysis; however, they do indicate that a previously unobserved web of cross-cultural relationships beyond ethnomusicological typologies may be uncovered through the study of what people sing across the world's diverse folk musics.
Figures & tables
Language
No. of songs
No. of wedding songs
No. of songs monoling. exps.
Czech
3320
211
2829
Dutch
3041
149
2612
Estonian
4515
100
3553
Korean
15861
480
7767
Ukrainian
1764
257
1585
Table 1: Number of songs with lyrics and number of wedding songs (Korean: marriage songs) in each corpus and number of songs that were used in the experiments from each corpus. The high reduction on the Korean part was caused by a large number of missing typology labels needed for the experiments
Language
# words evaluated
Mean importance
Mean correctness
Mean overall
Czech
764
0.84
0.97
3.67
Dutch
1396
0.48
0.93
4.18
Estonian
903
0.79
0.76
2.94
Korean
584
0.95
0.79
3.00
Ukrainian
1144
0.64
0.90
3.32
Table 2: MT manual evaluation results.
Corpus
Avg. weighted F1 score
Avg. # PCA components
no. effective training songs
Czech
0.26
411
1685
Dutch
0.53
230
499
Estonian
0.51
461
2415
Korean
0.63
488
3552
Ukrainian
0.52
293
1040
Weddings
0.63
216
500
Table 3: Classification results on monolingual corpora and on wedding songs (all languages, and the 3 well-translated ones): avg. weighted F1 score, and avg. number of PCA components explaining 95 % of the embeddings variance, and the number of training songs after subsampling.
Figure 1: Classification results: row-normalised confusion matrices for top 10 labels in each language, and on the languages for wedding songs. We report recall because we care most about the question: “Is this label hard to identify?”.
University of Luxembourg Faculty of Science, Technology and Medicine Luxembourg, Luxembourg · University of Warsaw Institute of Informatics Warsaw, Poland