Paralinguistic speech tasks are often considered relatively language-agnostic, as they rely on extralinguistic acoustic cues rather than lexical content. However, prior studies report performance degradation under cross-lingual conditions, indicating non-negligible language dependence. Still, these studies typically focus on isolated language pairs or task-specific settings, limiting comparability and preventing a systematic assessment of task-level language dependence. We introduce the Cross-Lingual Transfer Matrix (CLTM), a systematic method to quantify cross-lingual interactions between pairs of languages within a given task. We apply the CLTM to two paralinguistic tasks, gender identification and speaker verification, using a multilingual HuBERT-based encoder, to analyze how donor-language data affects target-language performance during fine-tuning. Our results reveal distinct transfer patterns across tasks and languages, reflecting systematic, language-dependent effects.
Figures & tables
Figure 1 : Typical learning curve for a single language, showing the dynamic interval and derivative regimes.
Figure 2 : Learning curves for both tasks, showing performance as a function of training samples for a representative subset of languages. The task-specific dynamic interval [N,2N] used to compute the CLTM is highlighted.
Figure 3 : Architecture for gender recognition.
Figure 4 : Speaker verification pipeline: SID training via a classification head, then embeddings are L2-normalized and compared with cosine similarity for verification.
Figure 5 : Reduced CLTMs (16 representative languages) for gender recognition and speaker verification. Colors show how much adding donor-language data affects performance on a target language compared to adding the same amount of target-language data.
Metric
Gender Recognition
Speaker Verification
RFD1
0.162
2.970
Asymrel
0.175
1.084
prop+
99.97%
8.93%
reciprocity+
99.93%
1.69%
cosrows
0.990
0.615
intra-family+
4.98%
41.68%
Table 1 : Aggregate CLTM diagnostics computed on the full 44×44 matrices for gender recognition and speaker verification.
Language pair
dcent (Euclidean)
CLTM values
Russian–Belarusian
1.37
1.35 / 1.46
Kurmanji–Central Kurdish
1.53
1.22 / 0.91
German–Portuguese
2.06
-2.02 / -0.32
Galician–Central Kurdish
2.39
-1.09 / -1.13
Table 2 : Euclidean centroid distances between language speaker embeddings and CLTM values for selected SV pairs.
Cross-lingual speaker verification (SV) systems typically exhibit performance degradation when enrollment and test utterances are spoken in different languages. However, standard evaluation protocols confound language mismatch with inter-speaker variability, as evaluation is generally performed with different speakers across languages. In this work, we introduce a bilingual same-speaker evaluation set for five Iberian languages, enabling analysis of cross-lingual SV under constant speaker identity. We apply this setup to a HuBERT-based SV system previously shown to exhibit strong language dependence, and analyze results using the Cross-Lingual Transfer Matrix (CLTM) to study pairwise cross-lingual transfer. Our results show that speaker-related variability accounts for part of the observed degradation, but language mismatch remains the main driver of cross-lingual performance loss. These findings provide a more precise characterization of language dependence in cross-lingual SV.
Pol Buitrago, Javier Hernando
1Barcelona Supercomputing Center (BSC), Spain · 2Universitat Polit`ecnica de Catalunya (UPC), Spain
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out ρ=0.705 and R2=0.49, beating a non-typological control at ρ=0.62, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available here.
Dalton Raphael Harmsen, Swier Garst, Thomas van Osch +2
AMOR/e Lab, Eindhoven University of Technology, Eindhoven, The Netherlands · SURF, Amsterdam, The Netherlands
Cross-lingual transfer has become a central paradigm for extending natural language processing (NLP) technologies to low-resource languages. By leveraging supervision from high-resource languages, multilingual language models can achieve strong task performance with little or no labeled target-language data. However, it remains unclear to what extent cross-lingual transfer can substitute for language-specific efforts. In this paper, we synthesize prior research findings and data collection results on Luxembourgish, which, despite its typological proximity to high-resource languages and its presence in a multilingual context, remains insufficiently represented in modern NLP technologies. Across findings, we observe a fundamental interdependence between cross-lingual transfer and language-specific efforts. Cross-lingual transfer can substantially improve target-language performance, but its success depends critically on the availability of sufficiently high-quality, task-aligned target-language data. At the same time, such resources, particularly in low-resource settings, are typically too limited in scale to drive strong performance on their own. Instead, such resources reach their full potential only when leveraged within a cross-lingual framework. We therefore argue that cross-lingual transfer and language-specific efforts should not be viewed as competing alternatives. Instead, they function as complementary components of a sustainable low-resource NLP pipeline. Based on these insights, we provide practical guidelines for integrating and balancing cross-lingual transfer with language-specific development in sustainable low-resource NLP pipelines.
Fred Philippy, Siwen Guo, Jacques Klein +1
1SnT, University of Luxembourg, Luxembourg · 2Luxembourg Institute of Science and Technology, Luxembourg