cs.CLSep 26, 2026

Language Distances are Practical for Equitable Cross-Lingual Transfer

Authors: York Hay Ng, Razan Ahsan Rifandi, Aditya Khan, En-Shiun Annie Lee

Abstract

Cross-lingual transfer is strongly conditional on how the source language is chosen, but it is impractical to determine the best candidate source for every target language, especially for low-resource target languages. Language distances are widely used to rank candidate sources due to their correlation with transfer efficacy and applicability in resource-sparse settings. However, the reliability of distance-based rankers across tasks and resource levels remains underexplored. We therefore present the first equity-focused evaluation of paradigms for ranking source languages, studying resource-level inequality and task inequality across ten cross-lingual tasks and two multilingual models. While both inequalities are most pronounced for individual language distances and an English-always baseline, they are substantially reduced by training-free composite distances, and nearly eliminated by trained rankers. We further demonstrate the reliability of rankers using language distances compared to rankers using language model internals. Overall, we find that language distances provide a practical basis for equitable and performant transfer language selection. We recommend using trained rankers when task-specific transfer evaluations are available, and composite distances otherwise.

Explore similar work

CardsList
  1. Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies

    Sep 30, 2026Dalton Raphael Harmsen, Swier Garst, Thomas van Osch +2Task Transferability EstimationLinguistic Typology

  2. Cross-Lingual Transfer for Machine Translation in Turkic Languages

    Jul 31, 2026Omer Burak Cinar, Mehmet Mert Dalkilic, Cagri ToramanMultilingual Language ModelsLow-Resource MT

  3. DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

    Aug 11, 2026Akriti Dhasmana, Aarohi Srivastava, David ChiangLearning to RankLow-Resource Speech Recognition