Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
Organizations: AMOR/e Lab, Eindhoven University of Technology, Eindhoven, The Netherlands · SURF, Amsterdam, The Netherlands
Abstract
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out and , beating a non-typological control at , which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.
Figures & tables
| Model | LOLO | L2LO | |
| similarity proxy | 0.037 | 0.029 | |
| ridge GLM (sim/script/gene/geo) | 0.162 | 0.014 | 0.154 |
| random forest (all features) | 0.615 | 0.350 | 0.648 |
| tuned RF (typology) | 0.705 | 0.492 | 0.702 |
| metadata RF (matched-capacity, 5 fields) | 0.581 | 0.264 | 0.611 |
| metadata RF (own-sweep tuned) | 0.620 | 0.375 | 0.621 |
| Split | |
|---|---|
| cross-script pairs | 0.728 |
| same-script pairs | 0.600 |
| Latin–Latin only | 0.509 |
| script-only model | 0.138 |
| LOLO (leave-one-language-out) | 0.705 |
| LOSO (leave-one-script-out) | 0.776 |
| top- | 1 | 5 | 10 | 25 | 50 | 100 | 200 | 387 |
|---|---|---|---|---|---|---|---|---|
| global-rank | 0.622 | 0.613 | 0.573 | 0.570 | 0.647 | 0.627 | 0.742 | 0.705 |
| leak-free | 0.622 | 0.627 | 0.573 | 0.679 | 0.657 | 0.716 | 0.710 | 0.705 |
| Operationalization | #1 source | #2 source |
|---|---|---|
| empirical (ATLAS, raw BTS mean) | English ( ) | Hebrew ( ) |
| on-scale reweighted RF | English ( ) | Hebrew ( ) |
| OOD-generalizer (cross-family) | English ( ) | Hebrew ( ) |
| debiased residual (mean-zero, off-scale) | Marathi ( ) | Filipino ( ) |
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
| Component | Exact configuration |
|---|---|
| Empirical target | the directed BTS transfer matrix digitized from Figure C.2 of the ATLAS paper, read at one decimal |
| Typological features | Grambank with 195 features and WALS with 192, keyed by Glottocode; the union carries 387 features |
| Modeling set | ATLAS-24: the 24 ATLAS languages with Grambank coverage of at least ; the floor is Grambank-only, so WALS adds feature columns but no languages; all directed pairs with , which is 552 |
| Input representation | , the one-hot code profiles of source and target plus a per-feature disagreement block, as defined above |
| Main model | random forest with 200 trees, random_state , and max_features ; under WALS alone max_features is log2 ; all remaining parameters at scikit-learn defaults |
| Ridge models | throughout: the similarity proxy, the GLM on similarity, script, genealogical, and geographic distances, the script-only model, and the feature-importance ranker; the bias ridge standardizes its inputs first |
| Analysis | Exact configuration |
|---|---|
| E1 (RQ1): held-out performance and capacity | |
| Tier comparison | four tiers under identical folds: the similarity-proxy ridge, the GLM ridge, the 200-tree RF at scikit-learn defaults, and the tuned RF; metadata single-feature ridges for script, family, and Wikipedia and their combination, each with within-fold standardization on its own coverage-restricted universe: script and Wikipedia cover 24 languages, the WALS family map covers 22 |
| Distance baselines | five non-typological distance baselines under the identical folds and metrics, each a single-column ridge except where noted, with within-fold standardization on the full 24-language universe: the lang2vec cosine over the concatenated imputed syntax, phonology, and inventory vectors; the URIEL design with six columns, genetic from the language-family vector, geographic as the great-circle distance on Glottolog coordinates normalized by half the circumference, syntactic, phonological, and inventory as one minus the cosine of the respective vectors, and featural as their mean; and three Jaccard overlaps over per-language Wikipedia lead-paragraph corpora of 2,000 random articles per edition with a committed sha256 manifest, namely byte-level BPE vocabularies of 8,000 entries, pooled character 1-to-3-gram sets over the 5,000 most frequent word types, and the 10,000 most frequent word types; lang2vec 1.1.2 imputed feature sets cover all 24 languages, and the learned embeddings are unusable because seven of the 24 languages are absent and the embedding file does not load under NumPy 2 |
| Leave- -languages-out | ; is the deterministic LOLO partition; for the tuned RF averages over 50 random disjoint partitions of the 24 languages and the remaining tiers over 25; a pair is held out when either endpoint is in the held-out group; the reported is the partition mean and the interval is the partition spread; seed 42 |
| Permutation null | fold-respecting: within each fold only the training labels are permuted and the held-out labels are untouched; ; ; seed 42 |
| Cluster bootstrap | 95% intervals by resampling the 24 source languages with replacement, , seed 42; paired bootstrap of against the script-only ridge and against the GLM on shared pairs; a language-level variant resamples the 24 languages and reweights each pair by the product of how often its two endpoints were drawn |
| ID | Feature | ID | Feature |
|---|---|---|---|
| Grambank (195 features) | |||
| GB020 | Are there definite or specific articles? | GB021 | Do indefinite nominals commonly have indefinite articles? |
| GB022 | Are there prenominal articles? | GB023 | Are there postnominal articles? |
| GB024 | What is the order of numeral and noun in the NP? | GB025 | What is the order of adnominal demonstrative and noun? |
| GB026 | Can adnominal property words occur discontinuously? | GB027 | Are nominal conjunction and comitative expressed by different elements? |
| GB028 | Is there a distinction between inclusive and exclusive? | GB030 | Is there a gender distinction in independent 3rd person pronouns? |
| Feature source | #feat | LOLO | LOLO | unseen-lang |
|---|---|---|---|---|
| Grambank (GB) | 195 | 0.780 | 0.607 | 0.494 |
| WALS | 192 | 0.648 | 0.412 | 0.307 |
| Grambank+WALS (operational) | 387 | 0.705 | 0.492 | 0.405 |
| Feature source | best max_features | LOLO | |
|---|---|---|---|
| Grambank | 0.3 | 0.716 | 0.511 |
| Grambank+WALS | sqrt | 0.673 | 0.447 |
| WALS | log2 | 0.648 | 0.412 |
| Feature design | #feat | LOLO | |
|---|---|---|---|
| Grambank | 195 | 0.780 | 0.607 |
| Grambank+WALS, real block | 387 | 0.705 | 0.492 |
| Grambank+WALS, block scrambled (10 perms) | 387 |
| Floor | language added | glm | full | tuned | tuned | unseen-lang | |
|---|---|---|---|---|---|---|---|
| 24 | – | ||||||
| 25 | Telugu | ||||||
| 26 | Thai | ||||||
| 27 | Indonesian | ||||||
| 28 | Central Kurdish |
| Model (same RF, same folds) | #feat | LOLO | |
|---|---|---|---|
| metadata RF (script/family/genus/geo/wiki) | 5 | 0.581 | 0.264 |
| metadata RF, own-sweep tuned | 5 | 0.620 | 0.375 |
| typology RF (Grambank+WALS) | 387 | 0.705 | 0.492 |
| typology RF, budget-matched to metadata | 5 | 0.597 | – |
| Factor | Mantel | Mantel | MRM coef | MRM |
|---|---|---|---|---|
| typological | 0.266 | 0.015 | 0.249 | 0.281 |
| script | 0.334 | 0.017 | 0.149 | 0.035 |
| genealogical | 0.226 | 0.051 | 0.098 | 0.693 |
| geographic | 0.169 | 0.078 | 0.029 | 0.803 |
| Axis | Group | grouped | matched LOLO |
|---|---|---|---|
| LOSO | Arabic | 0.459 | 0.683 |
| LOSO | Cyrillic | 0.935 | 0.749 |
| LOSO | Devanagari | 0.912 | 0.622 |
| LOSO | Greek | 0.935 | 0.572 |
| LOSO | Han | 0.730 | 0.819 |
| LOSO | Hangul | 0.793 | 0.643 |
| Exp. | RQ | Quantity | Estimate | 95% CI |
|---|---|---|---|---|
| E1 | RQ1 | LOLO ; permutation | ||
| E1 | RQ1 | LOLO ; language-level bootstrap 95% CI | ||
| E1 | RQ1 | pure-unseen-language (Grambank+WALS) | – | |
| E3 | RQ1 | LOSO / LOFO macro- | / | / |
| E4 | RQ1 | top-25 importance Jaccard (random ) | ||
| E5 | RQ2 | residual gap, Marathi English (bootstrap over targets); |
| Baseline | LOLO | true |
|---|---|---|
| character -gram overlap | 0.225 | 0.044 |
| subword tokenizer overlap | 0.222 | 0.044 |
| lang2vec featural cosine | 0.148 | 0.012 |
| word vocabulary overlap | 0.142 | 0.010 |
| URIEL six-distance ridge | 0.095 | |
| tuned RF (typology, reference) | 0.705 | 0.492 |
| Nuisance descriptor | std. weight | mean contribution (BTS) |
|---|---|---|
| source resource ( ) | ||
| script difference | ||
| genealogical distance | ||
| geographic distance | ||
| target resource ( ) |
| Bias features | (%) | LOLO | rank (full) | rank (truth) | top | |
| (1 feature) | ||||||
| geo | 0.013 | 34.8 | 0.719 | 0.989 | 0.984 | English |
| res s | 0.060 | 26.0 | 0.718 | 0.990 | 0.987 | English |
| scr | 0.051 | 28.0 | 0.717 | 0.979 | 0.984 | English |
| gene | 0.023 | 33.1 | 0.712 | 0.992 | 0.988 | English |
| res t | 0.000 | 36.8 | 0.705 | 0.989 | 0.990 | English |
| Rank | empirical | debiased | centrality | OOD-generalizer | ||||
|---|---|---|---|---|---|---|---|---|
| source | score | source | score | source | score | source | score | |
| 1 | English | 0.022 | Marathi | 0.361 | Ukrainian | 0.717 | English | -0.011 |
| 2 | Modern Hebrew | -0.074 | Filipino | 0.355 | Portuguese | 0.700 | Modern Hebrew | -0.065 |
| 3 | French | -0.109 | Modern Hebrew | 0.328 | Serbian-Croatian-Bosnian | 0.679 | Filipino | -0.114 |
| 4 | Filipino | -0.113 | Swahili | 0.235 | Catalan | 0.678 | Vietnamese | -0.114 |
| 5 | Standard Arabic | -0.113 | Standard Arabic | 0.198 | Italian | 0.677 | Standard Arabic | -0.120 |
| Rank | empirical | bias-removed | OOD-generalizer | |||
|---|---|---|---|---|---|---|
| source | count | source | count | source | count | |
| 1 | English | 8/23 | Marathi | 23/23 | Modern Hebrew | 3/20 |
| 2 | French | 5/23 | Filipino | 23/23 | English | 2/9 |
| 3 | Modern Hebrew | 3/23 | Modern Hebrew | 23/23 | French | 1/9 |
| 4 | Portuguese | 3/23 | Standard Arabic | 23/23 | Modern Greek | 1/9 |
| 5 | Catalan | 3/23 | Western Farsi | 23/23 | Filipino | 0/22 |
| Rank | empirical | bias-removed | OOD-generalizer | |||
|---|---|---|---|---|---|---|
| source | count | source | count | source | count | |
| 1 | English | 14/24 | Marathi | 10/24 | Modern Hebrew | 14/24 |
| 2 | Modern Hebrew | 6/24 | Filipino | 8/24 | Swahili | 10/24 |
| 3 | Western Farsi | 5/24 | Modern Hebrew | 6/24 | English | 9/24 |
| 4 | French | 4/24 | Swahili | 0/24 | Vietnamese | 9/24 |
| 5 | Filipino | 4/24 | Standard Arabic | 0/24 | Standard Arabic | 9/24 |