Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological features, and whether the prominence of high-resource source languages reflects typology or data quality and quantity. We show that typological databases contain cheap and dense signals about cross-lingual transfer. Our typology-only random forest on a 24-language prior-work transfer matrix scores leave-one-language-out ρ=0.705 and R2=0.49, beating a non-typological control at ρ=0.62, which verifies the ability of typology-only predictions to reconstruct costly measured cross-lingual transfer. The signal survives leave-one-script-out and leave-one-family-out protocols, so script and family confounding do not explain the effect. By decomposing the transfer into a typology term and a resource-and-script bias term, we find the best-source ranking sensitive to this bias. In contrast, typology is not affected by this bias, which makes it a zero-compute screening tool that replaces hundreds of training runs with a model fit. Our code is available \href{https://github.com/dharmsen/typo-x-ling-transfer}{here}.
Figures & tables
Model
LOLO ρ
R2
L2LO ρ
similarity proxy
0.037
−0.010
0.029
ridge GLM (sim/script/gene/geo)
0.162
0.014
0.154
random forest (all features)
0.615
0.350
0.648
tuned RF (typology)
0.705
0.492
0.702
metadata RF (matched-capacity, 5 fields)
0.581
0.264
0.611
metadata RF (own-sweep tuned)
0.620
0.375
0.621
Table 1: LOLO and leave-2-languages-out predictive performance on ATLAS-24 (E1).
Split
ρ
cross-script pairs
0.728
same-script pairs
0.600
Latin–Latin only
0.509
script-only model
0.138
LOLO (leave-one-language-out)
0.705
LOSO (leave-one-script-out)
0.776
Table 2: Cross-script robustness and grouped holdouts (E2, E3) on ATLAS-24.
top- k
1
5
10
25
50
100
200
387
global-rank ρ
0.622
0.613
0.573
0.570
0.647
0.627
0.742
0.705
leak-free ρ
0.622
0.627
0.573
0.679
0.657
0.716
0.710
0.705
Table 3: Top- k feature sweep for Experiment E4 under Grambank+WALS; the leak-free row selects its top- k within each fold.
Operationalization
#1 source
#2 source
empirical (ATLAS, raw BTS mean)
English ( +0.022 )
Hebrew ( −0.074 )
on-scale reweighted RF
English ( −0.005 )
Hebrew ( −0.070 )
OOD-generalizer (cross-family)
English ( −0.011 )
Hebrew ( −0.065 )
debiased residual (mean-zero, off-scale)
Marathi ( +0.36 )
Filipino ( +0.36 )
Table 4: “Best pre-training source” over a fixed candidate set (Experiment E5); rows are not comparable across scales.
Appendix figures & tables27 assets
Supplementary material from the paper’s appendix.
Appendix
Component
Exact configuration
Empirical target
the directed BTS transfer matrix digitized from Figure C.2 of the ATLAS paper, read at one decimal
Typological features
Grambank with 195 features and WALS with 192, keyed by Glottocode; the union carries 387 features
Modeling set
ATLAS-24: the 24 ATLAS languages with Grambank coverage of at least 0.70 ; the floor is Grambank-only, so WALS adds feature columns but no languages; all 24×23 directed pairs with s=t , which is 552
Input representation
[Ms,Mt,Δ] , the one-hot code profiles of source and target plus a per-feature disagreement block, as defined above
Main model
random forest with 200 trees, random_state =42 , and max_features =0.5 ; under WALS alone max_features is log2 ; all remaining parameters at scikit-learn defaults
Ridge models
α=1.0 throughout: the similarity proxy, the GLM on similarity, script, genealogical, and geographic distances, the script-only model, and the feature-importance ranker; the bias ridge standardizes its inputs first
Appendix
Table 5: Shared setup underlying all experiments in Table 6 . All typology experiments use the Grambank+WALS feature source and the one-decimal ATLAS target at n=24 , which is 552 directed pairs, unless a row of Table 6 states a deviation.
Analysis
Exact configuration
E1 (RQ1): held-out performance and capacity
Tier comparison
four tiers under identical folds: the similarity-proxy ridge, the GLM ridge, the 200-tree RF at scikit-learn defaults, and the tuned RF; metadata single-feature ridges for script, family, and Wikipedia and their combination, each with within-fold standardization on its own coverage-restricted universe: script and Wikipedia cover 24 languages, the WALS family map covers 22
Distance baselines
five non-typological distance baselines under the identical folds and metrics, each a single-column ridge except where noted, with within-fold standardization on the full 24-language universe: the lang2vec cosine over the concatenated imputed syntax, phonology, and inventory vectors; the URIEL design with six columns, genetic from the language-family vector, geographic as the great-circle distance on Glottolog coordinates normalized by half the circumference, syntactic, phonological, and inventory as one minus the cosine of the respective vectors, and featural as their mean; and three Jaccard overlaps over per-language Wikipedia lead-paragraph corpora of 2,000 random articles per edition with a committed sha256 manifest, namely byte-level BPE vocabularies of 8,000 entries, pooled character 1-to-3-gram sets over the 5,000 most frequent word types, and the 10,000 most frequent word types; lang2vec 1.1.2 imputed feature sets cover all 24 languages, and the learned embeddings are unusable because seven of the 24 languages are absent and the embedding file does not load under NumPy 2
Leave- k -languages-out
k∈{1,2,3} ; k=1 is the deterministic LOLO partition; for k≥2 the tuned RF averages over 50 random disjoint partitions of the 24 languages and the remaining tiers over 25; a pair is held out when either endpoint is in the held-out group; the reported ρ is the partition mean and the interval is the partition spread; seed 42
Permutation null
fold-respecting: within each fold only the training labels are permuted and the held-out labels are untouched; B=200 ; p=(1+#{null≥observed})/(B+1) ; seed 42
Cluster bootstrap
95% intervals by resampling the 24 source languages with replacement, B=2000 , seed 42; paired bootstrap of Δρ against the script-only ridge and against the GLM on shared pairs; a language-level variant resamples the 24 languages and reweights each pair by the product of how often its two endpoints were drawn
Appendix
Table 6: Exact configuration of the five experiments E1 to E5 and their sub-analyses. Each experiment addresses the research question named in its group header, with E1 to E4 probing RQ1 and E5 probing RQ2. Settings inherited from Table 5 are not repeated. Repository scripts preserve their own identifiers, so the experiment labels E1 to E5 are assigned here for readability.
ID
Feature
ID
Feature
Grambank (195 features)
GB020
Are there definite or specific articles?
GB021
Do indefinite nominals commonly have indefinite articles?
GB022
Are there prenominal articles?
GB023
Are there postnominal articles?
GB024
What is the order of numeral and noun in the NP?
GB025
What is the order of adnominal demonstrative and noun?
GB026
Can adnominal property words occur discontinuously?
GB027
Are nominal conjunction and comitative expressed by different elements?
GB028
Is there a distinction between inclusive and exclusive?
GB030
Is there a gender distinction in independent 3rd person pronouns?
Appendix
Table 7: The 387 typological features used in this work, in natural ID order within each source: 195 Grambank parameters and 192 WALS features.
Feature source
#feat
LOLO ρ
LOLO R2
unseen-lang ρ
Grambank (GB)
195
0.780
0.607
0.494
WALS
192
0.648
0.412
0.307
Grambank+WALS (operational)
387
0.705
0.492
0.405
Appendix
Table 8: Feature-source comparison under identical per-source tuning, from Experiment E1. The LOLO ordering is preserved on the pure-unseen-language metric, so the Grambank-only advantage is not overfitting. The LOLO columns use each source’s production configuration; the unseen-language column uses the per-source sweep winner. WALS runs at n=22 ; Grambank and Grambank+WALS run at n=24 .
Feature source
best max_features
LOLO ρ
R2
Grambank
0.3
0.716
0.511
Grambank+WALS
sqrt
0.673
0.447
WALS
log2
0.648
0.412
Appendix
Table 9: Feature-source comparison on the 22 languages shared with WALS (Experiment E1), each source after its own max_features mini-sweep on identical LOLO folds. The order of Table 8 is preserved on identical languages.
Feature design
#feat
LOLO ρ
R2
Grambank
195
0.780
0.607
Grambank+WALS, real block
387
0.705
0.492
Grambank+WALS, block scrambled (10 perms)
387
0.751±0.011
0.561±0.016
Appendix
Table 10: WALS-block permutation control (Experiment E1) under identical LOLO folds and a per-design max_features mini-sweep. The scrambled rows permute the WALS block across the 24 languages, preserving marginal levels and missingness; the mean and standard deviation are over ten permutations, every one of which scores above the real union.
Floor
n
language added
glm R2
full R2
tuned R2
tuned ρ
unseen-lang ρ
0.70
24
–
0.014
0.350
0.492
0.705
0.405
0.60
25
Telugu
0.003
0.507
0.529
0.731
0.411
0.50
26
Thai
−0.009
0.658
0.680
0.824
0.508
0.40
27
Indonesian
0.001
0.642
0.676
0.823
0.514
0.00
28
Central Kurdish
0.033
0.205
0.476
0.700
−0.012
Appendix
Table 11: Coverage-floor sweep of the 0.7 Grambank threshold that defines ATLAS-24. Every tier is refit at each floor under the production configuration, and each lowering admits the listed language, reaching all 28 ATLAS languages with a Grambank profile at floor 0.00 . The remaining ten of the 38 ATLAS languages, among them Spanish and German, are absent from Grambank and cannot enter at any floor. R2 is LOLO true R2 ; the final two columns are the tuned RF’s LOLO ρ and pure-unseen-language ρ , whose ATLAS-24 values are the headline 0.705 and 0.405 .
Model (same RF, same folds)
#feat
LOLO ρ
R2
metadata RF (script/family/genus/geo/wiki)
5
0.581
0.264
metadata RF, own-sweep tuned
5
0.620
0.375
typology RF (Grambank+WALS)
387
0.705
0.492
typology RF, budget-matched to metadata
5
0.597
–
Appendix
Table 12: Matched-capacity and matched-tuning controls from Experiment E1. A matched-capacity non-typological RF recovers ρ≈0.58 at the shared configuration and 0.620 under its own hyperparameter sweep, well above the single-feature ridges. Typology’s full-capacity increment is Δρ=0.085 under matched tuning and disappears at a matched five-feature budget, 0.597 against 0.620 , which locates the advantage in feature breadth.
Figure 1: Label-permutation null versus observed LOLO scores for the tuned random forest on Grambank+WALS features (ATLAS-24, B=200 replicates, seed 42). Left: Pearson ρ ; right: R2 . Bars show the fold-respecting permutation null, the dashed line its mean, and the solid line the observed score. The observed ρ exceeds all 200 null replicates ( p=0.005 ).
Factor
Mantel r
Mantel p
MRM coef
MRM p
typological
0.266
0.015
0.249
0.281
script
0.334
0.017
0.149
0.035
genealogical
0.226
0.051
0.098
0.693
geographic
0.169
0.078
0.029
0.803
Appendix
Table 13: Symmetric distance regressions on the 24×24 upper triangle, the cross-check within Experiment E3. Script is a strong marginal correlate but provides no incremental directed power, since the script-only directed model is near chance.
Axis
Group
grouped ρ
matched LOLO ρ
LOSO
Arabic
0.459
0.683
LOSO
Cyrillic
0.935
0.749
LOSO
Devanagari
0.912
0.622
LOSO
Greek
0.935
0.572
LOSO
Han
0.730
0.819
LOSO
Hangul
0.793
0.643
Appendix
Table 14: Composition-matched control for the grouped holdouts of Experiment E3. Grouped ρ retrains with every pair touching the group held out; matched LOLO scores the pooled LOLO predictions on exactly that group’s held-out pairs. Macro rows are unweighted means over groups; weighted rows average by held-out pair count.
Exp.
RQ
Quantity
Estimate
95% CI
E1
RQ1
LOLO ρ ; permutation p=0.005
0.705
[0.51,0.82]
E1
RQ1
LOLO ρ ; language-level bootstrap 95% CI
0.705
[0.50,0.85]
E1
RQ1
pure-unseen-language ρ (Grambank+WALS)
0.405
–
E3
RQ1
LOSO / LOFO macro- ρ
0.776 / 0.729
[0.67,0.88] / [0.59,0.83]
E4
RQ1
top-25 importance Jaccard (random ≈0.033 )
0.073
[0.069,0.078]
E5
RQ2
residual gap, Marathi − English (bootstrap over targets); p<0.001
+0.236
[0.19,0.28]
Appendix
Table 15: Inferential statistics, with each row labeled by the experiment and research question it addresses. The permutation null uses B=200 , and the cluster and paired bootstraps use B=2000 . The final row’s interval is the permuted-label null’s central 95% range, not a bootstrap interval.
Baseline
LOLO ρ
true R2
character n -gram overlap
0.225
0.044
subword tokenizer overlap
0.222
0.044
lang2vec featural cosine
0.148
0.012
word vocabulary overlap
0.142
0.010
URIEL six-distance ridge
0.095
−0.019
tuned RF (typology, reference)
0.705
0.492
Appendix
Table 16: Additional distance baselines under the leave-one-language-out protocol of Table 1 , on ATLAS-24. Every tier covers all 24 languages. The reference row repeats the tuned typology random forest from the main text.
Figure 2: Digitized ATLAS-38 ground-truth transfer matrix (Figure C.2 of the ATLAS paper). This axis order, family coloring, and BTS color scale are shared by all model and difference panels in this section.
Figure 3: LOLO predictions of the similarity tier (ridge on typological similarity alone). Held-out Pearson ρ=0.037 , MAE =0.277 BTS.
Figure 4: LOLO predictions of the GLM tier (ridge on similarity, script, genealogical, and geographic distances). Held-out Pearson ρ=0.162 , MAE =0.277 BTS.
Figure 5: LOLO predictions of the full random forest (all typological features, default hyperparameters). Held-out Pearson ρ=0.615 , MAE =0.203 BTS.
Figure 6: LOLO predictions of the tuned random forest (the production model of the main experiments; max_features =0.5 ). Held-out Pearson ρ=0.705 , MAE =0.185 BTS – this is the ρ=0.705 of the main ATLAS-24 result.
Figure 7: Ground truth minus the LOLO predictions of the similarity tier, on the shared symmetric difference scale. MAE 0.277 BTS.
Figure 8: Ground truth minus the LOLO predictions of the GLM tier. MAE 0.277 BTS.
Figure 9: Ground truth minus the LOLO predictions of the full random forest. MAE 0.203 BTS.
Figure 10: Ground truth minus the LOLO predictions of the tuned random forest. MAE 0.185 BTS. Note the shrinkage of the error relative to the weaker tiers: held-out error concentrates near the scale’s neutral midpoint.
Nuisance descriptor
std. weight
mean ∣ contribution ∣ (BTS)
source resource ( log )
+0.087
0.070
script difference
−0.077
0.071
genealogical distance
−0.034
0.027
geographic distance
−0.018
0.016
target resource ( log )
+0.002
0.002
Appendix
Table 17: Terms of the bias layer B^ , the ridge fit of BTS on five non-typological nuisance descriptors over the 552 ATLAS-24 directed pairs. Weights are on the standardized feature scale, and the mean absolute contribution is each feature’s contribution to B^ averaged over pairs. The layer explains R2=0.12 of BTS in-sample at correlation 0.35 , but approximately zero under leave-one-language-out cross-validation. Its weights are therefore descriptive of ATLAS-24, not a stable law. Its mean equals the BTS mean of −0.31 , so the residual T=BTS−B^ is mean-zero. The source-resource and script terms dominate, while the target-resource term is negligible, so the bias is driven by source-side resource and script sharing.
Bias features
R2(B,y)
∥ΔT∥ (%)
LOLO ρ
rank ρ (full)
rank ρ (truth)
top
k=1 (1 feature)
geo
0.013
34.8
0.719
0.989
0.984
English
res s
0.060
26.0
0.718
0.990
0.987
English
scr
0.051
28.0
0.717
0.979
0.984
English
gene
0.023
33.1
0.712
0.992
0.988
English
res t
0.000
36.8
0.705
0.989
0.990
English
Appendix
Table 18: Exhaustive feature-subset ablation of the ATLAS-24 bias layer B , grouped by subset size k . Each row refits B on that subset of the five nuisance descriptors and propagates to the debias residual T=BTS−B . Abbreviations: scr=script difference, gene=genealogical distance, geo=geographic distance, res s /res t = log1p Wikipedia resource of source/target. R2(B,y) is the bias-explained variance of BTS, and ∥ΔT∥ is the relative shift of T from the full-five residual. LOLO ρ is the leave-one-language-out Pearson of the debiased BTS predictor, an RF on T with the bias B added back. The production biased baseline is ρ=0.705 , and the no-bias-layer row refits within this pipeline, landing at 0.701 . Rank ρ is the Spearman correlation of the per-source pretraining ranking against the full-model ranking and against empirical BTS ground truth. The top source is English under every configuration; = marks a flip, and none is observed.
Rank
empirical
debiased
centrality
OOD-generalizer
source
score
source
score
source
score
source
score
1
English
0.022
Marathi
0.361
Ukrainian
0.717
English
-0.011
2
Modern Hebrew
-0.074
Filipino
0.355
Portuguese
0.700
Modern Hebrew
-0.065
3
French
-0.109
Modern Hebrew
0.328
Serbian-Croatian-Bosnian
0.679
Filipino
-0.114
4
Filipino
-0.113
Swahili
0.235
Catalan
0.678
Vietnamese
-0.114
5
Standard Arabic
-0.113
Standard Arabic
0.198
Italian
0.677
Standard Arabic
-0.120
Appendix
Table 19: Full 24-language ranking of the ATLAS-24 candidate set under each operationalization of Experiment E5. The candidate set is used as both sources and targets. The columns are the empirical raw-BTS mean, the debiased residual T=BTS−B^ , the mean typological similarity to the targets, and the OOD-generalizer mean BTS to cross-family targets. Scores are on each column’s own scale and are not comparable across columns; only the within-column order is meaningful. The on-scale reweighted random forest of Table 4 is omitted here because its full 24-language ranking tracks the empirical one at Spearman 0.974 , so its top sources are summarized in the main text only. The bold top row shows that the four methods disagree on the best source, which is the RQ2 result.
Rank
empirical
bias-removed
OOD-generalizer
source
count
source
count
source
count
1
English
8/23
Marathi
23/23
Modern Hebrew
3/20
2
French
5/23
Filipino
23/23
English
2/9
3
Modern Hebrew
3/23
Modern Hebrew
23/23
French
1/9
4
Portuguese
3/23
Standard Arabic
23/23
Modern Greek
1/9
5
Catalan
3/23
Western Farsi
23/23
Filipino
0/22
Appendix
Table 20: Full 24-language ranking of the ATLAS-24 candidate set by positive-transfer coverage: the number of target languages with a positive transfer score. The columns are the empirical raw BTS out of 23 targets, the bias-removed residual T=BTS−B^ out of 23 targets, and the OOD-generalizer raw BTS to cross-family targets, out of that source’s cross-family target count. Count ties are broken by the corresponding mean score in Table 19 . Counts are not comparable across columns; only the within-column order is meaningful.
Rank
empirical
bias-removed
OOD-generalizer
source
count
source
count
source
count
1
English
14/24
Marathi
10/24
Modern Hebrew
14/24
2
Modern Hebrew
6/24
Filipino
8/24
Swahili
10/24
3
Western Farsi
5/24
Modern Hebrew
6/24
English
9/24
4
French
4/24
Swahili
0/24
Vietnamese
9/24
5
Filipino
4/24
Standard Arabic
0/24
Standard Arabic
9/24
Appendix
Table 21: Full 24-language ranking of the ATLAS-24 candidate set by best-source coverage: the number of target languages for which the source attains the highest transfer score. The columns are the empirical raw BTS, the bias-removed residual T=BTS−B^ , and the raw BTS restricted to cross-family sources, the best outside-family option per target. A tie at a target’s maximum counts for every co-best source, so column counts sum to more than 24: empirically, 11 of 24 targets have tied maxima, up to 9-way. Count ties between sources are broken by the corresponding mean score in Table 19 . Only the within-column order is meaningful.
Typological features are widely used in multilingual NLP, and the prediction of such features holds downstream utility. However, existing methods to predict missing values lack interpretable justifications for predictions, while their performance across resource levels and feature types remains underexplored. Given LLMs' abilities in meta-linguistic reasoning and in providing rationales, we investigate LLMs' performance in typological feature prediction via an in-context learning approach with linguistic data from URIEL+ and Glottolog. We find that zero-shot prompting is insufficient, but when given phylogenetic and geographic neighbour evidence, LLMs substantially outperform all baselines without disadvantaging low-resource languages. We further find that most LLM rationales are consistent with the provided evidence, offering a step toward explainable typological feature prediction.
Qianwen Wang, York Hay Ng, Aditya Khan +1
University of Toronto, Canada · Columbia · Ontario Tech University, Canada
Cross-lingual transfer is central to low-resource machine translation, but its behavior within closely related language families remains insufficiently characterized. We study transfer among five Turkic languages; Turkish, Azerbaijani, Uzbek, Kazakh, and Kyrgyz; using pairwise transfer matrices. In this setting, each model is fine-tuned with one transfer source and evaluated on a different transfer target while the translation target remains the same. Across mT5 experiments, we find that transfer is strongest between closely related Turkic pairs, especially Turkish-Azerbaijani and Kazakh-Kyrgyz. We also show that transfer direction matters, and that the same transfer source-transfer target pair can behave differently when the translation target changes. Latinization improves BLEU and chrF in several script-mismatched settings, but its effect is not uniform across metrics. Additional analyses show that transfer sources are mostly stable across different datasets and model settings.
Omer Burak Cinar, Mehmet Mert Dalkilic, Cagri Toraman
Middle East Technical University · Computer Engineering Department
Cross-lingual transfer has become a central paradigm for extending natural language processing (NLP) technologies to low-resource languages. By leveraging supervision from high-resource languages, multilingual language models can achieve strong task performance with little or no labeled target-language data. However, it remains unclear to what extent cross-lingual transfer can substitute for language-specific efforts. In this paper, we synthesize prior research findings and data collection results on Luxembourgish, which, despite its typological proximity to high-resource languages and its presence in a multilingual context, remains insufficiently represented in modern NLP technologies. Across findings, we observe a fundamental interdependence between cross-lingual transfer and language-specific efforts. Cross-lingual transfer can substantially improve target-language performance, but its success depends critically on the availability of sufficiently high-quality, task-aligned target-language data. At the same time, such resources, particularly in low-resource settings, are typically too limited in scale to drive strong performance on their own. Instead, such resources reach their full potential only when leveraged within a cross-lingual framework. We therefore argue that cross-lingual transfer and language-specific efforts should not be viewed as competing alternatives. Instead, they function as complementary components of a sustainable low-resource NLP pipeline. Based on these insights, we provide practical guidelines for integrating and balancing cross-lingual transfer with language-specific development in sustainable low-resource NLP pipelines.
Fred Philippy, Siwen Guo, Jacques Klein +1
1SnT, University of Luxembourg, Luxembourg · 2Luxembourg Institute of Science and Technology, Luxembourg