Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.
Figures & tables
Split
Sentences
Tokens
Named entity spans
Train
15,001
118,622
2,720
Dev
7,498
59,028
1,364
Test
7,508
59,426
1,427
Total
30,007
237,076
5,511
Table 1: Summary statistics for RunyaNER.
DATE
LOC
ORG
PER
Overall
Model
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Afro-XLMR
0.760
0.790
0.770
0.870
0.890
0.880
0.740
0.720
0.730
0.850
0.820
0.830
0.818
0.833
0.826
XLM-R
0.750
0.710
0.730
0.870
0.880
0.870
0.780
0.670
0.720
0.830
0.820
0.820
0.823
0.798
0.810
mBERT
0.720
0.720
0.720
0.850
0.880
0.870
0.730
0.680
0.710
0.890
0.840
0.860
0.804
0.803
0.803
Table 2: Monolingual fine-tuning results on RunyaNER (Baseline). Precision (P), recall (R), and F 1 -score (F1) are reported for each entity type and overall.
Code
Language
Subgroup
Zero-shot F 1
Multilingual F 1
Afro-XLMR
mBERT
XLM-R
Mean
Afro-XLMR
mBERT
XLM-R
Mean
kin
Kinyarwanda
Bantu (Great Lakes)
0.619
0.576
0.557
0.584
0.820
0.813
0.814
0.816
lug
Luganda
Bantu (Great Lakes)
0.622
0.521
0.495
0.546
0.831
0.817
0.806
0.818
nya
Chichewa/Nyanja
Bantu (Southeast)
0.576
0.462
0.447
0.495
0.814
0.808
0.818
0.813
swa
Kiswahili
Bantu (Swahili)
0.584
0.500
0.341
0.475
0.827
0.812
0.810
0.816
wol
Wolof
Senegambian
0.464
0.481
0.450
0.465
0.821
0.807
0.813
0.814
Table 3: Zero-shot cross-lingual and multilingual (bilingual) fine-tuning F 1 results on the RunyaNER test set for all auxiliary languages, ordered by zero-shot mean F 1 .
mBERT
XLM-R
Afro-XLMR
Metric
0-shot
multi
0-shot
multi
0-shot
multi
LinguaMeta
0.24
0.34
0.34
0.18
0.56 ∗
0.25
URIEL
0.53 ∗
0.51 ∗
0.41
0.51 ∗
0.52 ∗
0.10
Cosine
0.595 ∗∗
0.571 ∗∗
0.638 ∗∗
0.634 ∗∗
0.488 ∗
0.239
SWD
0.635 ∗∗
0.561 ∗
0.635 ∗∗
0.337
0.466 ∗
0.340
Table 4: Spearman correlation ( ρ ) between similarity metrics and Runyankore NER F1 across n=20 auxiliary languages. Embedding-based ρ is the strongest layer. We boldface the highest ρ per metric type and underline the highest ρ overall. ∗ p<0.05 , ∗∗ p<0.01 (two-tailed).
Figure 1: Runyankore NER transfer performance by auxiliary language ordered according to LinguaMeta similarity relative to Runyankore. Languages are arranged from highest to lowest similarity. Solid lines indicate cross-lingual zero-shot transfer, while dashed lines indicate multilingual fine-tuning performance.
Figure 2: Runyankore NER transfer performance by auxiliary language ordered according to URIEL typological similarity relative to Runyankore.
Figure 3: Runyankore NER transfer performance by auxiliary language ordered according to prototype cosine similarity relative to Runyankore.
Figure 4: Runyankore NER transfer performance by auxiliary language ordered according to Sliced Wasserstein Distance (SWD) similarity relative to Runyankore. Lower SWD values indicate stronger representational similarity.
Similarity
Auxiliary
Zero-shot F 1
Multilingual F 1
metric
language(s)
Afro-XLMR
mBERT
XLM-R
Mean
Afro-XLMR
mBERT
XLM-R
Mean
Quartets
LinguaMeta
lug/kin/swa/luo
0.644
0.615
0.635
0.631
0.817
0.821
0.809
0.816
URIEL
tsn/nya/lug/swa
0.593
0.582
0.566
0.580
0.812
0.813
0.816
0.814
Cosine
lug/sna/kin/nya
0.651
0.619
0.653
0.641
0.825
0.815
0.814
0.818
SWD
kin/twi/luo/ibo
0.598
0.515
0.562
0.558
0.817
0.815
0.809
0.814
Table 5: Runyankore NER performance under different auxiliary language selection schemes. Top-ranked quartets are shown for each similarity scheme, alongside the most similar single auxiliary language per scheme. Bold indicates the highest value in each column.