Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-resource languages, but in the absence of target language benchmarks, it is unclear which auxiliary language selection strategy leads to the best transfer. We introduce RunyaNER, the first publicly available NER benchmark for the East African language Runyankore, and use it to investigate the choice of which languages to use for transfer. Created with a semi-automated pipeline and fully manually verified, RunyaNER contains over 237k annotated words across 30k sentences. We benchmark pretrained models on RunyaNER, establishing that our dataset is of sufficient quality and size to produce effective Runyankore NER models. We then use RunyaNER to investigate auxiliary language selection in cross-lingual zero-shot and multilingual fine-tuning settings. Our experiments show that while transfer performance is highly sensitive to auxiliary language selection, embedding-based measures computed from labelled training spans correlate more strongly with downstream transfer performance than traditional linguistic features based on metadata or typology. By releasing RunyaNER and providing a systematic analysis of auxiliary language selection strategies, this work contributes both a new benchmark resource and practical insights for multilingual transfer in low-resource settings.
Figures & tables
Split
Sentences
Tokens
Named entity spans
Train
15,001
118,622
2,720
Dev
7,498
59,028
1,364
Test
7,508
59,426
1,427
Total
30,007
237,076
5,511
Table 1: Summary statistics for RunyaNER.
DATE
LOC
ORG
PER
Overall
Model
P
R
F1
P
R
F1
P
R
F1
P
R
F1
P
R
F1
Afro-XLMR
0.760
0.790
0.770
0.870
0.890
0.880
0.740
0.720
0.730
0.850
0.820
0.830
0.818
0.833
0.826
XLM-R
0.750
0.710
0.730
0.870
0.880
0.870
0.780
0.670
0.720
0.830
0.820
0.820
0.823
0.798
0.810
mBERT
0.720
0.720
0.720
0.850
0.880
0.870
0.730
0.680
0.710
0.890
0.840
0.860
0.804
0.803
0.803
Table 2: Monolingual fine-tuning results on RunyaNER (Baseline). Precision (P), recall (R), and F 1 -score (F1) are reported for each entity type and overall.
Code
Language
Subgroup
Zero-shot F 1
Multilingual F 1
Afro-XLMR
mBERT
XLM-R
Mean
Afro-XLMR
mBERT
XLM-R
Mean
kin
Kinyarwanda
Bantu (Great Lakes)
0.619
0.576
0.557
0.584
0.820
0.813
0.814
0.816
lug
Luganda
Bantu (Great Lakes)
0.622
0.521
0.495
0.546
0.831
0.817
0.806
0.818
nya
Chichewa/Nyanja
Bantu (Southeast)
0.576
0.462
0.447
0.495
0.814
0.808
0.818
0.813
swa
Kiswahili
Bantu (Swahili)
0.584
0.500
0.341
0.475
0.827
0.812
0.810
0.816
wol
Wolof
Senegambian
0.464
0.481
0.450
0.465
0.821
0.807
0.813
0.814
Table 3: Zero-shot cross-lingual and multilingual (bilingual) fine-tuning F 1 results on the RunyaNER test set for all auxiliary languages, ordered by zero-shot mean F 1 .
mBERT
XLM-R
Afro-XLMR
Metric
0-shot
multi
0-shot
multi
0-shot
multi
LinguaMeta
0.24
0.34
0.34
0.18
0.56 ∗
0.25
URIEL
0.53 ∗
0.51 ∗
0.41
0.51 ∗
0.52 ∗
0.10
Cosine
0.595 ∗∗
0.571 ∗∗
0.638 ∗∗
0.634 ∗∗
0.488 ∗
0.239
SWD
0.635 ∗∗
0.561 ∗
0.635 ∗∗
0.337
0.466 ∗
0.340
Table 4: Spearman correlation ( ρ ) between similarity metrics and Runyankore NER F1 across n=20 auxiliary languages. Embedding-based ρ is the strongest layer. We boldface the highest ρ per metric type and underline the highest ρ overall. ∗ p<0.05 , ∗∗ p<0.01 (two-tailed).
Figure 1: Runyankore NER transfer performance by auxiliary language ordered according to LinguaMeta similarity relative to Runyankore. Languages are arranged from highest to lowest similarity. Solid lines indicate cross-lingual zero-shot transfer, while dashed lines indicate multilingual fine-tuning performance.
Figure 2: Runyankore NER transfer performance by auxiliary language ordered according to URIEL typological similarity relative to Runyankore.
Figure 3: Runyankore NER transfer performance by auxiliary language ordered according to prototype cosine similarity relative to Runyankore.
Figure 4: Runyankore NER transfer performance by auxiliary language ordered according to Sliced Wasserstein Distance (SWD) similarity relative to Runyankore. Lower SWD values indicate stronger representational similarity.
Similarity
Auxiliary
Zero-shot F 1
Multilingual F 1
metric
language(s)
Afro-XLMR
mBERT
XLM-R
Mean
Afro-XLMR
mBERT
XLM-R
Mean
Quartets
LinguaMeta
lug/kin/swa/luo
0.644
0.615
0.635
0.631
0.817
0.821
0.809
0.816
URIEL
tsn/nya/lug/swa
0.593
0.582
0.566
0.580
0.812
0.813
0.816
0.814
Cosine
lug/sna/kin/nya
0.651
0.619
0.653
0.641
0.825
0.815
0.814
0.818
SWD
kin/twi/luo/ibo
0.598
0.515
0.562
0.558
0.817
0.815
0.809
0.814
Table 5: Runyankore NER performance under different auxiliary language selection schemes. Top-ranked quartets are shown for each similarity scheme, alongside the most similar single auxiliary language per scheme. Bold indicates the highest value in each column.
Language is humanity's most consequential technology, yet for over a billion speakers across India's twenty-two constitutionally recognised languages, its digital layer remains structurally incomplete. Named Entity Recognition (NER), the foundational step in transforming raw text into machine-interpretable knowledge, has been studied exhaustively for English but remains largely unsolved across most Indic languages. This paper presents a rigorous comparative study of generative and encoder-based neural architectures for NER on all eleven languages of the Naamapadam benchmark. We evaluate five classic model families spanning sequence-to-sequence transformers and multilingual encoders; four decoder-only large language models (LLMs) fine-tuned with LoRA and 4-bit NF4 quantisation; and nine generative models in zero-to-5-shot inference. Under strict CoNLL span-level evaluation, encoder-based models (mBERT and XLM-R, both F1=0.675 on Hindi) substantially outperform every generative architecture in ten of eleven languages, with gaps of 7.5-40 percentage points against the strongest competitor (Gemma-2-2B: avg F1=0.427). The best few-shot result reaches only 28% of the encoder baseline. We identify three language clusters--encoder-dominant, partial-coverage, and failure-zone; and provide actionable deployment guidelines grounded in transfer learning and low-resource NLP principles.
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings.
Akriti Dhasmana, Aarohi Srivastava, David Chiang
Computer Science and Engineering University of Notre Dame Notre Dame, IN, USA
Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon. When fine-tuning these models for an unseen language, practitioners must choose a proxy language token, yet no principled method exists for this selection. We implemented an embedding initialization strategy where a language token is the average of embeddings from multiple typologically related languages already in the mod el. We evaluate this approach on Limbum-to-English translation using a parallel corpus of 8,837 sentence pairs from New Testament text and a bilingual dictionary. We compare models: NLLB-200 zero-shot (chrF2++ = 12.5), a Transformer trained from scratch (chrF2++ = 14.5), NLLB-200 fine-tuned with a Swahili proxy token (chrF2++ = 47.3), and NLLB-200 with our averaged embedding initialization (chrF2++ = 46.7). We find that the multi-language initialization achieves performance comparable to the best single-language proxy. Both NLLB-200 variants improve over the from-scratch baseline by over 32 chrF2++ points. These results show that multilingual transfer is the dominant factor in extremely low-resource Bantu translation while eliminating the need for heuristic proxy selection. However, all systems fail to preserve tonal diacritics, highlighting an open challenge. We make our dataset and code available to support further research.
Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu +3
Ruzivo Research Lab, African Leadership University · African Leadership University