Every text classifier for an African language begins with a budgeting question: how many labelled examples are needed, and can labels from other African languages stand in for them? We answer both questions empirically for 28 language-task pairs, news topic classification in 16 languages (MasakhaNEWS) and tweet sentiment in 12 languages (AfriSenti), using a character n-gram linear model that trains in seconds on two CPU cores with no pretrained weights and no accelerator. Monolingual learning curves at budgets from 25 to several thousand labels show that topic classification reaches 90% of its full-data macro-F1 with about 400 labels in the median language, while sentiment is still improving at the full training size in 11 of 12 languages and needs thousands of labels. Pooling the full training data of the other languages in the benchmark is worth a great deal at small budgets and nothing at large ones: at 25 target labels it adds 0.20 macro-F1 on average for news (up to 0.43 for Lingala) and 0.08 for sentiment, the gain decays to zero by 800 labels, and at full size pooling hurts in 9 of 16 and 8 of 12 languages. Twenty-five target labels plus pooled data match what 100 to 400 monolingual labels achieve for most news languages. A complete zero-shot transfer matrix shows that transfer without any target labels recovers a median of only 13% (news) and 4% (sentiment) of the gap between a majority-class predictor and the in-language model, with the exceptions explained by shared script (Amharic and Tigrinya), shared lexicon (English and Nigerian Pidgin, the Arabic dialects), or a shared label prior rather than by language family. We release code that regenerates every number from the public benchmark files and translate the results into concrete annotation guidance for teams building African-language classifiers without GPUs.
Figures & tables
Figure 1: MasakhaNEWS learning curves. Macro-F1 on the official test split against the number of labelled target-language training examples (log scale), for a model trained on those examples alone (monolingual) and for one trained on those examples plus the full training sets of the other 15 languages (pooled). Bars are ±1 standard deviation over three subsamples. The dotted line is the pooled model with zero target examples.
Figure 2: AfriSenti learning curves, same layout as Figure 1 . Sentiment is far from saturated at the full training size in every language except Amharic, whose curve is dominated by a train-test label shift (52% neutral in training, 67% negative in test).
full mono
100 labels
400 labels
full
n95
language
n
mono
zero
m
p
m
p
pooled
raw
fit
MasakhaNEWS
Amharic
1311
0.923
0.727
0.845
0.836
0.892
0.878
0.888
400
181
Tigrinya
947
0.666
0.430
0.366
0.516
0.552
0.630
0.711
947
866
Somali
1021
0.707
0.312
0.402
0.521
0.623
0.686
0.726
800
685
Oromo
1128
0.812
0.520
0.532
0.619
0.634
0.712
0.794
1128
1463
Table 1: Macro-F1 on the official test split. n is the full training size. zero is the pooled model with no target labels; m and p are the monolingual and pooled models at 100 and 400 target labels; full is the model trained on all target labels, alone or pooled. n95 is the smallest budget at which the monolingual model reaches 95% of its full-data score, measured (raw) and from a fitted power law (fit); a raw value equal to n means the curve had not flattened.
Figure 3: Gain from pooling, pooled minus monolingual macro-F1, as a function of the target-language budget. Thin lines are languages; the thick line is the mean. Pooling is worth about 0.20 macro-F1 at 25 news labels and 0.08 at 25 sentiment labels, and nothing by 800.
Figure 4: Zero-shot transfer. A model trained on the row language’s full training set is tested on the column language. Cells show the fraction of the gap between a majority-class predictor and the column language’s own full-data model that the transferred model recovers, so 0 means no better than guessing the most common label and 1 means as good as in-language training. Languages are ordered by script and family: Ge’ez (amh, tir), Cushitic (som, orm), Nigerian languages (hau, yor, ibo, pcm), colonial languages (eng, fra), Bantu (swa, lin, lug, run, sna, xho); Arabic dialects (amh is Ge’ez), Nigerian, and Bantu for AfriSenti.
In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 African languages) and AfriSenti (12 languages plus two never seen in training), a pooled threshold meets the 90% target on average but covers Somali at 77.5%, Tigrinya at 83.7%, and the two unseen languages at 77.5% and 81.2%. Estimating one threshold per language brings every language to between 89.1% and 91.0% without retraining, and it shows how unequal the cost of the promise is: keeping it means sending 43% of Somali news and over 80% of Amharic and Xitsonga tweets to a person, against under 8% of Nigerian Pidgin news. One or two hundred labels per language are enough and the models train in minutes on one CPU core, so the fix is affordable: calibrate, report, and budget human review one language at a time.
African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance. The study is a systematic sample-size scaling study of natural language inference (NLI) on 16 African languages based on the AfriXNLI benchmark. Under controlled conditions, two multilingual transformer models with roughly 0.6B parameters XLM-R Large fine-tuned on XNLI and AfroXLM-R Large are tested on sample sizes of between 50 and 500 labeled examples and average their results across random subsampling runs. As opposed to the usual belief of monotonic increase with increased data, we find a strongly language sensitive and often non-monotonic scaling behavior. Some languages show early saturation or decrease in performance with sample size as well as high variance in low resource regimes. These results indicate that the volume of data is not enough to guarantee stable profits to African NLI, creating the necessity of language sensitive datasets creation and stronger multi-lingual modelling strategies.
Anuj Tiwari, Oluwapelumi Ogunremu, Terry Oko-odion +2
Noida Institute of Engineering and Technology1 · ML Collective1,2,3,4,5
How can language learning systems be developed for languages that lack sufficient training resources? This challenge is increasingly faced by developers across the African continent who aim to build AI systems capable of understanding and responding in local languages. To address this gap, we introduce AFRILANGDICT, a collection of 194.7K African language-English dictionary entries designed as seed resources for generating language-learning materials, enabling us to automatically construct large-scale, diverse, and verifiable student-tutor question-answer interactions suitable for training AI-assisted language tutors. Using AFRILANGDICT, we build AFRILANGEDU, a dataset of 78.9K multi-turn training examples for Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO). Using AFRILANGEDU, we train language tutoring models collectively referred to as AFRILANGTUTOR. We fine-tune two multilingual LLMs: Llama-3-8B-IT and Gemma-3-12B-IT on AFRILANGEDU across 10 African languages and evaluate their performance. Our results show that models trained on AFRILANGEDU consistently outperform their base counterparts, and combining SFT and DPO yields substantial improvements, with gains ranging from 1.8% to 15.5% under LLM-as-a-judge evaluations across four criteria. To facilitate further research on low-resource languages, all resources are available at https://huggingface.co/afrilang-edu.
Tadesse Destaw Belay, Shahriar Kabir Nahin, Israel Abebe Azime +6
Instituto Politécnico Nacional, Mexico · University of South Florida, FL, USA · Saarland University, Germany +2