In the Global South, the lower-income countries of Africa, Asia, and Latin America where most of the world's languages are spoken, a deployed text classifier usually runs on ordinary CPUs, serves many languages with a single model, has few labeled examples in any of them, and relies on people to catch its mistakes. Such a system is only useful if it can promise how often it will be wrong: at most a fixed fraction of the labels it assigns on its own may be incorrect, and everything else must go to a person. Split conformal prediction delivers this promise through a single confidence threshold, normally estimated on validation data pooled across languages. We ask whether the promise reaches every language, and it does not. On MasakhaNEWS (16 African languages) and AfriSenti (12 languages plus two never seen in training), a pooled threshold meets the 90% target on average but covers Somali at 77.5%, Tigrinya at 83.7%, and the two unseen languages at 77.5% and 81.2%. Estimating one threshold per language brings every language to between 89.1% and 91.0% without retraining, and it shows how unequal the cost of the promise is: keeping it means sending 43% of Somali news and over 80% of Amharic and Xitsonga tweets to a person, against under 8% of Nigerian Pidgin news. One or two hundred labels per language are enough and the models train in minutes on one CPU core, so the fix is affordable: calibrate, report, and budget human review one language at a time.
Figures & tables
Figure 1: The deferral pipeline. One cheap model serves every language; a conformal threshold turns its probabilities into a prediction set, and only sets with exactly one label become automatic labels. The two policies differ only in where the threshold comes from.
Figure 2: Coverage (top) and deferral rate (bottom) per language at a 0.90 target (dashed), mean ± s.d. over 30 random splits. Trained languages are ordered by base accuracy; starred languages never appeared in training. A pooled threshold (red) over-covers well-resourced languages and under-covers hard ones (shaded region); a per-language threshold (blue) meets the target everywhere, at a very different deferral cost per language.