Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.
Figures & tables
Undiacritised
Reading
Tone pattern
Meaning
igba
igbá
mid–high
calabash
ìgbà
low–low
time, period
ìgbá
low–high
garden egg
igba
mid–mid
two hundred
Table 1: One undiacritised Yorùbá form, igba , and four valid readings.
Figure 1: Yo-ByT5 corpus and fine-tuning lineage. The MENYO-20k train split and the Biblica Open Yorùbá Contemporary Bible 2017 form the training set; the model is initialised from google/byt5-small and fine-tuned on it. The YAD test benchmark, on which Yo-ByT5 is evaluated, is drawn from the MENYO-20k test split.
Model
Repository / provider
Architecture
Params
Yo-ByT5
lazymonster/yobyt5-restoration
ByT5, byte-level
299.6M
mT5-base
Davlan/mT5_base_yoruba_adr
mT5, subword
582.4M
omowe-T5 all-und
Davlan/omowe-t5-small-diacritizer-all-und-full
T5-small, subword
76.9M
omowe-T5 menyo
Davlan/omowe-t5-small-diacritizer-menyo
T5-small, subword
76.9M
ByT5-small menyo
Davlan/byt5-small-diacritizer-menyo
ByT5, byte-level
299.6M
mT5-small menyo
Davlan/mt5-small-diacritizer-menyo
mT5, subword
300.2M
Table 2: Evaluated models: Yo-ByT5, five publicly released Yorùbá diacritisers, and one open-weight LLM baseline.
Model
Params
CER
WER
DER
DER-tone
DER-und.
WDER
BLEU
ChrF
Yo-ByT5
299.6M
3.48
14.86
10.14
8.93
6.15
15.64
0.6841
0.8431
mT5-base
582.4M
4.38
14.47
10.28
9.27
7.62
13.92
0.7051
0.8506
omowe-T5 all-und
76.9M
13.36
20.33
14.76
14.12
13.11
16.68
0.7247
0.8295
omowe-T5 menyo
76.9M
35.18
43.51
20.46
19.56
21.24
19.94
0.5404
0.7562
gpt-oss:120b *
≈ 120B
10.31
36.78
28.42
24.06
27.65
34.15
0.3801
0.6468
ByT5-small menyo
299.6M
30.84
44.98
40.86
38.96
35.38
45.71
0.3661
0.5829
Table 3: Comparison on the YAD test benchmark (3,330 sentences), corrected YAD test set, base-word Needleman-Wunsch alignment. Beam decoding ( num_beams = 5). Sorted by DER; best value per column in bold. * gpt-oss:120b is run one-shot at temperature 0, not beam search.
Figure 2: Word error typology across the seven evaluated models (corrected YAD test, greedy decoding). Blue hues: the three mark classes. Orange hues: the three text-altering classes. The text-altering share is stated beside each bar.
Figure 3: Base-text preservation by gpt-oss:120b against input sentence length (YAD test). The share of outputs that keep the base text falls from 53.4% on the shortest sentences to 3.0% on the longest.