cs.CLOct 1, 2026

Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yorùbá

Authors: Ahmad Samuel Gali, Shamsuddeen Hassan Muhammad

Organizations: University of Lagos · Bayero University Kano · Imperial College London

Abstract

Yorùbá is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Restoration (ADR) model fine-tuned from ByT5-small. We evaluate Yo-ByT5 alongside five publicly released Yorùbá ADR models and one open-weight large language model (LLM) on the YAD benchmark under a consistent protocol. Our results demonstrate that Yo-ByT5 matches the performance of the strongest existing model, mT5-base, with a DER of 10.14% and a CER of 3.48%. Furthermore, it exhibits superior text fidelity despite using approximately half the parameter count of mT5-base. We also release our training code and model outputs, as well as call for the development of a larger, purpose-built benchmark for Yorùbá diacritic restoration.

Figures & tables

Explore similar work

CardsList
  1. Koshur Diacritizer: A Byte-Level Sequence-to-Sequence Model for Kashmiri Diacritic Restoration

    Jun 14, 2026Haq Nawaz Malik, Nahfid Nissar, Faizan Iqbal

  2. Tone on a Budget: A Reference-Free Metric for Lexical Tone in Massively Multilingual Text-to-Speech

    Sep 13, 2026Moses Daudu, Adeola Enitan Bamidele, Honor-Jesus BezaleelSeed-Tts-Eval BenchmarkCharacter Error Rate

  3. Constrained CTC Decoding for Efficient Diacritic Restoration

    Jul 21, 2026Rufael Marew, Amr Keleg, Hanan AldarmakiGrapheme-To-Phoneme