The ÌròyìnSpeech Text Corpus: 24,905 Curated Yorùbá Sentences for Speech and Language Technology
Organizations: Stanford University · Niger-Volta LTI · Mila / McGill University
Abstract
ÌròyìnSpeech is a 42-hour, 80-speaker Yorùbá read-speech corpus whose audio has been distributed by ELRA since 2024. This paper describes the release of its text component: 24,905 unique, hand-verified, tone-marked Yorùbá sentences (275,897 tokens; 15,687 types), curated in 2022 as recording prompts. Roughly 11,000 sentences were adapted from openly licensed news material; the remainder were written in-house to broaden coverage beyond the religious translation that dominates existing Yorùbá corpora. Every sentence was checked by hand for tone-mark accuracy, edited for read-aloud clarity and a neutral register, and localised so that non-Yorùbá personal and place names appear in Yorùbá form. Preparing the text for release surfaced systematic Unicode normalisation failures affecting more than 60% of lines (with precomposed and decomposed forms of the same letter co-occurring within single sentences) which we document and correct. The corpus supports diacritic restoration, grapheme-to-phoneme conversion, TTS front-end development and orthographic research, and serves as a validated prompt set for new recording.
Figures & tables
| Field | Value |
|---|---|
| Object name | ÌròyìnSpeech Text Corpus |
| Format | UTF-8 plain text (.txt), CSV, JSON Lines |
| Creation dates | 2022-01 to 2022-12 (curation); 2026 (cleaning and packaging) |
| Dataset creators | Kọ́lá Túbọ̀sún, Anuoluwapo Aremu, Tolúlọpẹ́ Ògúnrẹ̀mí, Iroro Orife, David Ifeoluwa Adelani |
| Language | Yorùbá (ISO 639-3: yor) |
| Licence | CC BY-NC 4.0 |
| Measure | Value |
|---|---|
| Rows in the 2022 working file | 25,001 |
| Unique sentences after cleaning | 24,905 |
| Exact duplicates removed | 96 |
| Word tokens | 275,897 |
| Word types | 15,687 |
| Hapax legomena | 8,445 |