cs.CLOct 4, 2026

AraYoungVoices: A Diverse L1/L2 Corpus of Arabic Child and Adolescent Speech

Authors: Shammur Absar Chowdhury, Zien Sheikh Ali, Houssam Eddine-Othman Lachemat, Hamdy Mubarak

Organizations: Qatar Computing Research Institute, Qatar

Abstract

State-of-the-art ASR systems primarily target native adult speech, leading to substantial performance gaps for children, adolescents, and L2 speakers. We introduce AraYoungVoices, a 151.72-hour Arabic read-speech corpus from 286 speakers aged 7--18, comprising AraKids (7--12) and AraTeens (13--18). The corpus includes 146 native Arabic (L1) and 140 second-language (L2) speakers, with native speakers spanning Egyptian, Gulf, Levantine, and North African dialectal backgrounds and L2 speakers representing diverse linguistic backgrounds across the Americas, Asia, Africa, and Europe. We benchmark four pretrained ASR models under zero-shot and fine-tuned settings using unseen-speaker-&\&-unseen-prompt (USUP) and unseen-speaker-&\&-seen-prompt (USSP) evaluations. Results show that L2 speech remains substantially more challenging than L1 speech, with the largest errors observed mainly for younger L2 speakers. Age-specific fine-tuning improves the matched age group, while joint fine-tuning provides a stronger balance across populations. ASR hypotheses are also consistently closer to the standard reading prompt than to the verbatim transcription, particularly for L2 speech, suggesting partial normalization of reading deviations.

Figures & tables

Explore similar work

CardsList
  1. Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition

    Sep 28, 2026Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi +37Arabic Natural Language ProcessingArabic

  2. CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR

    Sep 28, 2026Bashar Talafha, Samar M. Magdy, Aisha Alansari +36DialectsMultilingual Automatic Speech Recognition