cs.CLOct 5, 2026

Automatic Speech Recognition for Low-Resource Sinhala: A Critical Review of Methods, Challenges, and Future Directions

Authors: Chanuka Dinuwan, Sanath Jayasena, Buddhika Karunarathne

Organizations: MSc in Data Science and Artificial Intelligence Department of Computer Science & Engineering University of Moratuwa, Sri Lanka · University of Moratuwa, Sri Lanka

Abstract

Automatic speech recognition (ASR) for low-resource languages remains a major challenge. Sinhala, the primary language of Sri Lanka with about 16 million speakers, illustrates the difficulty: agglutinative morphology, a 54-phoneme inventory, subject-object-verb (SOV) syntax and scarce annotated speech data limit both conventional and modern ASR systems. This paper presents the first critical review of Sinhala ASR research, tracing its development from Hidden Markov Models (HMMs) through deep neural networks to self-supervised pre-trained models such as wav2vec 2.0, XLS-R, Whisper and Massively Multilingual Speech (MMS). We compare existing Sinhala systems with related low-resource ASR work on Tamil, Malayalam and Hindi in terms of architecture, training data, word error rate (WER) and robustness to real-world acoustic conditions, and we assess self-supervised and transfer learning as responses to scarce labeled data. We show that most reported WERs are not directly comparable because they differ in corpus, data split and scoring, and that the only controlled comparison in the literature attributes an 18.1% relative WER reduction to corpus correction alone. We also discuss context-aware ASR that draws on phonological, syntactic and semantic knowledge. We identify six research gaps: (1) the lack of large annotated corpora covering multiple dialects and acoustic conditions; (2) weak contextual modeling of Sinhala morphosyntax; (3) high WER in real-world conditions; (4) the absence of standardized benchmarks; (5) the lack of parameter-efficient fine-tuning studies; and (6) the absence of annotated code-switched Sinhala-English speech resources. We outline a research agenda to address these gaps, intended as a roadmap for researchers working on Sinhala and other morphologically rich languages.

Explore similar work

CardsList
  1. From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition

    Jul 7, 2026Lukmal Ilyas, Nevidu JayatillekeMultilingual Automatic Speech RecognitionCross-Lingual Transfer

  2. Hard to Be Heard: Phoneme-Level ASR Analysis of Phonologically Complex, Low-Resource Endangered Languages

    Apr 20, 2026V. S. D. S. Mahesh Akavarapu, Michael Daniel, Gerhard JägerGrapheme-To-PhonemeWav2Vec

  3. Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models

    Jul 19, 2026Ganapati Das, Dwipen Laskar, Hasin Afzal Ahmed +4Whisper-Large-V3Automatic Speech Recognition