cs.CLDec 5, 2025

LMSpell: Spell Correction with Pre-Trained Language Models

Authors: Akesh Gunathilake, Nadil Karunarathna, Tharusha Bandaranayake, Nisansa de Silva, Surangika Ranathunga, Nevidu Jayatilleke

Organizations: Department of Computer Science and Engineering, University of Moratuwa, Katubedda, 10400, Sri Lanka · School of Mathematical and Computational Sciences, Massey University, Auckland, 102904, New Zealand

Abstract

Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of PLMs for spell correction across multiple languages, including low-resource languages. We show that even relatively small PLMs such as the 270M-parameter Gemma 3 and mBART50, when fine-tuned on a dataset of only 5k sentences, can outperform rule-based spell correctors, highlighting a practical pathway for building effective spell correction systems with limited data. We also present a case study with Sinhala to shed light on the plight of spell correction for LRLs.

Figures & tables

Appendix figures & tables6 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Contextual Tamil Spelling and Grammar Correction Using Progressively Fine-Tuned Sequence-to-Sequence Transformers

    Sep 3, 2026Karthikeyan A, Jaya Nirmala S, Sangeetha Sivanesan +4Grammatical Error CorrectionLow-Resource Language Processing

  2. Edit-level Majority Voting Mitigates Over-Correction in LLM-based Grammatical Error Correction

    May 13, 2026Takumi Goto, Yusuke Sakai, Taro WatanabeSelf-Consistency DecodingGrammatical Error Correction