LMSpell: Spell Correction with Pre-Trained Language Models
Organizations: Department of Computer Science and Engineering, University of Moratuwa, Katubedda, 10400, Sri Lanka · School of Mathematical and Computational Sciences, Massey University, Auckland, 102904, New Zealand
Abstract
Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of PLMs for spell correction across multiple languages, including low-resource languages. We show that even relatively small PLMs such as the 270M-parameter Gemma 3 and mBART50, when fine-tuned on a dataset of only 5k sentences, can outperform rule-based spell correctors, highlighting a practical pathway for building effective spell correction systems with limited data. We also present a case study with Sinhala to shed light on the plight of spell correction for LRLs.
Figures & tables
| Paper | Language | Architectures Compared | ||
| EO | DO | ED | ||
| 9160935 | English | ✓ | ✓ | ✓ |
| liu2024chinesespellingcorrectionrephrasing | Chinese | ✓ | ✓ | ✗ |
| martynov-etal-2024-methodology | Russian, English | ✓ | ✗ | ✓ |
| su2024ucsc | Chinese | ✓ | ✗ | ✓ |
| jiang-etal-2024-chinese | Chinese | ✓ | ✓ | ✗ |
| Paper | Language | Models Used |
| cheng-etal-2020-spellgcn | Chinese (5) | BERT |
| zhang-etal-2020-spelling | Chinese (5) | BERT |
| 9160935 | English (5) | GloVe, fastText, ELMo, GPT, GPT-2, BERT, RoBERTa, XLM-RoBERTa, BART, T5, XLNet |
| hu2020misspelling | English (5) | BERT |
| jayanthi-etal-2020-neuspell | Sinhala (2) | BERT |
| liu-etal-2021-plome | Chinese (5) | BERT |
| Model | Architecture | Params | # Langs |
| mT5 xue-etal-2021-mt5 | ED | 580M | 101 |
| mBART50 tang2020multilingual | ED | 680M | 50 |
| XLM-RoBERTa conneau-etal-2020-unsupervised | EO | 550M | 100 |
| Gemma 2 Instruct team2024gemma | DO | 9B | 1 |
| Gemma 3 1B Instruct team2025gemma | DO | 1B | 35 |
| Gemma 3 270m Instruct team2025gemma | DO | 270m | 35 |
| Language | Dataset | Number of Sentences | Family | Resource Level | PLMs | ||
| Train | Val. | Test | |||||
| Azerbaijani ( az ) | localdoc_2024 | 81,176 | 1,000 | 2,000 | Turkic | Low (Category 3) | mT5, XLM-R |
| Bulgarian ( bg ) | klouchek-batista-navarro-2024-bulgarian | 20,719 | 1,000 | 2,000 | Indo-European (Slavic) | Low (Category 3) | mT5, XLM-R |
| French ( fr ) | rasaboun_spellingcorrectionfrench ; fdemelo_french_news | 5,393 | 500 | 1,000 | Indo-European (Romance) | High (Category 5) | mT5, mBART50, XLM-R, Llama 3.1 |
| Hindi ( hi ) | etoori-etal-2018-automatic | 160,000 | 1,000 | 2,000 | Indo-European (Indo-Aryan) | Low (Category 3) | mT5, mBART50, XLM-R, Llama 3.1 |
| Korean ( ko ) | vitruv_err_spelling_kor | 77,000 | 1,000 | 2,000 | Koreanic | High (Category 4) | mT5, mBART50, XLM-R |
| Model | az (X, T) | bg , (X,T) | fr (X, T, B, L, L1) | hi (X, T, B, L, L1) | ko (X, T, B) | si (X, T, B) | vi (X, T) | id (X, T, B) | tk (X, T, B) | |||||||||
| Det | Corr | Det | Corr | Det | Corr | Det | Corr | Det | Corr | Det | Corr | Det | Corr | Det | Corr | Det | Corr | |
| XLM-R (X) | 24.72 | 16.41 | 11.10 | 8.24 | 57.02 | 51.84 | 74.64 | 72.58 | 76.16 | 65.24 | 12.79 | 11.52 | 25.51 | 25.80 | 47.75 | 43.83 | 21.68 | 17.28 |
| mT5 (T) | 8.31 | 0.91 | 5.92 | 3.13 | 7.99 | 2.36 | 76.50 | 64.28 | 25.81 | 0.93 | 5.55 | 3.31 | 12.13 | 5.90 | 46.34 | 41.31 | 3.24 | 1.27 |
| mBART (B) | 54.17 | 45.17 | 21.46 | 18.91 | 78.81 | 74.75 | 89.20 | 88.27 | 96.57 | 95.24 | 55.45 | 52.37 | 70.97 | 66.96 | 56.3 | 52.53 | 50.07 | 42.5 |
| Llama 3.1 (L) | 61.62 | 59.71 | 48.44 | 50.31 | 87.71 | 86.73 | 88.33 | 88.25 | 97.32 | 96.40 | 61.46 | 61.18 | 70.35 | 67.44 | 69.13 | 67.89 | 70.38 | 68.06 |
| Llama 3.2 1B (L1) | 36.81 | 36.22 | 20.78 | 23.33 | 77.43 | 75.00 | 68.20 | 68.26 | 88.18 | 87.28 | 54.93 | 55.16 | 49.80 | 45.50 | 53.92 | 51.48 | 41.35 | 38.81 |
| Lang | Training Data Set Size Model | 5000 | 51071 | 127677 | 255353 | 510706 | |||||
| D-F1 | C-F0.5 | D-F1 | C-F0.5 | D-F1 | C-F0.5 | D-F1 | C-F0.5 | D-F1 | C-F0.5 | ||
| si | XLMR | 12.79 | 11.52 | 24.30 | 22.55 | 46.78 | 45.24 | 52.14 | 50.98 | 64.07 | 63.45 |
| sinBert | 8.10 | 6.54 | 21.64 | 20.09 | 29.73 | 27.60 | 30.71 | 28.24 | 49.70 | 48.23 | |
| mT5 | 5.55 | 3.31 | 78.82 | 77.56 | 78.93 | 78.25 | 76.25 | 75.64 | 77.07 | 76.45 | |
| mBART50 | 55.45 | 52.37 | 68.22 | 66.62 | 73.27 | 72.37 | 71.90 | 71.27 | 74.86 | 74.40 | |
| Llama 3.1 8B | 61.46 | 61.18 | 69.87 | 69.63 | 72.58 | 72.87 | 74.59 | 75.38 | 75.07 | 75.90 | |
| Exp | Llama 3.1 | Gemma 2 | ||
| D-F1 | C-F0.5 | D-F1 | C-F0.5 | |
| ZS (B) | 51.13 | 38.58 | 4.17 | 4.06 |
| FS (B) | 48.73 | 49.73 | 14.20 | 13.03 |
| ZS (F) | 75.07 | 75.90 | 81.93 | 82.87 |
| FS (F) | 79.66 | 80.67 | 80.69 | 82.04 |
| RAG 1 | - | - | 79.32 | 80.87 |
| Domain | D-F1 | C-F0.5 |
| Government | 34.50 | 33.60 |
| Newspaper | 57.58 | 58.33 |
| Magazine | 44.23 | 44.23 |
| Socialmedia | 31.15 | 31.71 |
| Wikipedia | 36.36 | 36.36 |
| Model | Original Set | Synthetic Set | ||
| D-F1 | C-F0.5 | D-F1 | C-F0.5 | |
| mT5 | 5.55 | 3.31 | 1.68 | 0.76 |
| mBART50 | 55.45 | 52.37 | 49.61 | 47.30 |
| Gemma 2 | 62.52 | 62.31 | 66.89 | 69.12 |
| Llama 3.1 | 61.46 | 61.18 | 66.73 | 68.41 |
| Train Set | Test Set | Accuracy | Precision | Recall | F1 |
| Original | Original | 0.81 | 0.87 | 0.81 | 0.78 |
| Corrected | Original | 0.92 | 0.93 | 0.92 | 0.92 |
| Corrected | Corrected | 0.93 | 0.94 | 0.93 | 0.93 |
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
| Parameter | mT5 | mBART50 | XLM-R | Sinbert |
| Batch Size | 16 | 16 | 16 | 32 |
| Mixed Precision | bf16 | fp16 | fp16 | fp16 |
| Zero Stage | 2 | 2 | 2 | 2 |
| Max Seq Length (Train) | 128 | 128 | 128 | 128 |
| Patience | 5 | 5 | 5 | 5 |
| ZWJ Fix | Yes | Yes | Yes | No |
| Parameter | Gemma 2 9b | Gemma 3 1b | Gemma 3 270m | Llama 3.1 8b | Llama 3.2 1b |
| Batch Size | 4 | 8 | 4 | 4 | 8 |
| ZWJ Fix | No | No | No | No | No |
| Grad. Acc. Steps | 2 | 1 | 1 | 2 | 1 |
| Intial Lr. Rate | 1e-5 | 1e-5 | 1e-5 | 1e-5 | 1e-5 |
| R | 8 | 16 | None 5 5 5 Full Model is finetuned instead of using LORA | 8 | 16 |
| Flash Attention | Yes | Yes | No | Yes | Yes |
| Model | GPU Type | Time (hours) |
| sinBert | NVIDIA T4 (x4) | 8 |
| XLM-R | NVIDIA T4 (x4) | 22 |
| mT5 | NVIDIA T4 (x4) | 38 |
| mBART | NVIDIA T4 (x4) | 22 |
| Gemma 2 | NVIDIA L4 (x1) | 70 |
| Gemma 3 270M | NVIDIA T4 (x1) | 12 |
| Model | GPU Type | Time (hours) |
| sinBert | NVIDIA T4 (x2) | 0.2 |
| XLM-R | NVIDIA T4 (x2) | 0.5 |
| mT5 | NVIDIA T4 (x2) | 0.5 |
| mBART | NVIDIA T4 (x2) | 0.5 |
| Gemma 2 | NVIDIA T4 (x1) | 4 |
| Gemma 3 1B | NVIDIA T4 (x1) | 1 |