Organizations: Department of Computer Science and Engineering, University of Moratuwa, Katubedda, 10400, Sri Lanka · School of Mathematical and Computational Sciences, Massey University, Auckland, 102904, New Zealand
Spell correction is still a challenging problem for many languages, especially low-resource languages (LRLs). While pre-trained language models (PLMs) have been employed for spell correction, there has been no proper comparison across PLMs. We present the first empirical study on the effectiveness of the three types of PLMs for spell correction across multiple languages, including low-resource languages. We show that even relatively small PLMs such as the 270M-parameter Gemma 3 and mBART50, when fine-tuned on a dataset of only 5k sentences, can outperform rule-based spell correctors, highlighting a practical pathway for building effective spell correction systems with limited data. We also present a case study with Sinhala to shed light on the plight of spell correction for LRLs.
Figures & tables
Paper
Language
Architectures Compared
EO
DO
ED
9160935
English
✓
✓
✓
liu2024chinesespellingcorrectionrephrasing
Chinese
✓
✓
✗
martynov-etal-2024-methodology
Russian, English
✓
✗
✓
su2024ucsc
Chinese
✓
✗
✓
jiang-etal-2024-chinese
Chinese
✓
✓
✗
Table 1: Empirical studies on PLMs for spell correction. EO - Encoder-Only, DO - Decoder-Only, and ED - Encoder-Decoder
Table 2: Research that used PLMs for spell correction. The ranathunga-de-silva-2022-languages language category is given in parentheses after the name of each language.
Model
Architecture
Params
# Langs
mT5 xue-etal-2021-mt5
ED
580M
101
mBART50 tang2020multilingual
ED
680M
50
XLM-RoBERTa conneau-etal-2020-unsupervised
EO
550M
100
Gemma 2 Instruct team2024gemma
DO
9B
1
Gemma 3 1B Instruct team2025gemma
DO
1B
35
Gemma 3 270m Instruct team2025gemma
DO
270m
35
Table 3: PLMs used in the experiments. EO - Encoder-Only, DO - Decoder-Only, and ED - Encoder-Decoder
Table 4: Details of the languages. Resource level is according to ranathunga-de-silva-2022-languages ’s language categorization. The final column shows the PLMs that were pre-trained on data from each language
Model
az (X, T)
bg , (X,T)
fr (X, T, B, L, L1)
hi (X, T, B, L, L1)
ko (X, T, B)
si (X, T, B)
vi (X, T)
id (X, T, B)
tk (X, T, B)
Det
Corr
Det
Corr
Det
Corr
Det
Corr
Det
Corr
Det
Corr
Det
Corr
Det
Corr
Det
Corr
XLM-R (X)
24.72
16.41
11.10
8.24
57.02
51.84
74.64
72.58
76.16
65.24
12.79
11.52
25.51
25.80
47.75
43.83
21.68
17.28
mT5 (T)
8.31
0.91
5.92
3.13
7.99
2.36
76.50
64.28
25.81
0.93
5.55
3.31
12.13
5.90
46.34
41.31
3.24
1.27
mBART (B)
54.17
45.17
21.46
18.91
78.81
74.75
89.20
88.27
96.57
95.24
55.45
52.37
70.97
66.96
56.3
52.53
50.07
42.5
Llama 3.1 (L)
61.62
59.71
48.44
50.31
87.71
86.73
88.33
88.25
97.32
96.40
61.46
61.18
70.35
67.44
69.13
67.89
70.38
68.06
Llama 3.2 1B (L1)
36.81
36.22
20.78
23.33
77.43
75.00
68.20
68.26
88.18
87.28
54.93
55.16
49.80
45.50
53.92
51.48
41.35
38.81
Table 5: Performance of PLMs fine-tuned with 5k sentences from each language. For VI, the full dataset of 4,500 sentences was used). For each language, PLMs that include that language are indicated within brackets. Note that although Llama 3.2 and Gemma 3 are said to have been pre-trained on a multilingual dataset, languages coverage details are not publicly available. For each column, Bold indicates the best performance, and underline indicates the second-best.
Lang
Training Data Set Size Model
5000
51071
127677
255353
510706
D-F1
C-F0.5
D-F1
C-F0.5
D-F1
C-F0.5
D-F1
C-F0.5
D-F1
C-F0.5
si
XLMR
12.79
11.52
24.30
22.55
46.78
45.24
52.14
50.98
64.07
63.45
sinBert
8.10
6.54
21.64
20.09
29.73
27.60
30.71
28.24
49.70
48.23
mT5
5.55
3.31
78.82
77.56
78.93
78.25
76.25
75.64
77.07
76.45
mBART50
55.45
52.37
68.22
66.62
73.27
72.37
71.90
71.27
74.86
74.40
Llama 3.1 8B
61.46
61.18
69.87
69.63
72.58
72.87
74.59
75.38
75.07
75.90
Table 6: Performance across different dataset sizes for Sinhala and Hindi. Each dataset size corresponds to approximately 1%, 10%, 25%, 50%, 100% of the Sinhala training dataset.
Exp
Llama 3.1
Gemma 2
D-F1
C-F0.5
D-F1
C-F0.5
ZS (B)
51.13
38.58
4.17
4.06
FS (B)
48.73
49.73
14.20
13.03
ZS (F)
75.07
75.90
81.93
82.87
FS (F)
79.66
80.67
80.69
82.04
RAG 1
-
-
79.32
80.87
Table 7: Zero-shot (ZS), Few-shot (FS) and RAG results for the un-finetuned (B) and fine-tuned (F) LLMs.
Domain
D-F1
C-F0.5
Government
34.50
33.60
Newspaper
57.58
58.33
Magazine
44.23
44.23
Socialmedia
31.15
31.71
Wikipedia
36.36
36.36
Table 8: Domain-specific results for Gemma-2-full
Model
Original Set
Synthetic Set
D-F1
C-F0.5
D-F1
C-F0.5
mT5
5.55
3.31
1.68
0.76
mBART50
55.45
52.37
49.61
47.30
Gemma 2
62.52
62.31
66.89
69.12
Llama 3.1
61.46
61.18
66.73
68.41
Table 9: Original test set evaluated with models finetuned with original and synthetic error data.
Train Set
Test Set
Accuracy
Precision
Recall
F1
Original
Original
0.81
0.87
0.81
0.78
Corrected
Original
0.92
0.93
0.92
0.92
Corrected
Corrected
0.93
0.94
0.93
0.93
Table 10: Finetuning Performance of SinLlama aravinda2025sinllama on Sinhala text classification using datasets spell corrected using Gemma-2-full .
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 1: Interface of LMSpell
Figure 2: Evaluation workflow during hallucination
Parameter
mT5
mBART50
XLM-R
Sinbert
Batch Size
16
16
16
32
Mixed Precision
bf16
fp16
fp16
fp16
Zero Stage
2
2
2
2
Max Seq Length (Train)
128
128
128
128
Patience
5
5
5
5
ZWJ Fix
Yes
Yes
Yes
No
Appendix
Table 12: Training parameters for Encoder-based and Encoder-Decoder models.
Parameter
Gemma 2 9b
Gemma 3 1b
Gemma 3 270m
Llama 3.1 8b
Llama 3.2 1b
Batch Size
4
8
4
4
8
ZWJ Fix
No
No
No
No
No
Grad. Acc. Steps
2
1
1
2
1
Intial Lr. Rate
1e-5
1e-5
1e-5
1e-5
1e-5
R
8
16
None 5 5 5 Full Model is finetuned instead of using LORA
Tamil spell and grammar correction is challenging because Tamil is an agglutinative low-resource language with rich verbal morphology, complex sandhi (phonetic transformation) rules at word boundaries, and a script of 247 distinct letters. Prior work targets word-level surface errors with rule-based methods, statistical n-gram models, Minimum Edit Distance, or hybrid pipelines with a transformer re-ranker; such methods cannot reliably handle contextual errors - subject-verb agreement, tense consistency, or cross-word sandhi - which require sentence-level understanding. We propose an end-to-end sequence-to-sequence formulation and fine-tune mT5-small and mBART-50 on a synthetic corpus of up to 657,720 noisy-clean Tamil sentence pairs spanning ten error categories. Both backbones follow the same four-stage progressive schedule, each stage targeting one weakness: surface noise (v2), contextual grammar (v3), single-site sandhi (v4), and multi-site cross-word sandhi (v5). On a 1,000-sentence balanced diagnostic set verified disjoint from all training data, our best model, mBART-50 v5, reaches 69.3% top-1 exact-match accuracy, with 87.5% on sandhi and 43.5% on subject-verb agreement. The schedule is what produces these gains: subject-verb accuracy rises from 1.0% to 52.5% once contextual pairs are introduced, and sandhi from 0% to 87.5% once multi-site sandhi pairs are. We additionally quantify a precision-recall trade-off this literature has not reported: sandhi recall is paid for monotonically in identity accuracy. Finally, Tamil-LLaMA-7B-Instruct reaches 19.0% zero-shot and 24.7% with three demonstrations against a 20.0% copy baseline, showing that a Tamil-adapted instruction model does not transfer to specialised sentence-level correction without task-specific supervision.
Karthikeyan A, Jaya Nirmala S, Sangeetha Sivanesan +4
National Institute of Technology, Tiruchirappalli · Tamil University, Thanjavur
Grammatical error correction using large language models often suffers from the over-correction issue. To mitigate this, we propose a training-free inference method that performs edit-level majority voting over multiple candidates generated by a single model, without requiring model modifications or additional training. Across nine benchmarks covering English, Czech, German, Ukrainian, Korean, Hindi, and Romanian, the proposed method outperforms both greedy and MBR decoding in most cases. Moreover, it yields stable correction quality regardless of the instruction prompts used. We release two repository supporting GEC datasets loading and LLM inference.
Automatic speech recognition (ASR) has improved substantially in recent years, yet performance remains limited for low-resource languages. Large language models (LLMs) have shown promise for improving ASR through generative error correction (GER), but their effectiveness in low-resource settings remains underexplored. In addition, it remains unclear to what extent data contamination influences the reported improvements in LLM-based GER. This study investigates LLM-based GER for low-resource Frisian. In addition to a public corpus, we construct and use a Frisian offline dataset with non-public texts for evaluation to control for potential data contamination. Results show that GER improves ASR performance in most settings, with the best GPT-5.1 results surpassing oracle WERs. Comparable gains on the offline dataset indicate that improvements reflect true correction ability. We further provide a detailed error analysis revealing model correction patterns.