SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0
Authors: Nevidu Jayatilleke, Nisansa de Silva, Uthpala Nimanthi, Gagani Kulathilaka, Azra Safrullah, Johan Sofalas
Organizations: Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka · Research Department, Informatics Institute of Technology, Sri Lanka
SiDiaC-v.2.0 is the largest comprehensive Sinhala Diachronic Corpus to date, covering a period from 1800 CE to 1955 CE in terms of publication dates, and a historical span from the 5th to the 20th century CE in terms of written dates. The corpus consists of 229k words across 185 literary works that underwent thorough filtering, preprocessing, and copyright compliance checks, followed by extensive post-processing. Additionally, a subset of 59 documents totalling 65k words was annotated based on their written dates. Texts from the National Library of Sri Lanka were selected from the SiDiaC-v.1.0 non-filtered list, which was digitised using Google Document AI OCR. This was followed by post-processing to correct formatting issues, address code-mixing, include special tokens, and fix malformed tokens. The construction of SiDiaC-v.2.0 was informed by practices from other corpora, such as FarPaHC, SiDiaC-v.1.0, and CCOHA. This was particularly relevant for syntactic annotation and text normalisation strategies, given the shared characteristics of low-resource language status between Faroese and the similar cleaning strategies utilised in CCOHA. This corpus is categorised into two layers based on genres: primary and secondary. The primary categorisation is binary, assigning each book to either Non-Fiction or Fiction. The secondary categorisation is more detailed, grouping texts under specific genres such as Religious, History, Poetry, Language, and Medical. Despite facing challenges due to limited resources, SiDiaC-v.2.0 serves as a comprehensive resource for Sinhala NLP, building upon the work previously done in SiDiaC-v.1.0.
Figures & tables
Figure 1: Log token counts of the existing Diachronic corpora against the log of language resource level
Figure 2: Sequential Data Filtration Procedure
Figure 3: Examples of Sinhala Text Modernisation and Morpheme Segmentation in Document AI .
Figure 4: An example block of text with <eos> tokens from " Sinhala Bhasha Ithihasaya "
Figure 5: An example poem with <psi> tokens from " Yasodharaawatha ". Note that we have added <eos> tokens at the end of each poem line.
Table 1: Metadata records in SiDiaC-v.2.0 . △ Transliterated (Romanised) by Native Sinhala speakers ⋄ If the authors were unknown, they were labelled as Unknown ⋆ Always included as listed in the digital repository of Natlib † Included when applicable sourced from Sannasgala (2015) ; Soratha Thera (2011)
Corpus
Post-Processing
Date-based Filtering
Documents
Total Words
Unique Words
Total Sentences
SiDiaC-v.1.0
V.1.0
✓
46
58027
22837
-
V.2.0
✓
40
43959
16025
2970
SiDiaC-v.2.0
V.2.0
✗
185
229098
59331
11806
V.2.0
✓
59
64805
21776
4363
Table 2: Summary of Information in SiDiaC-v.1.0 and SiDiaC-v.2.0 .
Appendix figures & tables16 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
Language
Time Span
Token Count
Name
Class
ENGALL ( Davies, 2012 )
English
5
- 1800 – - 1999
8.50×1011
ENGFIC ( Davies, 2012 )
English
5
- 1800 – - 1999
7.50×1010
People in the News ( Hennig and Wilson, 2020 )
English
5
- 2000 – - 2019
1.65×1090
COHA ( Davies, 2012 )
English
5
- 1810 – - 2009
4.10×1080
EDGeS-English ( Bouma et al., 2020 )
English
5
- 1301 – - 2020
3.28×1080
Appendix
Table 3: Summary of the 80 diachronic corpora surveyed; including and sorted by the language class as defined by Ranathunga and de Silva (2022) .
Figure 6: Multi-class timeline of presence and intensity. The visualisation illustrates longitudinal shifts in language class activity, with vertical bar thickness serving as a proxy for rate magnitude.
Figure 7: Examples of malformed tokens in SiDiaC-v.1.0 and the way SiDiaC-v.2.0 has corrected them. These sample images were taken from the books 1 Anagathawanshaya: Methe Budu Siritha and 2 Sarartha Sangrahawa: Prathama Bhagaya .
Figure 8: Examples of code-mixed content in SiDiaC-v.1.0 . These sample texts were taken from the following books: 1 Jubili Warnanawa - English in Latin script, 2 Adhimasa Dheepanaya - Pali in Sinhala script, 3 Moggallana Panchika Pradeepaya - Sanskrit in Sinhala script.
Figure 9: An example of commentaries in SiDiaC-v.1.0 from " Sanna sahitha Salalihini Sandheshaya ".
Figure 10: An example of poetic suffix shifts from " Hansa Sandheshaya " and the way SiDiaC-v.1.0 and SiDiaC-v.2.0 handled them.
Figure 11: An example of multi-column texts from " Kavya Wajrayudhaya - Palamu Kotasa " and the way SiDiaC-v.1.0 and SiDiaC-v.2.0 handled them.
Figure 12: An example of content tables from " Sanskrutha Shabdhamalawa hewath Sanskrutha Nama Waranagilla " in SiDiaC-v.1.0 . Note that the words highlighted in red are malformed tokens.
Figure 13: Examples of footnotes from " † Sadhdharma Rathnawaliya - Prathama Bagaya " and " ⋆ Hansa Sandheshaya ".
Primary Category
Secondary Category
Total
Fiction
Non-Fiction
History
Language
Medical
Poetry
Religious
1800 - 1820
1
0
0
0
0
1
0
1
1820 - 1840
0
1
0
1
0
0
0
1
1840 - 1860
2
1
0
0
0
1
2
3
1860 - 1880
2
6
1
2
0
2
3
8
1880 - 1900
12
51
3
7
1
17
34
*62
Appendix
Table 4: Distribution of Books Across Issued Dates vs Genres in SiDiaC-v.2.0 . *The total count for the secondary category between 1880 - 1900, 1900 - 1920, and 1920 - 1940 CE is 62, 37, and 50, respectively, while the overall number of books in those periods is 63, 40, and 50. This discrepancy arises because the books ‘ Hithopadhesha Sannaya ’, ‘ Dhrawya Gunadharpana Sannaya ’, ‘ Asabandhi Sabandi ’, ‘ Ajuudha Neethiya ’ and ‘ Maadhanaa ’ were not classified under any of the five secondary categories.
Primary Category
Secondary Category
Total
Fiction
Non-Fiction
History
Language
Medical
Poetry
Religious
5th
0
1
0
0
1
0
0
1
12th
1
0
0
0
0
1
0
1
13th
1
12
1
3
0
1
8
13
14th
2
2
1
0
0
0
3
4
15th
4
4
0
2
0
3
3
8
Appendix
Table 5: Distribution of Books Across Written Centuries vs Genres in SiDiaC-v.2.0 - filtered . *The total count for the secondary category in the 20th century amounts to 16, while the overall number of books is 17. This discrepancy arises because the book ‘ Hithopadhesha Sannaya ’, which offers advice, was not classified under any of the five secondary categories.
History
Language
Medical
Poetry
Religious
Unclassified
Total
Fiction
2
0
0
41
5
2
50
Non-Fiction
16
17
5
13
81
3
135
Total
18
17
5
54
86
5
185
Appendix
Table 6: Frequency distribution of books in SiDiaC-v.2.0 categorised by primary and secondary genre classifications.
History
Language
Medical
Poetry
Religious
Unclassified
Total
Fiction
0
0
0
12
4
0
16
Non-Fiction
5
10
2
3
22
1
43
Total
5
10
2
15
26
1
59
Appendix
Table 7: Frequency distribution of books in SiDiaC-v.2.0 - filtered categorised by primary and secondary genre classifications.
Title
Issued Date
Author
Written Date
Genre
OCR Confidence ↑
Primary
Secondary
Adhimasa Sangrahawa
1903
Madhampe Dhammathilaka Himi
1850 - 1903
Non-Fiction
Religious
0.969200
Adhimasa Winishchaya
1904
Walikande Sri Sumangala Himi
1850 - 1904
Non-Fiction
Religious
0.997100
Anagathawanshaya: Methe Budu Siritha
1934
Watadhdhara Medhanandha Himi; Siri Parakumabahu Wilgammula Sangaraja Himi
1325 - 1333
Fiction
Religious
0.999200
Ashoka Shilalipi saha Prathimakarana Winishchaya
1919
D. E. Wickramasuriya
1916
Non-Fiction
History
0.998900
Dhaham Sarana
1931
Unknown
1220 - 1293
Fiction
Religious
0.989100
Appendix
Table 8: The metadata information for the 59 books used in the creation of SiDiaC-v.2.0 - filtered .
Table 9: Frequency Distribution of Neighbour Words for " " from the 13th to the 20th Century.
Table 10: Frequency Distribution of Neighbour Words for " " from the 13th to the 20th Century.
Sinhala is a morphologically rich abugida spoken by roughly 16 million people in Sri Lanka, and to date, there are no publicly available real-world datasets for page-level Sinhala OCR. All previous studies for assessing Sinhala OCR models have used artificially generated data. To bridge the gap, we introduce sinhala-ocr-lk-acts-1010, an annotated dataset of 1,010 page-level images and their transcriptions collected from Sri Lankan Legislative Acts published between 1981-1989 and 2000-2019, split into 707 training examples, 101 validation examples, and 202 testing examples. Three models based on deep learning-based visual language processing, namely DeepSeek-OCR V1, DeepSeek-OCR V2, and LightOnOCR-2-1B, are fine-tuned using QLoRA in 8 experiments conducted on consumer and cloud GPUs. LightOnOCR-2-1B is the top performer, achieving a CER of 1.05% across all test examples, outperforming state-of-the-art open-source OCR models such as Surya-OCR (8.84%) and Tesseract v5 (10.69%), as well as commercially available OCR models such as Google Document AI (2.06%). Our results suggest that LightOnOCR-2-1B outperforms other baselines on real-world OCR tasks and maintains consistent performance across all print periods, even when documents are severely degraded.
Avisha Dilhara, Nevidu Jayatilleke
School of Computing, Informatics Institute of Technology, Sri Lanka · Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka
We present a collection of open, machine-readable document datasets covering parliamentary proceedings, legal judgments, government publications, news, and tourism statistics from Sri Lanka. The collection currently comprises of 278,621 documents (80.7 GB) across 26 datasets in Sinhala, Tamil, and English. The datasets are updated daily and mirrored on GitHub and Hugging Face. These resources aim to support research in computational linguistics, legal analytics, socio-political studies, and multilingual natural language processing. We describe the data sources, collection pipeline, formats, and potential use cases, while discussing licensing and ethical considerations. This manuscript is at version v2026-07-02-0940.
Tracking semantic change in low-resource languages across extensive historical timelines presents significant challenges due to data scarcity and the limitations of static embedding alignments. This study investigates the diachronic evolution of the Sinhala language from the 13th to the 20th century using a multi-stage computational framework. We first align century-specific Word2Vec and FastText embeddings using Similarity Matrix Based Alignment (SMA) and Orthogonal Procrustes (OP) techniques, finding that OP alignment provides more stable neighbourhood tracking for identifying temporal similarity dips. To move beyond aggregate measures, we introduce a Bidirectional Semantic Impact Pruning approach using contextualised embeddings from a fine-tuned Llama-3.1-8B. By applying Leave-One-Out (LOO) diagnostics, we attempt to isolate influential sentences to distinguish between systemic semantic shifts and transient polysemic expansion. Our results show that semantic drift in the fine-tuned Llama-3.1-8B is not evenly distributed across all usages. Instead, a significant part of the change is driven by a smaller set of high-impact contextual instances, rather than gradual and uniform change across all occurrences. This work provides a preliminary framework for low-resource Sinhala diachronic analysis, highlighting the trade-offs between model sensitivity and data availability.
Nevidu Jayatilleke, Nisansa de Silva
Department of Computer Science & Engineering, University of Moratuwa, Sri Lanka