Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.
Figures & tables
Domain
Documents
Tokens
Sentences
Source
News
100,846
52,641,127
897,348
ThaiPBS website
Wikipedia
175,069
101,378,583
1,303,443
Thai Wikipedia (MediaWiki random sampling)
Spoken
2,505
30,238,809
266,944
YouTube transcripts (7+ channels)
Social Media
87,700
157,708,614
2,052,040
Wisesight + Pantip datasets
Total
366,120
341,967,133
4,519,775
Table 1: The four sub-corpora that make up ThaiTrees.
Setting
Value
Base model
PhayaThaiBERT
Training data
UD Thai-TUD
Split (sentences)
2,902 / 362 / 363
Split (word tokens)
62,011 / 7,521 / 7,683
Optimizer
AdamW
β1,β2,ϵ
0.9 , 0.999 , 10−8
Table 2: Fine-tuning configuration for the PhayaThaiBERT POS tagger ( phayathaibert-thai-pos-tagger ). The best checkpoint was selected by held-out accuracy; macro F1 is over the 15 UPOS classes present in the treebank.
Stage
Tool
Benchmark
Source
Word seg.
AttaCut
91 % WL-F1 (BEST)
Chormai et al.,2020
Sent. seg.
CRFcut
87 % / 82 % acc. (ORCHID / TED)
Chumpolsathien,2020
Dep. parsing
AttaParse 1.0
86.4 % UAS / 76.6 % LAS (TUD test) †
Sriwirote et al.,2025
POS tagging
PhayaThaiBERT (ours)
90.64 % / 81.34 % macro F1 (UD Thai-TUD)
this work
Table 3: In-domain benchmarks for the four pipeline stages. † AttaParse 1.0 = the no-POS graph-based PhayaThaiBERT configuration (row GP) of Sriwirote et al. (2025) , Table 3.
Rank
Word
Count
Freq/M
1
6,300,593
18,424.6
2
3,972,907
11,617.8
3
3,928,739
11,488.6
4
3,729,958
10,907.4
5
3,720,946
10,881.0
6
3,655,798
10,690.5
Table 4: Top 20 words across the corpus by overall frequency per million.
News
Wikipedia
Spoken
Social Media
Rank
Word
logOR
Word
logOR
Word
logOR
Word
logOR
1
2.32
2.42
4.07
3.21
2
2.30
2.35
3.28
2.41
3
2.29
2.32
3.25
2.24
4
2.16
2.29
3.11
2.19
5
2.07
2.29
2.74
2.17
Table 5: Top 20 words of each domain by log odds ratio against the union of the other three domains, among words occurring at least 10,000 times in both the target and reference and in more than ten documents of the target domain.
Rank
Pattern
Count
Freq/M
1
NOUN-nmod-NOUN
14,271,548
71,416.1
2
VERB-obj-NOUN
13,929,524
69,704.6
3
VERB-compound-VERB
9,735,805
48,718.9
4
NOUN-acl-VERB
9,083,706
45,455.7
5
NOUN-case-ADP
8,823,314
44,152.7
6
VERB-advmod-ADV
7,575,704
37,909.5
Table 6: Top 20 dependency-edge patterns across the corpus by overall frequency per million. Patterns are written head-relation-dependent .
News
Wikipedia
Spoken
Social Media
Rank
Pattern
logOR
Pattern
logOR
Pattern
logOR
Pattern
logOR
1
NOUN-appos-NUM
2.19
PROPN-appos-PROPN
2.93
PART-det-PART
2.80
ADV-punct-PUNCT
1.92
2
NUM-flat-PROPN
1.12
NOUN-appos-PROPN
1.79
PART-amod-PART
2.77
VERB-obl-PUNCT
1.91
3
VERB-compound-CCONJ
0.91
NOUN-nsubj-PROPN
1.65
PART-compound-PART
2.47
ADV-fixed-PART
1.73
4
VERB-cc-SCONJ
0.83
PROPN-conj-NOUN
1.65
CCONJ-fixed-SCONJ
2.23
VERB-discourse-PART
1.72
5
VERB-csubj-VERB
0.75
PROPN-conj-PROPN
1.63
PART-advmod-PART
2.17
ADV-advmod-PART
1.69
Table 7: Top 20 dependency-edge patterns of each domain by log odds ratio against the union of the other three domains. Patterns are written head-relation-dependent (POS of head, dependency relation, POS of dependent); only patterns with at least 10,000 edges in both the target and reference domains are considered.
Center for Language and Cognition (CLCG), University of Groningen · Georgetown University · Computational Linguistics, Department of Linguistics, Bielefeld University