ThaiTrees: Thai Syntactic Dependency Trees Across Domains
Organizations: Department of Linguistics Chulalongkorn University
Abstract
Studying syntactic patterns in naturally occurring language requires a large parsed corpus, but manual annotation is costly and difficult to scale. Thai has a manually annotated dependency treebank for training and evaluating parsers, but lacks a large automatically parsed corpus for quantitative syntactic research. We present ThaiTrees, a 342M-token corpus drawn from news, Wikipedia, spoken transcripts, and social media. We develop a reproducible pipeline for cleaning, processing, and parsing Thai text under the Universal Dependencies framework. The resulting corpus makes grammatical relations searchable and supports the study of syntactic distributions. We release a frequency lexicon and CoNLL-U parses in machine-readable formats suitable for both AI-assisted and conventional programmatic analysis.
Figures & tables
| Domain | Documents | Tokens | Sentences | Source |
|---|---|---|---|---|
| News | 100,846 | 52,641,127 | 897,348 | ThaiPBS website |
| Wikipedia | 175,069 | 101,378,583 | 1,303,443 | Thai Wikipedia (MediaWiki random sampling) |
| Spoken | 2,505 | 30,238,809 | 266,944 | YouTube transcripts (7+ channels) |
| Social Media | 87,700 | 157,708,614 | 2,052,040 | Wisesight + Pantip datasets |
| Total | 366,120 | 341,967,133 | 4,519,775 |
| Setting | Value |
|---|---|
| Base model | PhayaThaiBERT |
| Training data | UD Thai-TUD |
| Split (sentences) | 2,902 / 362 / 363 |
| Split (word tokens) | 62,011 / 7,521 / 7,683 |
| Optimizer | AdamW |
| , , |
| Stage | Tool | Benchmark | Source |
|---|---|---|---|
| Word seg. | AttaCut | 91 % WL-F1 (BEST) | Chormai et al.,2020 |
| Sent. seg. | CRFcut | 87 % / 82 % acc. (ORCHID / TED) | Chumpolsathien,2020 |
| Dep. parsing | AttaParse 1.0 | 86.4 % UAS / 76.6 % LAS (TUD test) | Sriwirote et al.,2025 |
| POS tagging | PhayaThaiBERT (ours) | 90.64 % / 81.34 % macro F1 (UD Thai-TUD) | this work |
| Rank | Word | Count | Freq/M |
|---|---|---|---|
| 1 | 6,300,593 | 18,424.6 | |
| 2 | 3,972,907 | 11,617.8 | |
| 3 | 3,928,739 | 11,488.6 | |
| 4 | 3,729,958 | 10,907.4 | |
| 5 | 3,720,946 | 10,881.0 | |
| 6 | 3,655,798 | 10,690.5 |
| News | Wikipedia | Spoken | Social Media | |||||
|---|---|---|---|---|---|---|---|---|
| Rank | Word | Word | Word | Word | ||||
| 1 | 2.32 | 2.42 | 4.07 | 3.21 | ||||
| 2 | 2.30 | 2.35 | 3.28 | 2.41 | ||||
| 3 | 2.29 | 2.32 | 3.25 | 2.24 | ||||
| 4 | 2.16 | 2.29 | 3.11 | 2.19 | ||||
| 5 | 2.07 | 2.29 | 2.74 | 2.17 | ||||
| Rank | Pattern | Count | Freq/M |
|---|---|---|---|
| 1 | NOUN-nmod-NOUN | 14,271,548 | 71,416.1 |
| 2 | VERB-obj-NOUN | 13,929,524 | 69,704.6 |
| 3 | VERB-compound-VERB | 9,735,805 | 48,718.9 |
| 4 | NOUN-acl-VERB | 9,083,706 | 45,455.7 |
| 5 | NOUN-case-ADP | 8,823,314 | 44,152.7 |
| 6 | VERB-advmod-ADV | 7,575,704 | 37,909.5 |
| News | Wikipedia | Spoken | Social Media | |||||
|---|---|---|---|---|---|---|---|---|
| Rank | Pattern | Pattern | Pattern | Pattern | ||||
| 1 | NOUN-appos-NUM | 2.19 | PROPN-appos-PROPN | 2.93 | PART-det-PART | 2.80 | ADV-punct-PUNCT | 1.92 |
| 2 | NUM-flat-PROPN | 1.12 | NOUN-appos-PROPN | 1.79 | PART-amod-PART | 2.77 | VERB-obl-PUNCT | 1.91 |
| 3 | VERB-compound-CCONJ | 0.91 | NOUN-nsubj-PROPN | 1.65 | PART-compound-PART | 2.47 | ADV-fixed-PART | 1.73 |
| 4 | VERB-cc-SCONJ | 0.83 | PROPN-conj-NOUN | 1.65 | CCONJ-fixed-SCONJ | 2.23 | VERB-discourse-PART | 1.72 |
| 5 | VERB-csubj-VERB | 0.75 | PROPN-conj-PROPN | 1.63 | PART-advmod-PART | 2.17 | ADV-advmod-PART | 1.69 |