Organizations: Key Laboratory of Linguistics, Chinese Academy of Social Sciences (University of Chinese Academy of Social Sciences), Beijing · Rixin College, Tsinghua University, Beijing · Institute of Linguistics, Chinese Academy of Social Sciences, Beijing · Institute of Ethnology and Anthropology, Chinese Academy of Social Sciences, Beijing · School of Software and Microelectronics, Peking University, Beijing
Tangut is an extinct language whose script does not explicitly mark word boundaries. We present the first systematic study of Tangut word segmentation using 2,750 expert-annotated segments (31,893 tokens), traditional lexicons, and unlabeled text. Our framework combines a reliability-calibrated lexicon-lattice representation, explicit distributional statistics, and a lightweight character encoder pretrained with MLM. In within-source five-fold cross-validation, the model integrating TangutEncoder, CRF, and external features obtains the numerically highest main-system mean F1 of 0.911 and substantially improves recall beyond the labeled training vocabulary. We further evaluate document-level transfer on 479 segments (4081 tokens) from five works absent from the annotated training corpus. You can access our project at https://github.com/jiangli-va/TangutSeg.
Figures & tables
Figure 1: Surviving fragment of Mahāratnakūṭa-sūtra , Scroll 68.
Category
Segments
Tokens
Types
Buddhist scriptures
234
3717
769
Secular documents
2516
28176
4046
Total
2750
31893
4433
Table 1: Statistics of the expert-annotated Tangut corpus.
Symbol
Feature group
Dim.
BIE
Lexicon-lattice indicators
11
R
Reliability of observed entries
3
P
Prior for unseen entries
3
M
Lexicographic metadata
3
Dictall
Full lexicon representation
20
Table 2: Composition of the 20-dimensional lexicon representation.
Symbol
Feature group
Dim.
Freq
Bigram frequency
2
Coo
Character association
2
Ent
Neighbor entropy
4
Distall
Full distributional representation
8
Table 3: Explicit distributional features extracted from unlabeled text.
Model
P
R
F 1
OOV-R
IV-R
Supervised baselines (RQ1)
Dict-corpus
0.845±0.005
0.890±0.003
0.867±0.004
0.219±0.014
0.954±0.002
CRF
0.880±0.004
0.888±0.003
0.884±0.003
0.415±0.026
0.933±0.003
BiLSTM–CRF
0.862±0.008
0.873±0.008
0.868±0.008
0.472±0.014
0.912±0.006
Transformer-Random
0.848±0.008
0.863±0.010
0.855±0.009
0.519±0.020
0.896±0.010
Lexicon-lattice ablation (RQ2)
Table 4: Overall segmentation results under five-fold cross-validation.
Secular
Religious
Model
F 1
OOV-R
IV-R
F 1
OOV-R
IV-R
Dict-corpus
0.869±0.004
0.229±0.017
0.956±0.003
0.855±0.019
0.132±0.028
0.936±0.010
CRF
0.887±0.010
0.430±0.035
0.934±0.007
0.839±0.011
0.262±0.067
0.909±0.009
BiLSTM–CRF
0.875±0.007
0.498±0.016
0.917±0.006
0.810±0.021
0.244±0.028
0.875±0.021
Transformer-Random
0.868±0.008
0.545±0.018
0.906±0.010
0.760±0.014
0.280±0.060
0.816±0.017
CRF+ Dictall
0.904±0.005
0.481±0.024
0.942±0.004
0.872±0.004
0.301±0.058
0.938±0.008
Table 5: Mean performance by document genre.
Model
F 1
CRF
0.686
CRF + Dictall
0.726
CRF + Dictall + Distall
0.747
Transformer-Random
0.642
Transformer-Char2Vec
0.652
TangutEncoder
0.722
Table 6: Segmentation results on five unseen works. Full results are provided in Appendix C.4 .
Appendix figures & tables18 assets
Supplementary material from the paper’s appendix.
Appendix
Audit unit
All
Religious
Secular
Exact segment ( ≥3 chars)
0
0
0
5-gram overlap (%)
0.37
2.66
0.02
10-gram overlap (%)
0.01
0.08
0.00
Appendix
Table 7: Normalized overlap between annotated and unlabeled text.
Figure 2: Annotation examples from the religious and secular portions of the corpus. Vertical bars indicate expert-annotated word boundaries.
Group
Labels
Core lexical categories
a, c, d, m, n, p, q, r, u, v
Fine-grained lexical labels
b, l, t, nb, nc, nh, nl, no, ns, mc, mo, rd, ri, rp
Table 9: Composition of the independent cross-document test set. OOV rates are computed over word tokens relative to the complete supervised training corpus. Word types in the total row are deduplicated across the five works.
Parameter
Value
L1 coefficient ( c1 )
1.0
L2 coefficient ( c2 )
10−3
Maximum L-BFGS iterations
200
Validation-based stopping
Not used
All possible transitions
Enabled
Character context window
±2 characters
Appendix
Table 10: Hyperparameters of the linear CRF. All reported CRF results use a fixed budget of 200 L-BFGS iterations.
Layers
Emb.
Hidden
Batch
Dropout
LR
F 1
2
100
256
512
0.3
5×10−4
0.8636
2
100
128
512
0.3
5×10−4
0.8595
2
100
64
512
0.3
5×10−4
0.8905
2
100
32
512
0.3
5×10−4
0.8873
2
100
16
512
0.3
5×10−4
0.8423
1
100
64
512
0.3
5×10−4
0.8638
Appendix
Table 11: Preliminary hyperparameter search for the supervised BiLSTM–CRF. Emb. denotes the character-embedding dimension, and Hidden denotes the concatenated output dimension of the two directions. The best segmentation F 1 is highlighted.
Parameter
Selected value
Character embedding size
100
BiLSTM layers
2
Hidden size
32 per direction
BiLSTM output size
64
Emission size
4
Model dropout
0.3
Appendix
Table 12: Selected architectural and training parameters of the BiLSTM–CRF. The maximum number of epochs is only an upper bound; the checkpoint with the lowest development loss is restored after early stopping.
Parameter
Value
Character embedding size
192
Maximum sequence length
128
Transformer layers
3
Attention heads
4
Dimension per head
48
Feed-forward size
768
Appendix
Table 13: Shared architecture of the Transformer–CRF variants. The two external-feature projections are included only in the corresponding fusion models.
Char2Vec parameter
Value
Training objective
Skip-gram
Vector size
192
Context window
5
Negative samples
10
Minimum frequency
1
Training epochs
30
Appendix
Table 14: Training parameters of the static Char2Vec initialization.
MLM parameter
Value
Masking ratio
0.15
Span-masking probability
0.50
Span length
2–4 characters
Replacement strategy
80/10/10
Optimizer
AdamW
Learning rate
3×10−4
Appendix
Table 15: Masked-language-model pretraining parameters of TangutEncoder. Pretraining completed the full 5,000 steps, and the checkpoint with the lowest validation loss was retained.
Fine-tuning parameter
Value
Frozen-encoder epochs
3
Optimizer
AdamW
Encoder learning rate
5×10−5
Task-head learning rate
5×10−4
Weight decay
0.01
Batch size
32 sentences
Appendix
Table 16: Downstream training parameters shared by Transformer–Random, Transformer–Char2Vec and TangutEncoder.
Association features
P
R
F 1
OOV-R
IV-R
Freq + dPMI
0.906±0.003
0.905±0.003
0.905±0.003
0.486±0.009
0.945±0.001
Freq + dPMI + Ent
0.907±0.004
0.906±0.004
0.907±0.004
0.491±0.016
0.947±0.002
Freq + Dice
0.905±0.003
0.905±0.004
0.905±0.003
0.485±0.010
0.945±0.002
Freq + Dice + Ent
0.907±0.004
0.906±0.004
0.906±0.004
0.489±0.016
0.946±0.002
Freq + t -score
0.905±0.004
0.904±0.004
0.904±0.004
0.483±0.014
0.944±0.003
Freq + t -score + Ent
0.907±0.003
0.906±0.004
0.907±0.004
0.490±0.013
0.946±0.002
Appendix
Table 17: Comparison of association measures under five-fold cross-validation.
Model
P
R
F 1
OOV-R
IV-R
Dict-corpus
.845
.890
.867
.219
.954
Dict-dictionary
.787
.728
.756
.666
.733
Dict-all
.821
.740
.779
.665
.748
Appendix
Table 18: Bidirectional maximum matching with different dictionary sources. Values are five-fold means.
Model
External features
P
R
F 1
OOV-R
IV-R
BiLSTM
0
0.862±0.008
0.873±0.008
0.868±0.008
0.472±0.014
0.912±0.006
BiLSTM+BIE
11
0.885±0.005
0.894±0.004
0.889±0.004
0.546±0.011
0.927±0.004
BiLSTM+BIE+R
14
0.891±0.003
0.898±0.004
0.894±0.004
0.546±0.009
0.931±0.003
BiLSTM+ Dictall
17
0.893±0.003
0.900±0.003
0.897±0.003
0.551±0.011
0.933±0.003
+ Dictall + IL
28
0.896±0.003
0.900±0.002
0.898±0.002
0.557±0.009
0.933±0.002
+ Dictall + IL + DD
30
0.897±0.004
0.904±0.003
0.900±0.003
0.554±0.006
0.938±0.002
Appendix
Table 19: BiLSTM–CRF variants with progressively richer external features.
Secular
Religious
Model
F 1
OOV-R
F 1
OOV-R
BiLSTM
0.875±0.007
0.498±0.016
0.810±0.021
0.244±0.028
BiLSTM+BIE
0.897±0.005
0.573±0.015
0.830±0.018
0.305±0.043
BiLSTM+BIE+R
0.903±0.004
0.574±0.014
0.829±0.013
0.286±0.053
BiLSTM+ Dictall
0.905±0.003
0.579±0.014
0.835±0.016
0.298±0.037
+ Dictall + IL
0.905±0.002
0.586±0.011
0.842±0.018
0.300±0.043
Appendix
Table 20: Genre-level results for the BiLSTM–CRF feature variants.
Pretraining
F 1
OOV-R
IV-R
MLM
0.911
0.608
0.946
MLM + WordRank
0.912
0.617
0.946
Appendix
Table 21: Downstream results of lexicon-aware continued pretraining. Both encoders are evaluated with the same dictionary and distributional features.
Model
P
R
F 1
POS Acc.
POS N
Overall (479 segments)
CRF
0.6410
0.7381
0.6861
–
–
CRF + Dictall
0.6850
0.7711
0.7255
–
–
CRF + Dictall + Distall
0.7114
0.7863
0.7470
–
–
Transformer-Random
0.5915
0.7008
0.6415
–
–
Transformer-Char2Vec
0.6015
0.7121
0.6522
–
–
Appendix
Table 22: Full segmentation and POS-tagging results on five unseen works. POS accuracy is calculated only over words whose predicted boundaries exactly match the gold segmentation; POS N denotes the number of such words.
Figure 3: The Tangut Digital Intelligence Platform.
We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair. Tangkhul is a severely under-resourced Tibeto-Burman language spoken primarily in Manipur, India, with virtually no prior natural language processing infrastructure. We describe two systems: (1) a primary system based on ByT5-large fine-tuned on 38,336 Tangkhul-English parallel sentence pairs, and (2) a contrastive system based on mT5-small fine-tuned on the same corpus. Our primary ByT5-large system achieves a corpus BLEU score of 39.97, chrF++ of 58.07, BERTScore F1 of 0.8104, and COMET (wmt22-comet-da) of 0.7302 on a held-out test set of 3,856 sentences. We further discuss the orthographic challenges specific to Tangkhul's Latin-script diacritics, the domain bias of our training corpus (which comprises biblical text, stories, and conversational data), and avenues for future improvement through data diversification and domain adaptation.
Automatic Text Recognition (ATR) now supplies digital humanities with large volumes of unstructured, heterogeneous, and often noisy text in ancient languages. Downstream semantic analysestext reuse identification, alignment, and semantic search-rely on sentence embeddings, yet existing methods transfer poorly to ancient languages: generic multilingual encoders underperform, specialized language models yield anisotropic representation spaces, and labeled similarity data is unavailable. We study two fully unsupervised strategies - TSDAE and contrastive sentence embedding (CSE) - that adapt a specialized token-level language model into a corpus-specific sentence encoder using only raw sentences. On the philologically central case of biblical reuse in patristic literature (2,935 expert-verified parallels in Latin and Ancient Greek, from Augustine, Jerome, and Athanasius), we decompose reuse identification into two separately evaluated tasks-binary detection and correspondence retrieval-and benchmark the adapted encoders against multilingual, specialized, distilled, and supervised fine-tuned baselines, as well as on artificially noised data simulating HTR artifacts and scribal abbreviations. The adapted encoders outperform all baselines on both tasks, with complementary profiles: TSDAE leads detection given a large in-domain corpus, while CSE leads retrieval, reaches its optimum with as few as 4-8k raw in-domain sentences-a few tens of seconds of training on a laptop GPU-and transfers across works and authors, including to noisy post-ATR text when retrained directly on it. UMAP atlases relate the geometric effect of each strategy to the measured gains, and the full pipeline-segmentation, fine-tuning, cross-corpus semantic search-is made available to non-specialists through the online tool Paraphrasis.
Th{é}otime de la Selle
ISC, HiSoMA, CNRS · Institut des Sources Chrétiennes, HiSoMA, CNRS, Lyon, France
Text recognition, or extracting electronic text from document images, has been indispensable for knowledge retrieval tasks, such as retrieval-augmented generation (RAG). For Khmer, extracted text is subject to an extra word segmentation step, as Khmer does not use any visible word delimiters to denote word boundaries. Thus, a recognition-then-segmentation pipeline for Khmer requires two separate sequential models; this is not only error-prone but also adds significant latency for large-scale document processing. This paper proposes a novel joint Khmer text recognition and word segmentation framework in a unified model. The proposed model, using a connectionist-temporal-classification (CTC) decoder for fast, parallel decoding, can be instructed to recognize Khmer text with (b=1) and without (b=0) word segmentation. Experimental results on different benchmark datasets of different document modalities (document, scene, and handwritten images) show that the proposed model can not only recognize characters in document images but also locate word boundaries, removing the need for an extra word segmentation step in a conventional sequential pipeline.
Marry Kong, Rina Buoy, Sovisal Chenda +3
Techo Startup Center, Ministry of Economy and Finance, Phnom Penh, Cambodia · Osaka Metropolitan University, Osaka, Japan