Quantitative Measurement of Language Distance among Closely Related Indo-European Languages Using Pretrained Language Models: A Case Study on the North Germanic Branch
Authors: Yiping Bai
Organizations: Guangdong Haiqixing Marine Technology Co., Ltd., Guangzhou 510000, China
Among closely related North Germanic languages, the quantification of language distance has traditionally relied on qualitative methods, lacking a unified multi-dimensional computational framework. Multilingual pretrained models based on the Transformer architecture can map texts from different languages into a shared vector space, enabling quantitative measurement of language distance. This paper focuses on the three North Germanic languages---Danish, Norwegian (Bokmål), and Swedish---and proposes a three-metric quantitative framework based on pretrained language models: (1)~sentence-level semantic distance, computed as cosine similarity between LaBSE and mBERT encodings of parallel sentences; (2)~orthographic fragmentation rate, measuring subword tokenization efficiency when cross-applying monolingual BERT vocabularies to parallel texts; (3)~MLM predictability, comparing prediction confidence and entropy in masked language modeling using mBERT across languages. Using 150 trilingual parallel sentence triplets from the Tatoeba corpus as controlled samples, we obtain consistent distance rankings on two independent models: LaBSE: da--no 0.012< no--sv 0.016< da--sv 0.020; mBERT: da--no 0.016< no--sv 0.045≈ da--sv 0.046. This ranking is consistent with the historical linguistic conclusion that ``400 years of Danish rule over Norway (1380--1814) led to highly cognate written languages.'' The three metrics---semantic, orthographic, and predictability---converge on the same conclusion, providing a reproducible computational framework for the quantitative study of distance among closely related languages, extensible in principle to more branches of the Indo-European language family, pending validation on additional language groups.
Figures & tables
#
Danish
Norwegian
Swedish
1
Dette er en midlertidig sætning.
Dette er en midlertidig setning.
Det här är en tillfällig mening.
3
Hvorfor kom du ikke?
Hvorfor kom du ikke?
Varför kom du inte?
7
Jeg bor ikke i Helsinki.
Jeg bor ikke i Helsingfors.
Jag bor inte i Helsingfors.
12
Hun spiller tennis hver dag.
Hun spiller tennis hver dag.
Hon spelar tennis varje dag.
15
Vi har en far.
Vi har en far.
Vi har en far.
Table 1: Sample parallel sentence triplets
Exp.
Model
Size
Hardware
Framework
① Sentence vector
LaBSE
∼ 1.8 GB
CPU
sentence-transformers 5.7.0
② Fragmentation
3 monolingual BERTs
hundreds of KB each
CPU
transformers 5.16.1
③ Dual-metric
mBERT
∼ 700 MB
T4 GPU
transformers 5.16.1, PyTorch 2.11.0
Table 2: Computing resources and model configuration
Metric
Aspect
Model
Logic
Output
Sentence vector distance
Semantics
LaBSE / mBERT
Closer vectors ⇒ closer languages
Cosine distance
Fragmentation rate
Orthography
3 monolingual BERT vocabs
More efficient cross-tokenization ⇒ closer scripts
Subtokens / words
MLM confidence
Predictability
mBERT MLM head
More accurate cross-prediction ⇒ more shared context
Confidence / entropy
Table 3: Overview of three-indicator experimental design
(a) LaBSE
da
no
sv
da
0.0000
0.0121
0.0204
no
0.0121
0.0000
0.0156
sv
0.0204
0.0156
0.0000
(b) mBERT
da
no
sv
Table 4: Distance matrices from two models ( 1−mediancosθ )
Pair
LaBSE
mBERT
da–no
0.9879 / 0.9725 / 0.0474
0.9840 / 0.9753 / 0.0245
no–sv
0.9844 / 0.9669 / 0.0525
0.9551 / 0.9476 / 0.0359
da–sv
0.9796 / 0.9636 / 0.0485
0.9538 / 0.9432 / 0.0410
Table 5: Language pair similarity statistics (median / mean / std)
Figure 1: LaBSE language distance heatmap ( 1−median cosine similarity)
Text \ Vocab
da
no
sv
da
1.2000
1.5000
1.7071
no
1.3693
1.5000
1.6000
sv
1.5000
1.5505
1.2000
Table 6: Fragmentation rate matrix (subtokens per word, median)
Department of Computer Science and Engineering, Ohio State University · Paul G. Allen School of Computer Science & Engineering, University of Washington