Quantitative Measurement of Language Distance among Closely Related Indo-European Languages Using Pretrained Language Models: A Case Study on the North Germanic Branch
Authors: Yiping Bai
Organizations: Guangdong Haiqixing Marine Technology Co., Ltd., Guangzhou 510000, China
Among closely related North Germanic languages, the quantification of language distance has traditionally relied on qualitative methods, lacking a unified multi-dimensional computational framework. Multilingual pretrained models based on the Transformer architecture can map texts from different languages into a shared vector space, enabling quantitative measurement of language distance. This paper focuses on the three North Germanic languages---Danish, Norwegian (Bokmål), and Swedish---and proposes a three-metric quantitative framework based on pretrained language models: (1)~sentence-level semantic distance, computed as cosine similarity between LaBSE and mBERT encodings of parallel sentences; (2)~orthographic fragmentation rate, measuring subword tokenization efficiency when cross-applying monolingual BERT vocabularies to parallel texts; (3)~MLM predictability, comparing prediction confidence and entropy in masked language modeling using mBERT across languages. Using 150 trilingual parallel sentence triplets from the Tatoeba corpus as controlled samples, we obtain consistent distance rankings on two independent models: LaBSE: da--no 0.012< no--sv 0.016< da--sv 0.020; mBERT: da--no 0.016< no--sv 0.045≈ da--sv 0.046. This ranking is consistent with the historical linguistic conclusion that ``400 years of Danish rule over Norway (1380--1814) led to highly cognate written languages.'' The three metrics---semantic, orthographic, and predictability---converge on the same conclusion, providing a reproducible computational framework for the quantitative study of distance among closely related languages, extensible in principle to more branches of the Indo-European language family, pending validation on additional language groups.
Figures & tables
#
Danish
Norwegian
Swedish
1
Dette er en midlertidig sætning.
Dette er en midlertidig setning.
Det här är en tillfällig mening.
3
Hvorfor kom du ikke?
Hvorfor kom du ikke?
Varför kom du inte?
7
Jeg bor ikke i Helsinki.
Jeg bor ikke i Helsingfors.
Jag bor inte i Helsingfors.
12
Hun spiller tennis hver dag.
Hun spiller tennis hver dag.
Hon spelar tennis varje dag.
15
Vi har en far.
Vi har en far.
Vi har en far.
Table 1: Sample parallel sentence triplets
Exp.
Model
Size
Hardware
Framework
① Sentence vector
LaBSE
∼ 1.8 GB
CPU
sentence-transformers 5.7.0
② Fragmentation
3 monolingual BERTs
hundreds of KB each
CPU
transformers 5.16.1
③ Dual-metric
mBERT
∼ 700 MB
T4 GPU
transformers 5.16.1, PyTorch 2.11.0
Table 2: Computing resources and model configuration
Metric
Aspect
Model
Logic
Output
Sentence vector distance
Semantics
LaBSE / mBERT
Closer vectors ⇒ closer languages
Cosine distance
Fragmentation rate
Orthography
3 monolingual BERT vocabs
More efficient cross-tokenization ⇒ closer scripts
Subtokens / words
MLM confidence
Predictability
mBERT MLM head
More accurate cross-prediction ⇒ more shared context
Confidence / entropy
Table 3: Overview of three-indicator experimental design
(a) LaBSE
da
no
sv
da
0.0000
0.0121
0.0204
no
0.0121
0.0000
0.0156
sv
0.0204
0.0156
0.0000
(b) mBERT
da
no
sv
Table 4: Distance matrices from two models ( 1−mediancosθ )
Pair
LaBSE
mBERT
da–no
0.9879 / 0.9725 / 0.0474
0.9840 / 0.9753 / 0.0245
no–sv
0.9844 / 0.9669 / 0.0525
0.9551 / 0.9476 / 0.0359
da–sv
0.9796 / 0.9636 / 0.0485
0.9538 / 0.9432 / 0.0410
Table 5: Language pair similarity statistics (median / mean / std)
Figure 1: LaBSE language distance heatmap ( 1−median cosine similarity)
Text \ Vocab
da
no
sv
da
1.2000
1.5000
1.7071
no
1.3693
1.5000
1.6000
sv
1.5000
1.5505
1.2000
Table 6: Fragmentation rate matrix (subtokens per word, median)
Crosslingual evaluation of language models that enables fair comparisons remains a fundamental challenge in multilingual NLP. Existing studies adopt a variety of downstream tasks and intrinsic metrics with different theoretical justifications, yet there has been little empirical investigation into whether these approaches yield meaningful crosslingual conclusions. We systematically examine crosslingual evaluation approaches using controlled monolingual language models trained on parallel data with varying tokenizer vocabulary sizes and model sizes, and further validate our findings on multilingual LLMs. We further discuss challenges in achieving comparable downstream evaluation across languages. Our results show that several widely used normalized metrics introduce crosslinguistic biases rooted in tokenization, encoding, and orthographic differences. In contrast, sentence-level negative log-likelihood computed over semantically equivalent sequences provides more meaningful and consistent crosslingual comparisons.
Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), but direct comparisons are rare; the focus has been on encoder-only models, with most work adapting existing pretrained models. We systematically compare different input representations in autoregressive multilingual pretraining, comparing orthographic text, IPA, and romanization in a controlled setup across three scales (467M, 709M, and 1.03B) on eight languages in four typologically motivated pairs. Across a wide range of downstream tasks on seen and unseen languages, romanized pretraining yields the strongest cross-lingual transfer, and the advantage over text widens with scale. IPA improves over text in most settings but trails romanization. Surprisingly, finetuning a text-pretrained model on romanized data hurts performance on languages already covered by the base model, only marginally helping when the model lacks script coverage. Our results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.
Muge Zhang, Aaron Jencks, Krishna Badikela +2
Department of Computer Science and Engineering, Ohio State University · Paul G. Allen School of Computer Science & Engineering, University of Washington
With the advent of word representations, word similarity tasks are becoming increasing popular as an evaluation metric for the quality of the representations. In this paper, we present manually annotated monolingual word similarity datasets of six Indian languages - Urdu, Telugu, Marathi, Punjabi, Tamil and Gujarati. These languages are most spoken Indian languages worldwide after Hindi and Bengali. For the construction of these datasets, our approach relies on translation and re-annotation of word similarity datasets of English. We also present baseline scores for word representation models using state-of-the-art techniques for Urdu, Telugu and Marathi by evaluating them on newly created word similarity datasets.
Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee +2
International Institute of Information Technology Hyderabad, Telangana, India