cs.CLSep 28, 2026

Quantitative Measurement of Language Distance among Closely Related Indo-European Languages Using Pretrained Language Models: A Case Study on the North Germanic Branch

Authors: Yiping Bai

Organizations: Guangdong Haiqixing Marine Technology Co., Ltd., Guangzhou 510000, China

Abstract

Among closely related North Germanic languages, the quantification of language distance has traditionally relied on qualitative methods, lacking a unified multi-dimensional computational framework. Multilingual pretrained models based on the Transformer architecture can map texts from different languages into a shared vector space, enabling quantitative measurement of language distance. This paper focuses on the three North Germanic languages---Danish, Norwegian (Bokmål), and Swedish---and proposes a three-metric quantitative framework based on pretrained language models: (1)~sentence-level semantic distance, computed as cosine similarity between LaBSE and mBERT encodings of parallel sentences; (2)~orthographic fragmentation rate, measuring subword tokenization efficiency when cross-applying monolingual BERT vocabularies to parallel texts; (3)~MLM predictability, comparing prediction confidence and entropy in masked language modeling using mBERT across languages. Using 150 trilingual parallel sentence triplets from the Tatoeba corpus as controlled samples, we obtain consistent distance rankings on two independent models: LaBSE: da--no 0.012<0.012 < no--sv 0.016<0.016 < da--sv 0.0200.020; mBERT: da--no 0.016<0.016 < no--sv 0.045≈0.045 \approx da--sv 0.0460.046. This ranking is consistent with the historical linguistic conclusion that ``400 years of Danish rule over Norway (1380--1814) led to highly cognate written languages.'' The three metrics---semantic, orthographic, and predictability---converge on the same conclusion, providing a reproducible computational framework for the quantitative study of distance among closely related languages, extensible in principle to more branches of the Indo-European language family, pending validation on additional language groups.

Figures & tables

Explore similar work

CardsList
  1. Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

    Aug 25, 2026Xiulin Yang, Ethan Gotlieb Wilcox, Catherine ArnettFruit

  2. One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

    Aug 26, 2026Muge Zhang, Aaron Jencks, Krishna Badikela +2Multilingual Language ModelsAccents

  3. Word Similarity Datasets for Indian Languages: Annotation and Baseline Systems

    Sep 28, 2026Syed S. Akhtar, Arihant Gupta, Avijit Vajpayee +2Indian LanguagesHindi