Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
Figures & tables
Figure 1
Figure 1: Percentage of instances with matching input and output scripts for QIRIM dataset, per model.
Figure 2: Script forcing results. Scores represent the percentage of cases where the script of the generated text matched the requested output script. Axes represent language-script pairs formatted as L-S, where L indicates the output language (S: Serbian, F: French, E: English, K: Kazakh) and S indicates the script (L: Latin, C: Cyrillic, A: Arabic). The language in parentheses indicates the prompt language.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Lang-script
Total words
Types
Sentences
Kz-ar
27,805,939
810,230
57,827
Kz-cyr
4,656,620
301,614
371,902
Kz-lat
430
371
100
Sr-cyr
15,255
4,424
150
Sr-lat
569,206
89,537
2,919
Cr-cyr
517,595
109,847
77,688
Appendix
Table 1: Summary statistics of datasets
Figure 3: Size of the vocabulary (number of non-character tokens) per language-script combination, for each model. (with Zipf)
Config
gpt-4o-m
gpt-4o
Deepseek-chat
Haiku
Sonnet
Gemini-flash
Gemini
F
C
S
F
C
S
F
C
S
F
C
S
F
C
S
F
C
S
F
C
S
F-C vs F-L
-4
0
-28
-8
-4
-56
-16
-14
-99
-17
-18
-75
-18
-17
-50
-7
-4
-73
-13
-16
-61
S-C vs S-L (S-L)
-3
-9
-10
-2
-5
0
-4
-4
0
-2
-8
9
-6
-8
0
-12
-15
0
0
-6
-3
S-C vs S-L (S-C)
-8
-11
-60
-6
-9
0
-4
-5
-1
6
1
-5
-4
-6
0
-6
-10
-1
-1
-2
-36
S-C vs S-L (E)
6
5
-51
2
-1
-1
0
4
0
-9
-10
0
-3
-4
0
-24
-29
-44
-10
-8
-44
K-C vs K-A (K)
-21
-13
-8
-15
-9
-6
4
5
-1
-9
7
-7
-2
4
-1
10
9
-2
-5
1
-1
Appendix
Table 2: Quality score differences across script configurations of LLMs. F: Factuality, C: Completeness, S: Script fidelity. Lines name are language-script pairs formatted as L-S, where L indicates the output language (S: Serbian, F: French, E: English, K: Kazakh) and S indicates the script (L: Latin, C: Cyrillic, A: Arabic). The language in parentheses indicates the prompt language.
Config
DeepSeek-R1-Distill-Llama-8B
Llama-2-7B-Chat
Mistral-7B-Instruct
Mixtral-8x7B-Instruct
F
C
S
F
C
S
F
C
S
F
C
S
F-C vs F-L
5
2
-99
-1
-2
-100
-4
2
-84
-4
-7
-93
S-C vs S-L (S-L)
0
-4
-42
-5
-3
-34
2
0
68
-2
-3
-36
S-C vs S-L (S-C)
2
2
-63
-4
-2
-99
-11
-10
-99
4
8
-95
S-C vs S-L (E)
1
4
-96
-1
1
-99
-11
-10
-63
2
-2
-89
K-C vs K-A (K)
3
6
-26
0
0
-55
0
0
-85
10
7
11
Appendix
Table 3: Quality score differences across script configurations of SLMs. F: Factuality, C: Completeness, S: Script fidelity. Lines name are language-script pairs formatted as L-S, where L indicates the output language (S: Serbian, F: French, E: English, K: Kazakh) and S indicates the script (L: Latin, C: Cyrillic, A: Arabic). The language in parentheses indicates the prompt language.
Figure 4: Perplexity distributions across script configurations for each SLM. Each panel shows perplexity values for different language-script combinations, with the language of the prompt in parenthesis.
Language Script Type
Number of Languages
Mono-scriptic
78
Multi-scriptic
10
Mixed
9
Appendix
Table 4: Number of languages in different script-based categories based on Tatoeba and literature. Mono-scriptic (multi-scriptic) are those languages for which both sources agree that they use a single (multiple) script(s), respectively. Mixed category languages use different number of scripts (mono in Tatoeba and multi in the literature, or vice versa).
Mixed Category
Multi-script
Category
Bulgarian
Bosnian
Kurdish (Central)
Chavacano
Javanese
Kunming
Korean
Kazakh
Marathi
Kurmanji
Appendix
Table 5: Left: Mixed category languages according to Tatoeba and literature (using different number of scripts in Tatoeba with respect to the literature). Right: Multi-script category containing languages which both Tatoeba and literature consider multi-scriptic.
Table 6: Literature sources (left) explicitly mentioning multi-scriptic feature of a language (right).
Figure 5: Percentage of languages by script usage. Blue bars represent languages for which the model generates text in a single script, while orange bars represent languages for which the model uses two or more scripts. The input groups correspond to the language categories defined in Sub-Section D.1 .
Figure 6: Distribution of output scripts for Latin (right column) and Cyrillic (left column) input prompts for open-source SLMs. Lighter colors indicate higher frequency
Figure 7: Distribution of output scripts for Latin (right column) and Cyrillic (left column) input prompts for proprietary LLMs. Lighter colors indicate higher frequency
Figure 15
Nb
Language prompt
Script prompt
Language Question
Requested language output
Requested script output
1
French
Latin
French
French
Latin
2
French
Latin
French
French
Cyrillic
3
English
Latin
French
French
Latin
4
English
Latin
Kazakh
Kazakh
Cyrillic
5
English
Latin
Kazakh
Kazakh
Latin
6
English
Latin
Kazakh
Kazakh
Arabic
Appendix
Table 7: Detail of configurations for the script forcing experiment (see Section 2 ).
Model Type
Adaptation
Forcing
Mixtral (Quantized)
∼ 15 hrs
∼ 5 hrs
Other SLMs
∼ 7 hrs
∼ 2 hrs
LLMs (via API)
∼ 1-5 hr
≤ 1-3 hrs
Appendix
Table 8: Approximate execution times per 1,000 prompts
Model
Adaptation
Forcing
GPT-4o-mini
∼ $0.5
∼ $1
GPT-4o
∼ $3
∼ $10
DeepSeek-Chat
∼ $0.5
∼ $2
Claude-3.5-Haiku
∼ $1.5
∼ $5
Claude-Sonnet-4
∼ $5
∼ $20
Gemini-Flash
∼ $1
∼ $3
Appendix
Table 9: Approximate API costs for each experimental configuration.
Many languages are written in multiple scripts, requiring large language models (LLMs) to generate equivalent linguistic content in distinct orthographic forms. While prior work suggests that LLMs route information through shared latent representations, how they internally mediate script variation remains poorly understood. We study this question by first examining per-layer output distributions with the logit lens, which reveals consistent latent romanization during transliteration, and then through representational and mechanistic analyses of script generation. At the representational level, we show that scripts of the same language become increasingly separable across layers and that a simple linear steering direction can flip a model's output script while largely preserving semantic content. The vector generalizes to writing systems unseen during construction: it flips non-Latin output to Latin, but maps Latin output into varied non-Latin scripts. At the mechanistic level, we localize a small set of late-layer attention heads that causally mediate script choice. These heads transfer across unrelated languages and writing systems, suggesting that script routing is implemented by language-agnostic components. Across both analyses, we observe a consistent directional asymmetry: non-Latin output is produced by a compact, identifiable gate, while Latin-script output emerges from diffuse contributions across the model. Overall, our findings hint that LLMs organize script variation around shared latent representations while exhibiting a privileged substrate toward Latin script.
Daniil Gurgurov, Alan Saji, Katharina Trinley +2
D German Research Center for AI (DFKI) · I Indian Institute of Technology, Madras · H University of Hamburg +1
In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a clear asymmetry: both the input script and the instructed output script are encoded in the earliest layers of the network, while, in contrast, commitment to the actual output script emerges only in the final layers, with the model's intermediate representations defaulting to Latin throughout most of the layers. This two-stage process is confirmed by logit-lens analyses, which show that script commitment consistently occurs at the very last layers of the LLMs. Together with the weaker script-following performance observed in smaller models, these results form a converging body of evidence linking script commitment to model depth, with broader implications for the design of sufficiently deep, inclusive multilingual architectures.
Large language models process the world's writing systems with radical inequality. We constructed the Digital Script Representation Index (DSRI), a seven-axis measure of digital support, and applied it to the 300 writing systems of the Global Script Database (Fukui, 2026). Only 29 scripts (9.7%) are fully supported by contemporary digital infrastructure; among 158 living scripts, 60 (38.0%) lack complete support. Tokenizer efficiency varies by a factor of 31.7 across 45 scripts measured with parallel text. A serial mediation model -- imperial intervention to speaker population to web corpus to tokenizer efficiency -- is consistent with full mediation, with the direct effect of empire indistinguishable from zero (beta = -0.22, p = 0.39) and structural equation model fit indices indistinguishable from saturation at n = 45; the bias-corrected bootstrap CI grazes zero, and we treat the mediation as suggestive rather than confirmatory. Across four independent LLM families (Claude, GPT-4o, Grok, DeepSeek; 12,000 API calls), base-rate-deviation error patterns converge at Spearman rho = 0.85-0.98 (all p < 0.002). 172 script-feature items are answered identically wrong by all four models; over-attribution outnumbers under-recognition 3.9:1, and "used for religion" alone concentrates 43.6% of convergent errors (enrichment 4.1x). With religion excluded as a sensitivity check, the cross-architecture convergence is preserved (mean rho = 0.87 on nine features) and the over-attribution asymmetry persists at 1.77:1 (n = 97, binomial p = 0.008), indicating multi-channeled rather than single-channeled bias. The findings are consistent with an interpretation in which the structural inequalities historical empires inflicted on script communities persist in contemporary language models through the shared training corpus rather than through any individual model's design choices.
Hiroki Fukui
Research Institute of Criminal Psychiatry / Sex Offender Medical Center · Department of Neuropsychiatry, Kyoto University