Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.
Figures & tables
Figure 1
Figure 1: Percentage of instances with matching input and output scripts for QIRIM dataset, per model.
Figure 2: Script forcing results. Scores represent the percentage of cases where the script of the generated text matched the requested output script. Axes represent language-script pairs formatted as L-S, where L indicates the output language (S: Serbian, F: French, E: English, K: Kazakh) and S indicates the script (L: Latin, C: Cyrillic, A: Arabic). The language in parentheses indicates the prompt language.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Lang-script
Total words
Types
Sentences
Kz-ar
27,805,939
810,230
57,827
Kz-cyr
4,656,620
301,614
371,902
Kz-lat
430
371
100
Sr-cyr
15,255
4,424
150
Sr-lat
569,206
89,537
2,919
Cr-cyr
517,595
109,847
77,688
Appendix
Table 1: Summary statistics of datasets
Figure 3: Size of the vocabulary (number of non-character tokens) per language-script combination, for each model. (with Zipf)
Config
gpt-4o-m
gpt-4o
Deepseek-chat
Haiku
Sonnet
Gemini-flash
Gemini
F
C
S
F
C
S
F
C
S
F
C
S
F
C
S
F
C
S
F
C
S
F-C vs F-L
-4
0
-28
-8
-4
-56
-16
-14
-99
-17
-18
-75
-18
-17
-50
-7
-4
-73
-13
-16
-61
S-C vs S-L (S-L)
-3
-9
-10
-2
-5
0
-4
-4
0
-2
-8
9
-6
-8
0
-12
-15
0
0
-6
-3
S-C vs S-L (S-C)
-8
-11
-60
-6
-9
0
-4
-5
-1
6
1
-5
-4
-6
0
-6
-10
-1
-1
-2
-36
S-C vs S-L (E)
6
5
-51
2
-1
-1
0
4
0
-9
-10
0
-3
-4
0
-24
-29
-44
-10
-8
-44
K-C vs K-A (K)
-21
-13
-8
-15
-9
-6
4
5
-1
-9
7
-7
-2
4
-1
10
9
-2
-5
1
-1
Appendix
Table 2: Quality score differences across script configurations of LLMs. F: Factuality, C: Completeness, S: Script fidelity. Lines name are language-script pairs formatted as L-S, where L indicates the output language (S: Serbian, F: French, E: English, K: Kazakh) and S indicates the script (L: Latin, C: Cyrillic, A: Arabic). The language in parentheses indicates the prompt language.
Config
DeepSeek-R1-Distill-Llama-8B
Llama-2-7B-Chat
Mistral-7B-Instruct
Mixtral-8x7B-Instruct
F
C
S
F
C
S
F
C
S
F
C
S
F-C vs F-L
5
2
-99
-1
-2
-100
-4
2
-84
-4
-7
-93
S-C vs S-L (S-L)
0
-4
-42
-5
-3
-34
2
0
68
-2
-3
-36
S-C vs S-L (S-C)
2
2
-63
-4
-2
-99
-11
-10
-99
4
8
-95
S-C vs S-L (E)
1
4
-96
-1
1
-99
-11
-10
-63
2
-2
-89
K-C vs K-A (K)
3
6
-26
0
0
-55
0
0
-85
10
7
11
Appendix
Table 3: Quality score differences across script configurations of SLMs. F: Factuality, C: Completeness, S: Script fidelity. Lines name are language-script pairs formatted as L-S, where L indicates the output language (S: Serbian, F: French, E: English, K: Kazakh) and S indicates the script (L: Latin, C: Cyrillic, A: Arabic). The language in parentheses indicates the prompt language.
Figure 4: Perplexity distributions across script configurations for each SLM. Each panel shows perplexity values for different language-script combinations, with the language of the prompt in parenthesis.
Language Script Type
Number of Languages
Mono-scriptic
78
Multi-scriptic
10
Mixed
9
Appendix
Table 4: Number of languages in different script-based categories based on Tatoeba and literature. Mono-scriptic (multi-scriptic) are those languages for which both sources agree that they use a single (multiple) script(s), respectively. Mixed category languages use different number of scripts (mono in Tatoeba and multi in the literature, or vice versa).
Mixed Category
Multi-script
Category
Bulgarian
Bosnian
Kurdish (Central)
Chavacano
Javanese
Kunming
Korean
Kazakh
Marathi
Kurmanji
Appendix
Table 5: Left: Mixed category languages according to Tatoeba and literature (using different number of scripts in Tatoeba with respect to the literature). Right: Multi-script category containing languages which both Tatoeba and literature consider multi-scriptic.
Table 6: Literature sources (left) explicitly mentioning multi-scriptic feature of a language (right).
Figure 5: Percentage of languages by script usage. Blue bars represent languages for which the model generates text in a single script, while orange bars represent languages for which the model uses two or more scripts. The input groups correspond to the language categories defined in Sub-Section D.1 .
Figure 6: Distribution of output scripts for Latin (right column) and Cyrillic (left column) input prompts for open-source SLMs. Lighter colors indicate higher frequency
Figure 7: Distribution of output scripts for Latin (right column) and Cyrillic (left column) input prompts for proprietary LLMs. Lighter colors indicate higher frequency
Figure 15
Nb
Language prompt
Script prompt
Language Question
Requested language output
Requested script output
1
French
Latin
French
French
Latin
2
French
Latin
French
French
Cyrillic
3
English
Latin
French
French
Latin
4
English
Latin
Kazakh
Kazakh
Cyrillic
5
English
Latin
Kazakh
Kazakh
Latin
6
English
Latin
Kazakh
Kazakh
Arabic
Appendix
Table 7: Detail of configurations for the script forcing experiment (see Section 2 ).
Model Type
Adaptation
Forcing
Mixtral (Quantized)
∼ 15 hrs
∼ 5 hrs
Other SLMs
∼ 7 hrs
∼ 2 hrs
LLMs (via API)
∼ 1-5 hr
≤ 1-3 hrs
Appendix
Table 8: Approximate execution times per 1,000 prompts
Model
Adaptation
Forcing
GPT-4o-mini
∼ $0.5
∼ $1
GPT-4o
∼ $3
∼ $10
DeepSeek-Chat
∼ $0.5
∼ $2
Claude-3.5-Haiku
∼ $1.5
∼ $5
Claude-Sonnet-4
∼ $5
∼ $20
Gemini-Flash
∼ $1
∼ $3
Appendix
Table 9: Approximate API costs for each experimental configuration.