cs.CLOct 6, 2026

Are Language Models Script-Aware?

Authors: David Kletz, Sandra Mitrović, Ljiljana Dolamić, Fabio Rinaldi

Organizations: SUPSI, IDSIA, Switzerland · armasuisse, Science & Technology, Switzerland

Abstract

Language models frequently generate outputs in unintended languages or scripts, a phenomenon known as off-target generation. While existing research has focused on language selection, the dimension of script knowledge remains understudied: before any linguistic understanding can occur, users must recognize the graphic symbols in a model's response. We investigate whether Small and Large Language Models (SLMs and LLMs) possess script knowledge by testing them on multi-scriptic languages. Through two complementary experiments, we evaluate whether models (1) adapt their output script to match the input, and (2) follow explicit instructions to generate text in a specified script. The models we tested demonstrate substantial script knowledge: they all achieve a near-perfect Latin script fidelity (more than 98%) and follow script instructions with high frequency. Nevertheless, we notice differences between LLMs and SLMs, with higher scores for LLMs including for non-standard script combinations.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Latin Substrate: How Language Models Represent and Mediate Script Choice

    May 29, 2026Daniil Gurgurov, Alan Saji, Katharina Trinley +2Representational CapacityChoice

  2. Script Choice in LLMs: Evidence for Late-Layer Commitment

    Sep 23, 2026David Kletz, Sandra Mitrović, Itay Sabato +2Cryptographic Merkle Tree CommitmentsCommitment

  3. The Digital Afterlife of Empires: Four Language Models Converge on the Same Imperial Cartography of Writing

    Apr 30, 2026Hiroki FukuiWritingAsymmetry