In this paper, we investigate how script knowledge is distributed across the layers of LLMs using two complementary interpretability methods: logistic regression probing and logit-lens analysis. Our probing experiments reveal a clear asymmetry: both the input script and the instructed output script are encoded in the earliest layers of the network, while, in contrast, commitment to the actual output script emerges only in the final layers, with the model's intermediate representations defaulting to Latin throughout most of the layers. This two-stage process is confirmed by logit-lens analyses, which show that script commitment consistently occurs at the very last layers of the LLMs. Together with the weaker script-following performance observed in smaller models, these results form a converging body of evidence linking script commitment to model depth, with broader implications for the design of sufficiently deep, inclusive multilingual architectures.
Figures & tables
Figure 1: Balanced accuracy of the requested-output script probe across layers, for all six Qwen 2.5 model sizes.
Figure 2: Balanced accuracy of the used-output script probe across layers.
Figure 3: Average first layer of script commitment by target script and model size. Error bars show standard deviation across prompts. Hollow profiles show the maximum each model can reach (number of layers).
Layer
Prediction
h_out
’ К’
h h34 _out
’ К’
h h33 _out
’-capital’
h h32 _out
’-capital’
h h31 _out
’-capital’
Table 1: Logit-lens top-1 prediction at the generated token position across the last 5 layers of Qwen2.5-3B-Instruct when asked to write “Capital” using the Cyrillic script. The ’ К’ character is a Cyrillic character.
Figure 4: Probability mass assigned to tokens of Latin, target and any other script across layers.
Appendix figures & tables6 assets
Supplementary material from the paper’s appendix.
Appendix
Model
Params
Layers
Hidden dim
Attn / KV heads
Qwen2.5-0.5B-Instruct
∼ 0.5B
24
896
14 / 2
Qwen2.5-1.5B-Instruct
∼ 1.5B
28
1536
12 / 2
Qwen2.5-3B-Instruct
∼ 3B
36
2048
16 / 2
Qwen2.5-7B-Instruct
∼ 7B
28
3584
28 / 4
Qwen2.5-14B-Instruct
∼ 14B
48
5120
40 / 8
Qwen2.5-32B-Instruct
∼ 32B
64
5120
40 / 8
Appendix
Table 2: Qwen2.5 models used.
Model
Strict
Lax
Qwen2.5-0.5B-Instruct
0.641
0.797
Qwen2.5-1.5B-Instruct
0.719
0.846
Qwen2.5-3B-Instruct
0.831
0.920
Qwen2.5-7B-Instruct
0.799
0.898
Qwen2.5-14B-Instruct
0.766
0.838
Qwen2.5-32B-Instruct
0.754
0.833
Appendix
Table 3: Script-following rates.
Figure 5: Balanced accuracy of the input-language probe across layers, for all six Qwen models.
Figure 6: Per-script breakdown across all six Qwen models. Left: correct-script generation rate by target script. Center: output distribution by target script, broken down into correct, wrong, and neutral outputs. Right: average first commitment layer by target script for correct outputs only.
Figure 7: Logit margin between the top-1 token and the highest-ranked token from different script, across layers.
Figure 8: Breakdown of wrong-script generations by the script actually produced, per model.
Many languages are written in multiple scripts, requiring large language models (LLMs) to generate equivalent linguistic content in distinct orthographic forms. While prior work suggests that LLMs route information through shared latent representations, how they internally mediate script variation remains poorly understood. We study this question by first examining per-layer output distributions with the logit lens, which reveals consistent latent romanization during transliteration, and then through representational and mechanistic analyses of script generation. At the representational level, we show that scripts of the same language become increasingly separable across layers and that a simple linear steering direction can flip a model's output script while largely preserving semantic content. The vector generalizes to writing systems unseen during construction: it flips non-Latin output to Latin, but maps Latin output into varied non-Latin scripts. At the mechanistic level, we localize a small set of late-layer attention heads that causally mediate script choice. These heads transfer across unrelated languages and writing systems, suggesting that script routing is implemented by language-agnostic components. Across both analyses, we observe a consistent directional asymmetry: non-Latin output is produced by a compact, identifiable gate, while Latin-script output emerges from diffuse contributions across the model. Overall, our findings hint that LLMs organize script variation around shared latent representations while exhibiting a privileged substrate toward Latin script.
Daniil Gurgurov, Alan Saji, Katharina Trinley +2
D German Research Center for AI (DFKI) · I Indian Institute of Technology, Madras · H University of Hamburg +1
Large language models (LLMs) perform inference by following a fixed depth and order, non-recurrent execution of all layers. We reveal the wide existence of training-free, flexible, dynamic program-of-layers (PoLar), where pretrained layers can be packed as modules and then skipped or looped to form a customized program for each input. For most inputs, substantially shorter program executions can achieve the same or better accuracy, while incorrect predictions of the original LLM can be corrected by alternative programs with fewer layers. These observations indicate that inference admits multiple valid latent computations beyond the standard forward pass. To efficiently achieve PoLar in practice, we propose a lightweight PoLar prediction network, which learns to generate execution programs that dynamically skip or repeat pretrained layers for each input. Experiments on mathematical reasoning benchmarks demonstrate that PoLar consistently improves accuracy over standard inference and prior dynamic-depth methods, often while executing fewer layers, and that these gains persist under out-of-distribution evaluation. Our results suggest that fixed-depth execution captures only a narrow subset of an LLM's latent reasoning capacity.
Ziyue Li, Yang Li, Tianyi Zhou
University of Maryland, College Park, MD, USA · 2MBZUAI, Abu Dhabi, UAE.
A language model can fail a syntactic test in two distinct ways: by not encoding the relevant structure, or by encoding it but failing to use it at the output. Behavioral evaluation alone cannot tell these apart. We propose a three-level evaluation framework (behavioral deployment, LM-head readout, and probe recoverability) measured on the same items under the same binary decision. Using a compact trilingual (English, Chinese, German) control-dependency benchmark, we find that probe recoverability exceeds or equals LM-head readout, which in turn exceeds or equals behavioral deployment, across seven models and all three languages in the aggregate. The recoverability surplus is never negative across all 14 (model, task) conditions. The disconnect concentrates in subject-control, where a nearest-noun heuristic gives the wrong answer. The single largest gap (0.653) appears on Qwen3-0.6B Instruct in question answering. The gap persists at Qwen3-14B Instruct. Instruction tuning degrades deployment more than encoding in percentage terms. We rule out option-position bias, late-layer erasure, output-formatting artifacts, and probe-training variance. The pattern is consistent with decoding that favors surface shortcuts, and the behavior-probe gap measures the strength of that preference. Activation patching shows the gap is layer-localized. Under instruction tuning, the LM-head-decoded layer shifts approximately ten layers later than the probe-decoded layer. These findings argue that behavioral evaluation understates what models encode, while probing alone overstates what they deploy.
Zhenyan Lu, He Wang, Xiaohui Huang
College of International Studies, National University of Defense Technology, Nanjing, China