Large language models exhibit strong multilingual capabilities, yet significant performance gaps persist between dominant and non-dominant languages. Prior work attributes this gap to imbalances between shared and language-specific neurons in multilingual representations. We propose Cross-Lingual Activation Steering (CLAS), a training-free inference-time intervention that selectively modulates neuron activations. We evaluate CLAS on classification and generation benchmarks, achieving average improvements of 2.3% (Acc.) and 3.4% (F1) respectively, while maintaining high-resource language performance. We discover that effective transfer operates through functional divergence rather than strict alignment; performance gains correlate with increased language cluster separation. Our results demonstrate that targeted activation steering can unlock latent multilingual capacity in existing models without modification to model weights.
Figures & tables
Figure 1: Overview of CLAS, with two high level steps (1) Neuron Categorization and (2) Activation Steering.
Figure 2: Distribution of types of neuron per layer across models. Llama has a total of 32 layers and Qwen has 28 layers. The bridge layers (purple shade) are selected near the final layers where the partial-shared language activations are the highest.
Llama
Qwen
Base
Int μ
INCLINE
CLAS (Ours)
Base
Int μ
INCLINE
CLAS (Ours)
XNLI
42.42
42.19
39.19
45.03 (+2.61)
59.08
59.05
41.34
59.29 (+0.21)
XStoryCloze
83.38
82.97
79.05
85.58 (+2.20)
78.68
78.96
61.08
79.70 (+1.02)
PAWS-X
65.70
62.87
59.58
66.00 (+0.30)
78.58
79.03
68.30
79.07 (+0.49)
HarryPotter
51.59
47.84
50.75
52.59 (+1.00)
53.67
49.58
28.42
55.50 (+1.83)
XQuAD
26.85
26.98
33.50
32.10 (+5.25)
32.73
21.78
18.38
33.94 (+1.21)
Table 1: The performance on different benchmarks across Llama and Qwen. The best performing model is in bold . The improvements of CLAS over base are shown within parentheses.
en
ar
es
eu
hi
id
my
ru
sw
te
zh
Avg.
Llama
Base
92.26
88.09
91.07
70.02
90.07
87.49
74.26
91.40
71.01
79.68
90.73
83.38
Int μ
92.39
89.01
91.33
71.74
86.30
84.98
74.39
91.79
71.08
78.89
90.14
82.97
INCLINE
92.39
90.21
95.10
67.77
84.58
88.88
52.75
92.52
71.14
52.81
94.77
79.05
CLAS
92.26
85.44
94.77
74.92
89.61
92.12
73.13
94.04
78.09
78.95
94.71
85.58
Qwen
Table 2: Performance on XStoryCloze (accuracy) across multiple languages and models. For each language, bold represents the cases where CLAS beats the base model.
en
ar
de
el
es
hi
ro
ru
th
tr
vi
zh
Avg.
Llama
Base
23.47
23.20
38.54
25.18
34.23
29.47
30.01
27.32
19.23
23.74
34.91
12.84
26.85
Int μ
23.47
21.25
33.08
26.68
35.31
29.39
36.98
27.25
15.83
21.64
40.00
12.90
26.98
INCLINE
34.96
20.39
45.79
35.68
44.88
34.02
34.13
21.87
26.87
44.18
36.01
23.19
33.50
CLAS
23.47
27.03
50.86
23.71
50.57
28.32
41.84
34.87
21.90
30.38
37.49
14.77
32.10
Qwen
Table 3: The performance (F1 scores) on the XQuAD dataset across multiple languages and models. For each language, bold represents the cases where CLAS beats the base model.
Figure 3: Cosine similarity with English across languages on each task using Llama model.
Figure 4: t-SNE visualization of multilingual representations before (left) and after (right) CLAS intervention. Each point represents a sentence embedding, with colors denoting languages. English, marked with ( red ) stars, is the anchor language.
Figure 5: Impact of α across XStoryCloze and XQuAD. The left column shows the performance for the Llama ( blue ) and the right column for the Qwen ( pink ).
Question
Llama (base)
CLAS
Was benannte die Generalkonferenz für Maß und Gewicht 1960 nach Tesla?
SI-Maßeinheit der magnetischen Flussdichte. stop. Die SI-Maßeinheit der magnetischen Flussdichte
SI-Maßeinheit der magnetischen Flussdichte.
(What did the General Conference on Weights and Measures name after Tesla in 1960?)
(SI unit of magnetic flux density. stop. The SI unit of magnetic flux density)
(SI unit of magnetic flux density.)
Die Lutherbibel beeinflusste wessen englische Übersetzung der Bibel?
William Tyndale Die Lutherbibel beeinflusste die englische Bibel von William Tyndale. Tyndale war ein eng
William Tyndale.
(The Luther Bible influenced whose English translation of the Bible?)
(William Tyndale’s English Bible was influenced by the Luther Bible. Tyndale was a close)
(William Tyndale.)
Table 4: Comparison of outputs for German XQuAD samples. English translations provided via Google Translate.
Figure 6: Multilingual performance comparison across different Llama layers (early layers 0-6, middle layers 6-16, and bottom layers 24-30).
Appendix figures & tables5 assets
Supplementary material from the paper’s appendix.
Appendix
Dataset
∣L∣
#Samp.
Metric
XNLI (natural language inference)
15
5010
Acc.
XStoryCloze (commonsense reasoning)
11
1511
Acc.
PAWS-X (paraphrase identification)
4
2000
Acc.
HarryPotter (MCQA)
5
300
Acc.
XQuAD (question answering)
12
1190
F1
Appendix
Table 5: Summary of evaluation datasets. #Samp. represents the number of samples per language.
en
ar
bg
de
el
es
fr
hi
ru
sw
th
tr
ur
vi
zh
Avg.
Llama
Llama
52.63
38.88
42.28
43.65
40.58
46.95
44.53
41.78
43.75
41.10
40.68
44.17
36.09
44.81
44.63
42.42
Int μ
52.63
38.12
40.68
45.77
41.3
46.59
43.11
41.10
43.79
41.42
44.39
43.81
36.37
42.85
41.38
42.19
INCLINE
52.63
39.24
38.74
35.77
38.08
38.62
35.75
44.25
33.51
35.51
42.87
39.60
33.35
46.89
46.43
39.19
CLAS
52.63
39.10
41.68
49.42
44.85
50.38
47.96
41.74
45.27
38.36
40.98
49.36
44.25
46.41
50.72
45.03
Qwen
Appendix
Table 6: The performance on XNLI across multiple languages and models.
en
de
es
fr
Avg.
Llama
Base
68.20
63.45
64.45
69.20
65.70
Int μ
68.20
62.20
54.70
71.70
62.87
INCLINE
69.00
66.45
54.55
57.75
59.58
CLAS
68.20
66.60
65.70
65.70
66.00
Qwen
Appendix
Table 7: Performance on the PAWS-X across multiple languages and models.
en
de
fr
it
es
Avg.
Llama
Base
65.00
50.67
50.67
48.33
56.67
51.59
Int μ
65.33
53.67
44.33
46.67
46.67
47.84
INCLINE
61.00
52.00
49.00
48.67
53.33
50.75
CLAS
65.00
53.67
52.00
47.67
57.00
52.59
Qwen
Appendix
Table 8: Accuracy on the multilingual HarryPotter Quiz across multiple languages and models.
Figure 7: Impact of the hyperparameter α on performance across datasets. The left column shows the performance for the Llama model, while the right column shows the performance for the Qwen model.