Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from six families on MultiBLiMP, a syntactic minimal-pair benchmark covering 101 languages, using four evaluation methods. We report three principal findings. First, post-training degrades grammatical competence, but the magnitude of this effect is reduced unevenly by model scale, while low-resource languages bear the highest cost. Second, post-trained models retain grammatical knowledge they cannot articulate through explicit prompting, yet this is measurable only in high-resource languages, because near-chance baselines in low-resource settings leave little knowledge to hide. Third, native-language prompting recovers otherwise hidden competence on low-resource languages, demonstrating that only high-resource languages can be probed directly from unprompted probabilities. We conclude that multilingual grammatical evaluation must adopt language-informed, multi-paradigm protocols to avoid systematically underestimating low-resource abilities.
Figures & tables
Evaluation Method
direct
meta
native prompt
English prompt
Gemma 3 (27B)
base
84.17 ± 19.4
73.17 ± 18.7
86.21 ± 18.1
85.77 ± 19.3
PT
79.46 ± 18.9
74.02 ± 17.9
81.63 ± 17.5
79.69 ± 17.1
Gemma 4 (31B)
base
85.56 ± 19.1
84.38 ± 16.1
88.36 ± 16.9
88.13 ± 17.9
PT
76.09 ± 17.2
82.33 ± 18.1
78.06 ± 16.9
79.77 ± 17.0
Llama 3 (8B)
base
80.88 ± 18.4
54.51 ± 9.9
83.24 ± 19.2
82.74 ± 18.7
Table 1: Average model accuracy by evaluation method and model variant. Results are reported for the largest model in each model family, a detailed breakdown of all model scores is provided in Appendix B . Error range denotes standard deviation.
Resource Tier
Low Zero
Mid
High
Super High
Gemma 3 (27B)
base
66.72 ± 22.1
86.60 ± 9.5
96.77 ± 2.3
98.60 ± 0.6
PT
61.70 ± 19.5
81.20 ± 10.6
92.49 ± 4.3
95.15 ± 1.6
Gemma 4 (31B)
base
68.61 ± 22.4
88.22 ± 8.9
97.77 ± 2.0
98.86 ± 0.7
PT
61.08 ± 18.2
76.96 ± 10.7
86.95 ± 6.1
91.79 ± 2.6
Llama 3 (8B)
base
64.96 ± 20.5
82.89 ± 10.7
91.91 ± 5.4
97.11 ± 1.7
Table 2: Average model accuracy by resource tier and model variant, for the direct evaluation method.
Figure 1: Base vs. PT accuracy for Gemma 3 across 101 languages. Each point is one (language, model size) combination; color indicates resource tier. The dashed diagonal marks parity; points below the diagonal indicate linguistic alignment tax.
Figure 2: Normalized linguistic alignment tax for Gemma 3 by resource tier and model size. Shaded bands are 95% bootstrap confidence intervals over the languages in each tier (10,000 resamples). 12 languages where Accbase≤0.5 have been excluded from all sizes.
Figure 3: Normalized linguistic alignment tax for the largest model in each family, measured using the Direct method, and averaged equally across languages with above-chance base accuracy (91–93 of 101, varying by model)
Figure 4: Direct vs. Meta accuracy for Gemma 3 under the PT setting. Each point is one ⟨ language, model size ⟩ combination; color indicates resource tier. The dashed diagonal marks parity; points above the diagonal indicate a positive inarticulate gap.
Figure 5: Mean inarticulate gap for Gemma 3 by resource tier and model size. Gap is computed as AccDirect(PT)−AccMeta(PT) , averaged equally across languages within each tier. Shaded bands are 95% bootstrap confidence intervals over the languages in each tier (10,000 resamples).
Figure 6: Mean inarticulate gap for the largest model in each family, computed as AccDirect(PT)−AccMeta(PT) , averaged equally across 101 languages. Models are ordered by total parameter count.
Figure 7: Prompt Gain set out against language frequency across all evaluated models, which are colored by size in billions of parameters.
Figure 8: Share of items solved by the base model that are lost after post-training, as a function of the base-model margin, for Gemma 4 (31B) under Direct evaluation. Bands are 95% Wilson intervals. Items whose stored margin rounds to zero (0.9%) fall outside the first bin.
Appendix figures & tables12 assets
Supplementary material from the paper’s appendix.
Appendix
Super High
High
Mid
Low/Zero
ISO
Lang
N
ISO
Lang
N
ISO
Lang
N
ISO
Lang
N
deu
German
2298
arb
Arabic
1215
amh
Amharic
112
abk
Abkhazian
40
eng
English
770
bel
Belarusian
2570
bre
Breton
260
aln
Gheg Alb.
677
fra
French
2548
ben
Bengali
21
bua
Buriat
103
apu
Apurinã
28
ita
Italian
2999
bul
Bulgarian
2458
cym
Welsh
1120
aqz
Akuntsu
14
pol
Polish
3272
cat
Catalan
2284
fao
Faroese
232
azz
H-P Nahuatl
207
Appendix
Table 3: All 101 languages in the MultiBLiMP benchmark, organized by resource tier. N = number of minimal pairs. SH = Super High ( > 10 11 CC tokens), H = High (10 9 –10 11 ), M = Mid (10 7 –10 9 ), LZ = Low/Zero ( < 10 7 ).
Evaluation Method
Size
Variant
direct
meta
native prompt
English prompt
Gemma 3
1.0
PT
67.98
54.66
70.52
69.75
base
77.34
49.58
79.52
79.05
4.0
PT
74.43
65.55
75.26
73.97
base
81.38
60.03
82.84
82.49
12.0
PT
77.91
72.16
79.87
77.53
Appendix
Table 4: Model accuracy by method, size, and base/PT variant.
Figure 9: Performance of TinyAya (blue star) against all other models up to 9B parameters. TinyAya demonstrates a strong skew towards the Super High resource languages, and underperforms on low- and mid-resource languages.
Gemma 3
Gemma 4
Llama 3
Qwen 3
Qwen 3.5
TinyAya
ISO
N
1B
4B
12B
27B
E2B
E4B
26B-A4B
31B
1B
3B
8B
0.6B
1.7B
4B
8B
14B
30B-A3B
0.8B
2B
4B
9B
35B-A3B
3.35B
abk
40
57.5
60.0
50.0
80.0
75.0
80.0
57.5
50.0
75.0
70.0
57.5
52.5
62.5
47.5
67.5
55.0
80.0
75.0
70.0
57.5
80.0
42.5
72.5
aln
677
75.3
75.6
80.8
81.8
76.5
79.9
84.5
86.6
70.9
69.9
74.6
70.3
71.3
74.5
69.4
70.0
73.7
68.4
72.1
71.6
75.2
75.6
74.2
amh
112
92.9
94.6
98.2
98.2
96.4
92.9
99.1
92.9
94.6
94.6
99.1
90.2
90.2
88.4
96.4
97.3
97.3
86.6
95.5
92.9
86.6
96.4
92.0
apu
28
92.9
96.4
96.4
96.4
96.4
92.9
92.9
96.4
96.4
92.9
96.4
96.4
92.9
96.4
96.4
92.9
96.4
92.9
96.4
100.0
96.4
96.4
92.9
aqz
14
35.7
28.6
21.4
21.4
35.7
42.9
14.3
28.6
21.4
42.9
35.7
42.9
21.4
21.4
57.1
28.6
42.9
42.9
42.9
42.9
28.6
28.6
35.7
Appendix
Table 5: Mean Direct accuracy (%) — Base models.
Gemma 3
Gemma 4
Llama 3
Qwen 3
Qwen 3.5
TinyAya
ISO
N
1B-it
4B-it
12B-it
27B-it
E2B-it
E4B-it
26B-A4B-it
31B-it
1B-it
3B-it
8B-it
0.6B-it
1.7B-it
4B-it
8B-it
14B-it
30B-A3B-it
0.8B-it
2B-it
4B-it
9B-it
35B-A3B-it
3.35B-it
abk
40
75.0
30.0
47.5
72.5
62.5
65.0
72.5
60.0
72.5
75.0
72.5
42.5
45.0
50.0
72.5
50.0
70.0
75.0
62.5
62.5
67.5
52.5
60.0
aln
677
62.3
62.5
75.3
74.7
65.0
68.0
65.6
67.4
68.1
66.3
66.9
68.0
68.5
69.4
70.3
66.9
68.2
69.3
73.0
73.1
74.2
77.2
71.3
amh
112
83.0
87.5
88.4
93.8
93.8
95.5
92.9
91.1
94.6
91.1
98.2
94.6
92.9
95.5
97.3
93.8
97.3
82.1
92.0
93.8
97.3
94.6
89.3
apu
28
57.1
85.7
82.1
71.4
89.3
75.0
71.4
64.3
96.4
85.7
96.4
96.4
92.9
100.0
96.4
96.4
96.4
96.4
96.4
96.4
96.4
96.4
85.7
aqz
14
28.6
50.0
28.6
28.6
21.4
42.9
50.0
28.6
28.6
21.4
50.0
28.6
35.7
35.7
28.6
50.0
21.4
28.6
50.0
21.4
57.1
57.1
50.0
Appendix
Table 6: Mean Direct accuracy (%) — PT models.
Family
N
βCC
p
R2
Gemma 3
356
0.0610
<0.001
0.651
Gemma 4
356
0.0613
<0.001
0.630
Llama 3.1
89
0.0578
<0.001
0.629
Llama 3.2
178
0.0535
<0.001
0.569
Qwen 3
534
0.0552
<0.001
0.535
Qwen 3.5
445
0.0555
<0.001
0.591
Appendix
Table 7: Within-family OLS: Accuracy∼log10(CCcount)+Size_Order . Size_Order is the rank of the model size within its family, starting at 0. All models are base, Direct method.
Figure 10: Direct accuracy against the mean log-probability difference between the grammatical and ungrammatical sentence, one point per language, for Gemma 3 (27B, base). Colour indicates resource tier. The dashed line marks equal probability.
Tier
Items
p10
p25
Median
p75
p90
Near-tie
Super High
18,260
3.5
5.5
8.0
11.0
15.0
2.6%
High
65,010
2.0
4.5
7.0
10.0
14.0
5.1%
Mid
23,936
−2.0
1.0
5.0
8.5
14.0
12.9%
Low/Zero
14,099
−5.0
0.0
4.0
8.0
14.0
14.2%
Appendix
Table 8: Distribution of the per-item log-probability difference logP(Scor)−logP(Sinc) by resource tier, for Gemma 3 (27B, base) under Direct evaluation. Percentiles are taken over items. The last column is the share of items on which the two sentences are separated by at most one natural-log unit in either direction.
Family
Super High
High
Mid
Low/Zero
All
Gemma 3 (27B)
3.45
4.28
5.40
5.02
4.71
Gemma 4 (31B)
7.07
10.81
11.27
7.53
9.47
Llama 3 (8B)
1.80
2.03
1.25
1.54
1.69
Qwen 3 (30B-A3B)
2.38
2.30
3.05
1.68
2.23
Qwen 3.5 (35B-A3B)
0.94
1.20
1.38
0.63
1.02
TinyAya (3.35B)
4.40
2.90
2.95
0.26
2.07
Appendix
Table 9: Raw linguistic alignment tax, Accbase−AccPT in percentage points, by resource tier for the largest model in each family (Direct method). Unlike the normalized tax, no languages are excluded; values are averaged equally across the languages in each tier (7, 38, 20 and 36 languages) and across all 101 languages in the last column.
Dropped
n
Smallest N
1B
4B
12B
27B
0
24
26
35.9 [24.2, 47.4]
33.0 [21.2, 43.7]
12.0 [ − 10.7, 28.1]
20.6 [12.8, 28.1]
1
23
28
36.7 [24.8, 48.5]
32.7 [20.6, 44.2]
10.5 [ − 12.4, 26.9]
20.1 [12.2, 27.9]
2
22
34
34.6 [23.1, 46.4]
33.2 [20.5, 44.7]
9.5 [ − 14.2, 26.9]
18.6 [10.9, 26.0]
3
21
50
32.0 [20.9, 43.2]
37.5 [28.8, 47.1]
8.4 [ − 17.6, 26.7]
17.5 [ 0 9.6, 25.2]
4
20
86
34.2 [23.2, 45.0]
39.6 [31.7, 48.7]
8.4 [ − 18.3, 27.3]
17.3 [ 0 9.2, 25.4]
5
19
99
35.2 [23.3, 46.4]
40.1 [31.9, 50.0]
7.9 [ − 20.7, 27.7]
18.7 [10.3, 26.8]
Appendix
Table 10: Low/Zero normalized alignment tax for Gemma 3 (Direct method) after repeatedly dropping the language with the smallest test set. Row k uses the 24−k Low/Zero languages with the largest test sets, so every row is a subset of the row above it; the models and the items are otherwise unchanged. “Smallest N ” is the size of the smallest test set still included. Cells are the mean over the retained languages with a 95% bootstrap confidence interval (10,000 resamples). The 12 Low/Zero languages with Accbase≤0.5 are excluded from all rows, as in § 4.3 .
Family
Super High
High
Mid
Low/Zero
Gemma 3 (27B)
0.038
0.044
0.094
0.047
Gemma 4 (31B)
−0.066
−0.091
−0.044
−0.042
Llama 3 (8B)
0.233
0.272
0.292
0.120
Qwen 3 (30B-A3B)
0.074
0.088
0.133
0.085
Qwen 3.5 (35B-A3B)
0.003
0.004
0.020
0.039
TinyAya (3.35B)
0.277
0.244
0.209
0.129
Appendix
Table 11: Mean inarticulate gap by resource tier and family (largest model only).
Resource Tier
Size
Variant
Low Zero
Mid
High
Super High
Gemma 3
1.0
PT
3.50
15.32
19.83
22.85
base
12.13
23.98
40.97
47.25
4.0
PT
6.20
12.08
9.91
7.93
base
11.52
28.72
25.67
27.41
12.0
PT
4.49
10.54
4.32
6.22
Appendix
Table 12: Mean inarticulate gap (Acc Direct− Acc Meta , %) by resource tier for all model variants. Bold=lowest gap per column.