sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak
Authors: Marek Šuppa, Ivan Vykopal, Andrej Ridzik, Kristián Sopkovič, Natália Kňažeková, Jaroslav Kopčan, Miroslav Blšták, Viktória Ondrejová, +4 more
Organizations: Comenius University in Bratislava, Slovakia · Cisco Systems · Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia · Brno University of Technology, Czechia · Technical University of Košice, Slovakia
Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data (ρ≥0.98), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently (ρ=0.72). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench
Figures & tables
Figure 1: Overview of the sk-bench design: ten skill categories, 30 evaluation tasks. Chips with a gold fill, heavy gold border and a leading star ( ⋆ ) are the 11 resources introduced or first packaged for generative-LLM evaluation by sk-bench, listed first within each panel. Plain chips with a hairline border are the 19 pre-existing datasets we integrate into the harness. Novelty (the chip itself) and provenance (the N / MT / MT ∗ / G badge at the right: native Slovak, translated from English, translated then adapted with a custom Slovak registry, LLM-generated over a Slovak source) are independent of each other—MultiWikiQA-SK is generated but pre-existing, Autoškola is native and new. Tasks with bilingual templates are run with both English and Slovak prompts. IFEval-SK and Slohry are Slovak-only. These ten skill categories organise the datasets. The results tables instead use six reporting groups defined by answer format (Table 1 , § 6 ).
Model
Knowledge & reasoning
Reading & QA
Classification
Generation
Instruction following
Grammar & morphology
Overall
< 2B
Granite 4.0-H 1B
39.8 ± 0.6
28.2 ± 0.6
50.1 ± 7.8
14.8 ± 0.9
14.9
6.0 ± 0.8
25.6 ± 1.8
Qwen3 1.7B
35.0 ± 5.5
30.9 ± 2.8
52.4 ± 7.8
12.8 ± 0.9
23.5
12.3 ± 1.9
27.8 ± 2.4
2–9B
Ministral 3 8B
60.4 ± 1.3
39.2 ± 6.2
70.6 ± 4.4
13.5 ± 0.8
24.0
37.1 ± 1.0
40.8 ± 2.3
Ministral 3 8B (R)
54.9 ± 1.9
37.8 ± 4.0
61.9 ± 7.2
13.1 ± 1.1
39.9
32.1 ± 2.8
39.9 ± 2.4
Table 1: No open-weights model reaches the proprietary APIs on Slovak. Main sk-bench leaderboard (Slovak, 0-shot): top two open-weights models per size tier and top three proprietary APIs. Cells are the unweighted mean of the group’s per-task primary metrics ( ×100 ). Best and second-best are over all 55 models, so a mark appears only where a shown model holds the record. Overall averages the six groups. Grey ± is the standard deviation across the three prompt templates. Full leaderboard and per-task scores: Tables 14 and 15 .
Figure 2: Open-weights Slovak scores plateau past ∼ 20B: overall sk-bench score vs. parameter count (log scale). The frontier (blue) flattens past gpt-oss 20B/Qwen3.5 27B—no larger open model, up to the 1.1T Kimi K2.6, beats gpt-oss 120B (60.0). The dashed line marks the best closed model (Gemini 3.1 Pro, 72.6). A fully-labelled version is in Appendix H (Figure 7 ).
Within parameter band
Contrast
Overall
< 2B
2–9B
9–32B
32–200B
> 200B
Prop.
top-10
NLI (N vs MT)
0.99
0.93
0.99
0.98
0.96 †
0.90 †
0.94 †
10/10
MCQ (N vs MT)
0.98
0.93
0.96
0.98
0.93 †
0.40 †
0.71 †
7/10
QA (N vs G)
0.72
0.17
0.43
0.74
0.64 †
0.80 †
0.94 †
6/10
Table 2: Translation is a safe ranking proxy for closed-form Slovak skills, but question provenance is not. Spearman ρ between the native and translated/generated sk-bench leaderboards—overall, within parameter-size bands, and as top-10 rank overlap—over all 55 models (zero-shot, Slovak prompts, accuracy for NLI/MCQ, QEM for QA). Overall 95% bootstrap CIs: NLI [0.97,1.00] , MCQ [0.94,0.99] , QA [0.57,0.83] . † band has n≤7 , so its within-band ρ is high-variance (wide CIs) and read only qualitatively. N=native, MT=translated, G=LLM-generated.
Figure 3: Test-time reasoning transfers to Slovak, and concentrates on instruction following. Per-group sk-bench change (Slovak, 0-shot) from thinking mode for the five ≥ 9B models, each gaining +8.5 to +12.5 overall. The effect is flat on free-form generation.
Appendix figures & tables21 assets
Supplementary material from the paper’s appendix.
Appendix
Figure 4: RQ1 data-provenance contrasts: each point is one of the 55 models, scored on the native source ( x ) and its translated or LLM-generated counterpart ( y ). The dashed line is y=x . Points hug the diagonal for NLI and multiple choice (translation preserves the ranking) but scatter off it for QA (human-authored vs. generated questions disagree), concentrated among weaker models. Per-tier ρ and top-10 overlap are in Table 2 .
Overall
Δ by group
Lineage
Base
+Alp.
Δ
Know.
Class.
Instr.
Gram.
Mistral-SK 7B
19.0
14.9
-4.0
-1.1
-16.9
-0.4
+4.9
Qwen3-SK 14B
33.5
43.8
+10.3
+26.6
+16.5
+13.2
-19.8
TildeOpen-SK 30B
31.9
32.8
+0.8
-0.6
-29.2
+11.5
+30.7
Appendix
Table 3: Effect of Slovak Alpaca-style instruction-tuning on three base models (Slovak, 0-shot), sorted by size. Base and +Alp. are the overall macro-average of the base Slovak LM and its instruction-tuned counterpart. The Δ columns are tuned − base for the overall score and the four most affected reporting groups (Knowledge & reasoning, Classification, Instruction following, Grammar & morphology). The net effect rises with base capability—negative for the weakest base (in bold), flat for the mid one, clearly positive for the strongest—while instruction following (the tuned skill) improves or holds for all three and the other groups swing in both directions.
Figure 5: The Qwen3-14B lineage: base, after Slovak continued pretraining, and after the subsequent Slovak instruction-tuning step, per reporting group (overall score beside each label). Adaptation trades broad capability for a single grammar gain. Instruction tuning recovers knowledge, reading, and instruction following, and reverses the grammar gain. Scores are the six group averages from Table 14 .
Figure 6: RQ2: per-reporting-group scores (Slovak, 0-shot, axes 0 – 100 ) of three Slovak base checkpoints before (Slovak base) and after ( + Alpaca) a small Slovak instruction-tuning step. For Qwen3-SK, whose upstream multilingual model (Qwen3-14B) was also evaluated, the dashed polygon shows that released model before Slovak adaptation. The upstream bases of Mistral-SK and TildeOpen-SK are not in the sweep. The effect varies by base and by group: Alpaca broadly lifts Qwen3-SK (knowledge, reading, classification, instruction-following) at the cost of grammar, is mixed for TildeOpen-SK, and mostly shrinks the weakest (Mistral-SK), and Slovak adaptation itself drops Qwen3-14B before Alpaca partly restores it (§ 6.2 , with per-group deltas in Table 3 ).
Overall
Δ by group
Model
Std
Think
Δ
Instr.
Gram.
Gen.
Qwen3 0.6B
21.5
27.2
+5.7
+19.6
+1.1
-0.3
Qwen3.5 0.8B†
24.0
7.2
-16.8
+10.1
+3.6
-9.4
Qwen3 1.7B
27.8
40.0
+12.2
+23.8
+10.2
-1.0
Qwen3.5 2B†
29.4
19.8
-9.7
+7.1
+12.5
-10.8
Qwen3.5 9B
47.9
56.4
+8.5
+21.2
+28.8
-3.5
Appendix
Table 4: Effect of test-time reasoning (“thinking” mode), Slovak, 0-shot, for the models evaluated in both modes (sorted by size). Std and Think are the overall macro-average. The Δ columns are thinking − standard for the overall score and for the three most affected reporting groups (Instruction following, Grammar & morphology, Generation). Reasoning runs under an 8,192-token generation cap. The two smallest models (†) exhaust it inside the reasoning trace without emitting a parseable answer, so their answer-extraction groups collapse while instruction and grammar still improve (negative overall Δ in bold). Full per-task thinking scores are in Table 16 .
Reporting group
SK
EN
Δ
Knowledge & Reasoning
58.1
58.3
+0.1
Reading & QA
39.4
41.0
+1.6
Classification
61.3
61.1
-0.2
Generation
13.2
13.1
-0.1
Grammar & Morphology
33.5
30.9
-2.6
Overall
49.9
50.3
+0.4
Appendix
Table 5: Prompt-language effect: mean Slovak- vs. English-prompt score per reporting group, macro-averaged across all evaluated models over the tasks available in both languages (IFEval-SK and Slohry are Slovak-only and excluded). Δ=EN−SK . A positive value means the English prompt scores higher. Full per-task English scores are in Table 17 .
skLEP
BCM
MERA
HuLU
INCL
sk-bench
Language
sk
cs
ru
hu
multi
sk
Generative LLM eval
✓
✓
✓
✓
Native paired contrast
✓
Open + closed, one harness
∼
✓
Grammar & morphology
∼
✓
Instruction-following
∼
✓
Appendix
Table 6: How sk-bench differs from the closest Slavic, Central-European, and multilingual LLM benchmarks. The lower block lists the design choices that set sk-bench apart. Cells mark what each benchmark actually provides ( ∼ = partial). BenCzechMark (BCM) and MERA score by log-likelihood , which closed APIs do not expose, so neither ranks open- and closed-weights models on one leaderboard. sk-bench’s generate-and-match scoring does. INCL = INCLUDE. Of sk-bench’s 30 tasks, 11 are resources introduced or first packaged here (Table 10 ). We do not tabulate the same count for the other suites, whose new-versus-integrated split we cannot establish uniformly from their papers.
Context
SK
EN
Δ
4,096
58.1
60.5
-2.4
8,192
55.6
56.4
-0.7
16,384
54.4
53.4
+1.1
Appendix
Table 7: OneRuler-SK by context length: mean retrieval accuracy ( ×100 , 0-shot) over the RULER-style subtasks at 4,096 / 8,192 / 16,384 tokens, macro-averaged across all evaluated models, for the matched Slovak and English haystacks. Because SK and EN share the same model and the same length, Δ=SK−EN isolates the language effect from the context-window limit. The gap does not widen with length—it closes and slightly reverses by 16k—so up to 16k there is no Slovak-specific long-context penalty. 16k is the longest length evaluated.
Slovak
English
Reporting group
0-shot
3-shot
0-shot
3-shot
Knowledge & reasoning
60.6
59.4
61.2
60.4
Reading & QA
29.4
32.6
30.7
33.6
Classification
61.9
66.8
61.3
67.0
Grammar & morphology
35.7
36.8
28.6
27.0
Overall
46.9
48.9
45.4
47.0
Appendix
Table 8: Zero-shot vs. 3-shot, macro-averaged over the 52 models with few-shot results, for the four reporting groups of the main leaderboard (Table 1 ) that have a few-shot counterpart. Generation (summarisation) and instruction following were run zero-shot only and are therefore omitted. The Overall row is the macro-average over these four groups. Few-shot helps most on classification and reading & QA, slightly hurts knowledge & reasoning, and moves the overall little. The pattern is the same under Slovak and English prompts.
Model
Input Tokens [M]
Completion Tokens [M]
Total Tokens [M]
Cost
Evaluation Time [hours]
Open-source Models
ibm-granite/granite-4.0-h-350m
457.03
7.27
464.29
N/A
4.17
ibm-granite/granite-4.0-350m
457.03
10.83
467.85
N/A
1.97
Qwen/Qwen3-0.6B
439.36
14.13
453.49
N/A
2.49
Qwen/Qwen3.5-0.8B
369.88
13.04
382.92
N/A
3.16
google/gemma-3-1b-it
361.46
8.33
369.79
N/A
2.22
Appendix
Table 9: Resource consumption and evaluation time across models on the sk-bench benchmark. Token counts are shown in millions (M). * denotes models evaluated only in the zero-shot setting due to API constraints. Evaluation time is measured in hours. “—” indicates data not available or not applicable.
Dataset
∣test∣
Prov.
Skill
Licence
Description
SentiSK
4,677
N
sentiment
CC BY-SA 4.0
Slovak user-text sentiment corpus used in binary and 3-class settings. Neutral examples are excluded from the binary score.
MovieReview
2,008
N
sentiment
—
Slovak film- and series-review sentiment with labels derived from user star ratings ( Štefánik et al., 2023 ) , included as a second native binary sentiment source. See Appendix D.3 .
skLEP-sentiment
1,042
N
sentiment
Mixed
The sentiment subtask from skLEP, retained as a small validation cross-check against the larger sentiment resources.
ToxicSK
881
N
safety classification
CC BY-SA 3.0
Slovak toxicity-detection corpus scored as yes/no classification.
skLEP-hate
1,319
N
safety classification
Mixed
skLEP hate-speech subtask, paired with ToxicSK to separate toxicity from explicit hate-speech labelling.
NLI-SK-annotated
597
N
NLI
—
Human-annotated Slovak 3-way entailment dataset after filtering the non-NLI “Other” category.
Appendix
Table 10: Dataset catalogue for sk-bench. ∣test∣ is the evaluation test-set size used by the released harness. Prov. : N=Slovak-native, MT=translated from English, MT ∗ =translated and adapted with a custom Slovak instruction registry, G=generated over Slovak source documents. Resources marked ★ are the eleven introduced or first packaged with sk-bench, at three levels of contribution: built from scratch (IFEval-SK, Chiby, Slohry, Slovak Trivia, OneRuler-SK), newly curated from public institutional sources (Autoškola, e-testy-sk, Slovak-tests, Scio), and existing expert resources first packaged for generative-LLM evaluation (SKJ1, Word Analogies). Licence is the licence declared on the source artefact (“Mixed” = the bundled release carries per-task licences; “—” = none declared).
Skill
Native side
Contrast side
NLI
NLI-SK-annotated
skLEP-NLI (MT)
Multiple-choice
e-testy-sk ★ ,
Okapi-MMLU-SK (MT)
knowledge
Slovak-tests ★
Question answering
skLEP-QA (human)
MultiWikiQA-SK (G)
Grammar & morph.
Chiby ★ , SKJ1 ★ ,
translation cannot
Slohry ★
produce these
Appendix
Table 11: The paired provenance contrasts that structure the suite. Each row holds a skill fixed and varies how the evaluation data was produced. ★ marks a resource introduced or first packaged with sk-bench. MT = machine-translated, G = LLM-generated. The grammar row has no contrast side: translation cannot produce these items, so they must be built natively. § 6.1 measures how far the rankings disagree in each row.
Model
Params.
Licence
Reported cutoff
# Langs
HF SHA
Ref.
ibm-granite/granite-4.0-h-350m
350M
Apache 2.0
—
12
3b17b717b8f2f5d305b0a92c1491e239aeda19c8
( Granite Team, IBM, 2024 )
ibm-granite/granite-4.0-350m
350M
Apache 2.0
—
12
bd8a1497065c0d6ba1ef19af6b0d2b14bacf71c2
( Granite Team, IBM, 2024 )
Qwen/Qwen3-0.6B
0.6B
Apache 2.0
—
100+
c1899de289a04d12100db370d81485cdf75e47ca
( Yang et al., 2025 )
Qwen/Qwen3.5-0.8B
0.8B
Apache 2.0
—
201
2fc06364715b967f1860aea9cf38778875588b17
( Yang et al., 2025 )
google/gemma-3-1b-it
1B
gemma
—
140+
dcc83ea841ab6100d6b47a070329e1ba4cf78752
( Team et al., 2025a )
ibm-granite/granite-4.0-h-1b
1B
Apache 2.0
—
12
d18cca4c121edb87d022116d281ce212c9136f57
( Granite Team, IBM, 2024 )
Appendix
Table 12: Summary of benchmarked models. The Ref. column cites each model’s family technical report. “—” indicates models for which no technical report is available, including our own Slovak-adapted/fine-tuned models as well as community Slovak models sourced from Hugging Face.
Group (examples)
Source published
Split previously released?
Underlying source public?
Newly packaged here (Chiby, Slohry, OneRuler-SK, Scio, Autoškola, e-testy-sk, Slovak-tests, Word Analogies, Slovak Trivia)
2024–2026
No — evaluation split first assembled for sk-bench
Often yes (public exams, Wikipedia, dictionary)
Previously released native (NLI-SK-annotated, SK-QuAD, skLEP, SlovakSum, SME-Sum)
Table 13: Data availability by group. “Source published” is when the underlying material became available, not when sk-bench packaged it. “Split previously released” asks whether an evaluation split over that material was already circulating before this work. No contamination measurement was performed. The table documents exposure opportunity, not exposure.
Figure 7: Open-weights size–score Pareto frontier, with every frontier model labelled. Overall sk-bench score against parameter count (log scale). The blue line is the frontier, the best score reachable at or below each size. Gains are steep up to ∼ 9B, then the frontier flattens: past gpt-oss 20B and Qwen3.5 27B no open model beats gpt-oss 120B (60.0), and the 675B–1.1T MoEs sit below it. The dashed line marks the best closed-weights model (Gemini 3.1 Pro, 72.6).
Model
Knowledge & reasoning
Reading & QA
Classification
Generation
Instruction following
Grammar & morphology
Overall
< 2B
Granite 4.0 350M
27.5 ± 3.1
10.1 ± 0.6
13.8 ± 10.3
9.7 ± 3.4
10.8
1.8 ± 1.3
12.3 ± 1.9
Granite 4.0-H 350M
25.5 ± 5.7
9.4 ± 0.9
27.3 ± 10.1
11.7 ± 2.0
15.3
1.3 ± 0.0
15.1 ± 3.2
Qwen3 0.6B
33.1 ± 5.2
20.9 ± 3.5
43.3 ± 7.5
10.6 ± 0.2
16.8
4.2 ± 1.3
21.5 ± 1.6
Qwen3.5 0.8B
33.4 ± 2.8
26.0 ± 0.3
53.5 ± 7.8
9.7 ± 0.2
14.3
7.0 ± 2.5
24.0 ± 1.3
Gemma 3 1B
35.1 ± 2.0
19.9 ± 3.0
54.4 ± 4.1
11.8 ± 1.0
12.3
3.0 ± 0.8
22.7 ± 0.6
Appendix
Table 14: Full sk-bench leaderboard: all 55 models by reporting group, grouped by size tier (Slovak, 0-shot). Each cell is the unweighted mean of that group’s per-task primary metrics ( ×100 ), with best and second-best per column. Overall is the unweighted mean of the six group averages. The small grey ± value is the standard deviation across the three prompt templates. The body leaderboard (Table 1 ) shows the top two open models per size tier and top three proprietary APIs. Per-task scores are in Table 15 .
Model
ARC
MMLU
G-MMLU
TruthQA
SCIO
FinExam
Driving
COPA
SKJ1
Trivia
Belebele
Chiby
WikiQA
ETest-M
ETest-Q
SKTest-M
SKTest-Q
OneRuler
skLEP-QA
NLI
skLEP-NLI
skLEP-RTE
Senti-3
Senti-2
skLEP-Sent
MovieRev
Toxic
skLEP-Hate
SlovakSum
SME-Sum
IFEval
Analogies
Slohry
Avg
< 2B
Granite 4.0 350M
26.5
27.7
26.7
21.9
25.9
25.9
33.2
48.7
12.7
26.1
30.0
0.9
3.6
19.4
0.2
25.4
1.3
7.2
3.0
7.3
7.1
5.2
28.5
22.0
23.8
25.6
3.1
1.4
9.8
9.6
10.8
0.3
3.2
15.9
Granite 4.0-H 350M
24.1
26.6
25.6
13.8
21.8
22.3
32.2
50.5
15.3
22.4
26.7
2.7
1.7
20.0
0.2
26.9
0.8
0.5
4.9
4.0
3.9
14.9
43.3
67.4
10.3
52.5
41.2
8.2
10.8
12.6
15.3
0.0
2.6
19.0
Qwen3 0.6B
25.9
25.9
26.2
65.3
21.4
31.0
61.1
49.4
3.3
21.1
34.2
26.8
1.2
20.9
0.5
29.0
1.5
56.0
17.6
23.4
21.2
36.1
28.3
75.8
33.8
58.6
56.3
56.6
11.6
9.7
16.8
2.6
5.9
28.9
Qwen3.5 0.8B
41.1
33.1
35.8
28.5
28.7
34.5
54.7
50.1
2.7
24.6
47.1
43.9
14.9
25.4
0.9
36.8
1.3
50.5
13.4
24.7
22.8
56.7
67.5
85.2
57.6
69.0
57.7
40.2
10.9
8.5
14.3
7.3
6.7
33.2
Gemma 3 1B
27.3
28.4
32.8
52.8
23.5
35.0
70.0
51.8
2.7
26.9
34.0
34.7
13.7
21.7
2.6
35.9
2.1
13.9
20.1
22.2
24.2
51.7
65.2
81.5
82.0
73.1
56.1
33.3
12.0
11.7
12.3
2.0
4.0
32.2
Appendix
Table 15: sk-bench results across all tasks. Slovak (SK), 0-shot. Each cell is the task’s primary metric ( ×100 ): accuracy for multiple-choice and classification, quasi-exact-match (qem) for open QA, ROUGE-L for summarisation, and prompt-level strict accuracy for IFEval. Best and second-best per column. Avg is the unweighted mean of the columns shown.
Model
ARC
MMLU
G-MMLU
TruthQA
SCIO
FinExam
Driving
COPA
SKJ1
Trivia
Belebele
Chiby
WikiQA
ETest-M
ETest-Q
SKTest-M
SKTest-Q
OneRuler
skLEP-QA
NLI
skLEP-NLI
skLEP-RTE
Senti-3
Senti-2
skLEP-Sent
MovieRev
Toxic
skLEP-Hate
SlovakSum
SME-Sum
IFEval
Analogies
Slohry
Avg
< 2B
Qwen3 0.6B
39.1
34.5
33.2
29.3
30.6
40.4
53.7
53.1
7.0
26.7
46.7
46.1
10.3
26.8
2.9
31.1
2.4
49.3
18.7
36.9
38.6
64.9
35.1
47.2
79.5
65.5
54.8
31.8
11.2
9.6
36.3
4.5
6.1
33.5
Qwen3.5 0.8B
7.1
2.8
2.3
0.6
2.5
1.7
3.6
1.9
0.0
0.3
1.2
0.1
0.0
4.1
0.1
2.3
0.8
0.5
0.0
2.5
3.2
0.3
2.9
4.5
15.2
6.1
3.9
2.5
0.3
0.2
24.4
2.3
18.9
3.6
Qwen3 1.7B
67.5
51.6
53.9
36.9
49.5
49.6
59.4
68.8
22.0
31.4
72.0
51.1
25.1
42.4
14.8
43.7
5.6
93.2
23.5
54.8
57.2
76.1
49.3
63.4
92.0
81.1
71.2
68.7
12.0
11.7
47.3
22.2
22.8
48.2
2–9B
Qwen3.5 2B
42.9
31.2
25.8
10.9
23.9
24.7
31.9
13.5
1.3
4.9
29.9
18.8
21.3
18.5
1.4
13.5
1.4
26.9
23.5
20.6
18.2
23.8
20.3
26.2
62.6
41.7
26.4
24.3
0.5
0.3
25.9
23.7
25.7
21.4
Appendix
Table 16: sk-bench results for thinking-mode models. Slovak (SK), 0-shot. Each cell is the task’s primary metric ( ×100 ): accuracy for multiple-choice and classification, quasi-exact-match (qem) for open QA, ROUGE-L for summarisation, and prompt-level strict accuracy for IFEval. Best and second-best per column. Avg is the unweighted mean of the columns shown.
Model
ARC
MMLU
G-MMLU
TruthQA
SCIO
FinExam
Driving
COPA
SKJ1
Trivia
Belebele
Chiby
WikiQA
ETest-M
ETest-Q
SKTest-M
SKTest-Q
OneRuler
skLEP-QA
NLI
skLEP-NLI
skLEP-RTE
Senti-3
Senti-2
skLEP-Sent
MovieRev
Toxic
skLEP-Hate
SlovakSum
SME-Sum
Analogies
Avg
< 2B
Granite 4.0 350M
26.3
27.0
25.4
7.5
27.7
24.7
19.0
50.3
0.0
26.5
32.0
50.3
6.4
20.6
0.6
28.1
1.3
16.8
8.5
34.4
31.2
53.8
46.6
74.8
59.9
70.8
46.7
72.7
8.2
9.0
0.9
29.3
Granite 4.0-H 350M
25.5
27.8
25.9
7.6
24.3
25.3
25.7
50.7
48.3
22.9
28.7
50.2
6.3
22.4
0.6
20.2
0.8
24.7
14.3
25.2
23.7
44.5
62.5
87.4
30.4
65.4
47.3
67.0
11.1
15.1
0.3
30.1
Qwen3 0.6B
29.9
29.5
29.0
30.2
23.8
32.6
65.5
48.5
7.7
24.1
39.6
50.7
8.1
24.7
0.3
30.5
1.1
52.3
18.6
32.3
29.6
52.5
65.9
66.8
46.2
55.7
53.3
28.0
12.4
12.5
2.0
32.4
Qwen3.5 0.8B
40.5
32.8
37.8
48.9
27.6
37.1
60.5
50.6
0.0
26.1
53.5
51.4
19.5
25.8
1.0
37.6
1.7
37.3
15.0
31.6
28.7
60.0
64.8
85.8
57.2
67.2
64.4
67.6
12.4
11.9
1.1
37.3
Gemma 3 1B
32.0
31.3
36.9
39.8
26.7
35.6
65.8
51.5
1.7
28.6
37.1
52.0
14.5
24.1
1.7
35.9
2.2
20.1
15.2
44.4
43.3
65.3
66.7
87.2
60.1
72.5
63.6
46.8
10.1
10.2
1.3
36.3
Appendix
Table 17: sk-bench results across all tasks. English (EN) prompts, 0-shot — the English-template counterpart of Table 15 . IFEval-SK and Slohry are Slovak-only by construction and omitted. Each cell is the task’s primary metric ( ×100 ), with best and second-best per column. Avg is the unweighted mean of the columns shown.
Model
ARC
MMLU
G-MMLU
Scio
Driving
COPA
SKTest-M
SKTest-Q
skLEP-QA
skLEP-NLI
skLEP-RTE
Senti-3
Senti-2
skLEP-Sent
MovieRev
Toxic
skLEP-Hate
Analogies
Slohry
Avg
< 2B
Granite 4.0 350M
24.6
27.8
27.1
23.9
24.9
49.1
26.2
1.0
10.1
24.9
42.2
16.1
21.2
36.8
48.0
41.8
62.9
0.5
1.5
26.9
Granite 4.0-H 350M
24.3
25.4
26.0
24.4
59.0
49.9
30.7
1.5
9.4
26.9
49.0
67.9
85.6
22.8
58.8
31.1
44.6
0.1
0.4
33.6
Qwen3 0.6B
35.5
31.8
32.7
26.5
36.5
43.5
33.9
1.1
15.1
35.7
45.2
47.0
84.6
51.8
59.1
54.9
46.9
0.7
8.8
36.4
Qwen3.5 0.8B
40.1
33.6
34.2
26.1
32.3
50.5
29.8
2.1
11.6
36.6
57.2
51.6
78.1
76.2
73.7
54.6
33.6
4.1
9.0
38.7
Gemma 3 1B
28.2
30.0
33.0
25.0
65.0
51.5
33.1
1.8
17.0
35.4
58.6
49.9
61.5
91.5
80.1
54.1
47.8
0.6
3.4
40.4
Appendix
Table 18: sk-bench results across all tasks. Slovak (SK), 3-shot — the few-shot counterpart of Table 15 , restricted to the tasks and models that were run with 3-shot exemplars (random exemplars from each task’s train/validation split). Summarisation, instruction following, and the other tasks run zero-shot only are omitted, as are the models without few-shot results. Each cell is the task’s primary metric ( ×100 ), with best and second-best per column. Avg is the unweighted mean of the columns shown.
Model
ARC
MMLU
G-MMLU
Scio
Driving
COPA
SKTest-M
SKTest-Q
skLEP-QA
skLEP-NLI
skLEP-RTE
Senti-3
Senti-2
skLEP-Sent
MovieRev
Toxic
skLEP-Hate
Avg
< 2B
Granite 4.0 350M
26.3
28.0
29.2
25.3
22.2
49.9
28.6
1.3
11.4
34.2
50.3
12.7
70.0
90.1
70.3
46.4
72.7
0.2
37.2
Granite 4.0-H 350M
27.1
27.6
28.1
28.9
31.3
50.6
30.7
1.1
11.1
31.0
55.1
55.1
87.3
64.1
60.0
47.0
72.5
0.2
39.4
Qwen3 0.6B
36.8
34.1
34.3
29.6
30.7
53.3
31.8
2.1
15.5
36.7
55.2
44.1
81.0
83.0
67.6
55.7
63.7
0.4
42.0
Qwen3.5 0.8B
43.3
35.0
37.6
28.6
48.7
51.9
31.6
1.7
14.6
36.7
60.4
48.7
76.6
89.6
74.7
57.9
71.5
0.6
45.0
Gemma 3 1B
30.5
29.9
33.6
24.0
67.9
52.4
35.8
2.2
18.1
40.3
61.5
50.6
64.4
91.5
78.0
72.7
67.4
0.9
45.7
Appendix
Table 19: sk-bench results across all tasks. English (EN) prompts, 3-shot — the English-template, few-shot counterpart of Table 15 , restricted to the tasks and models run with 3-shot exemplars. Summarisation, instruction following, the Slovak-only tasks (IFEval-SK, Slohry), and the other tasks run zero-shot only are omitted, as are the models without few-shot results. Each cell is the task’s primary metric ( ×100 ), with best and second-best per column. Avg is the unweighted mean of the columns shown.
We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4× the depth of existing multilingual benchmark coverage for Slovak. Our evaluation of 31 embedding models reveals that large instruction-tuned multilingual models achieve the strongest performance, while existing Slovak-specific models trained for NLU tasks transfer poorly to embedding tasks. To address the need for efficient, locally-deployable Slovak embeddings, we develop \texttt{e5-sk-small} (45M parameters) and \texttt{e5-sk-large} (365M) by applying vocabulary trimming and fine-tuning to Multilingual E5 models. Despite size reductions of up to 62%, our open-source models achieve competitive performance with proprietary APIs while remaining locally deployable for semantic search and retrieval-augmented generation (RAG). We release the benchmark, models, datasets, and code openly, hoping our approach offers a replicable path for other under-resourced languages.
Marek Šuppa, Andrej Ridzik, Daniel Hládek +2
Comenius University in Bratislava, Slovakia · Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia · Technical University of Košice, Slovakia +1
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two natively authored datasets−KyrgyzMMLU and KyrgyzRC−together with carefully translated and manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. We evaluate 26 open- and closed-source LLMs under zero-shot and few-shot settings, analyzing model performance, cross-lingual transfer, and the impact of translation artifacts on evaluation reliability. Across families and tasks, model rankings transfer broadly from English to Kyrgyz on WinoGrande and BoolQ, and to a lesser extent on MMLU, while HellaSwag exhibits a substantial English-Kyrgyz performance gap consistent with translation-induced plausibility shifts. Few-shot prompting improves several open-source models on reading comprehension but behaves inconsistently for proprietary models on translated tasks. We publicly release all datasets, evaluation code, and per-model results, and integrate the Kyrgyz tasks into a widely used multilingual evaluation framework to support future research on Kyrgyz NLP.
Timur Turatali, Aida Turdubaeva, Rustem Izmailov +2
1The Cramer Project · Institute of IT, Kyrgyz State Technical University named after I. Razzakov · School of Computer Science, University of Windsor +2
Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally. We evaluate four open-weight instruction-tuned models on SomaliBench v0, a native-author-verified benchmark of 100 harmful-intent prompts paired across English and Somali. Each of Llama-3.1-8B-Instruct, Gemma-2-9B-Instruct, Qwen-2.5-7B-Instruct, and Aya-23-8B is run locally with temperature 0 and the same English "helpful, harmless, and honest" (HHH) system prompt. A pinned Claude Sonnet snapshot (claude-sonnet-4-5-20250929) classifies each response as refused, complied, or unclear; the native author spot-checks a stratified 80-row sample. We find large English-to-Somali refusal gaps for all four models: Llama-3.1-8B (0.90; 95% bootstrap CI [0.85, 0.96]), Aya-23-8B (0.75 [0.67, 0.83]), Qwen-2.5-7B (0.69 [0.59, 0.78]), and Gemma-2-9B (0.38 [0.27, 0.49]). For three models, the dominant Somali non-refusal mode is not fluent harmful compliance but unclear output: empty, wrong-language, or incoherent generations. The native verification spot-check achieves 100% agreement with the judge (Cohen's kappa = 1.00) on the 80 sampled rows. We report aggregate refusal rates, category gaps, and reliability statistics only; raw model generations are retained locally and are not released.