cs.CLOct 6, 2026

sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak

Authors: Marek Šuppa, Ivan Vykopal, Andrej Ridzik, Kristián Sopkovič, Natália Kňažeková, Jaroslav Kopčan, Miroslav Blšták, Viktória Ondrejová, +4 more

Organizations: Comenius University in Bratislava, Slovakia · Cisco Systems · Kempelen Institute of Intelligent Technologies, Bratislava, Slovakia · Brno University of Technology, Czechia · Technical University of Košice, Slovakia

Abstract

Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data (ρ≥0.98ρ\geq0.98), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently (ρ=0.72ρ=0.72). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench

Figures & tables

Appendix figures & tables21 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation

    Jun 11, 2026Marek Šuppa, Andrej Ridzik, Daniel Hládek +2Multilingual BenchmarkText Embeddings

  2. KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

    Jul 19, 2026Timur Turatali, Aida Turdubaeva, Rustem Izmailov +2Multilingual Benchmark