TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
Organizations: Department of Computer Engineering, Boğaziçi University, Türkiye · VNGRS, Istanbul, Türkiye · Technical University of Munich (TUM), Munich, Germany
Abstract
The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish lacks a monolingual encoder trained from scratch with the advances consolidated in ModernBERT (rotary positional embeddings, FlashAttention, refined normalization). We introduce TabiBERT, a monolingual Turkish encoder based on the ModernBERT architecture, pretrained from scratch for one trillion tokens sampled from an 86.58B-token multi-domain corpus of web text (72%), scientific publications (19%), source code (6%) and mathematical content (0.3%). The model supports a context length of 8,192 tokens, sixteen times that of existing Turkish BERT models, and inherits the ModernBERT architecture's efficiency at long context. For rigorous and reproducible evaluation we introduce TabiBench, a benchmark of 27 datasets across eight task categories with standardized splits and evaluation protocols, summarized as a GLUE-style macro-average on a 0-100 scale. TabiBERT leads the Turkish models in five of eight categories and BERTurk, the previous best, in six of eight; the gains concentrate on question answering (+9.55 F1) and code retrieval (+2.41 NDCG@10), while the four short-text categories are near saturation. Its average of 77.28 exceeds BERTurk's 75.66; the multilingual mmBERT reaches 78.98 with twice the parameters and three times the training tokens, at 41% more tokens per Turkish input. We release model weights, training configurations and evaluation code as a transparent and reproducible foundation for future Turkish encoder research.
Figures & tables
| Model | # params (M) | Data size (GB) | Context length | Code and math | Flash Attention | T abi B ench score |
|---|---|---|---|---|---|---|
| T urkish BERT weet | ✗ | ✗ | ||||
| ytu-c osmos- BERT | ✗ | ✗ | ||||
| BERT urk | ✗ | ✗ | ||||
| T abi BERT | 8192 | ✓ | ✓ | 77.28 |
| Corpus | Type | # docs | # tokens (B) | Sampling rate |
| FineWeb-2 Turkish | Web | |||
| FineWeb-2 English | Web | |||
| Wikipedia-TR | Web | |||
| Dergipark | Scientific | |||
| Yöktez | Scientific | |||
| Books | Literary |
| Dataset | Provenance | Metric | #Labels | Train | Val | Test |
|---|---|---|---|---|---|---|
| Text classification | ||||||
| News Cat ( Guney, 2024 ) | Native | Macro-F1 | 5 | |||
| Product Reviews ( Barmanbay, 2023 ) | Native | Macro-F1 | 2 | |||
| Bil Tweet News ( Toraman, 2017 ; Toraman, 2022 ) | Native | Macro-F1 | 4 | |||
| Gender Hate Speech Turkish ( Toraman et al., 2022 ; Şahinuç et al., 2023 ) | Native | Macro-F1 | 3 | |||
| Token classification | ||||||
| Parameter | Values |
| Learning Rate | , , , |
| Weight Decay | , |
| Batch Size | 16, 32 |
| Epoch | 10, with early stopping |
| Model | # of params | Text clf. | Token clf. | STS | NLI | QA | Academic understanding | Information retrieval | Code retrieval | Total avg. |
| M | macro-F1 | micro-F1 | Pearson | macro-F1 | F1 | macro-F1 | NDCG@10 | NDCG@10 | ( T abi B ench) | |
| T urkish BERT weet | ||||||||||
| ytu-c osmos- BERT | 84.25 | |||||||||
| BERT urk | 93.67 | 85.33 | ||||||||
| T abi BERT | 84.51 | 69.71 | 70.06 | 75.44 | 56.95 | 77.28 | ||||
| mm BERT |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Model | NewsCat | BilTweetNews | GenderHateSpeech | ProductReviews | Weighted avg. |
| Test size | |||||
| T urkish BERT weet | |||||
| ytu-c osmos- BERT | 97.21 | 71.02 | 85.04 | 84.25 | |
| BERT urk | 57.87 | ||||
| T abi BERT | |||||
| mm BERT |
| Model | WikiNER | WikiANN-TR | PosUD-BOUN | PosUD-IMST | Weighted avg. |
| Test size | |||||
| T urkish BERT weet | |||||
| ytu-c osmos- BERT | 95.41 | ||||
| BERT urk | 79.44 | 94.60 | 93.67 | ||
| T abi BERT | 90.40 | ||||
| mm BERT |
| Model | SICK-TR | STSb-TR | Weighted avg. |
| Test size | |||
| T urkish BERT weet | |||
| ytu-c osmos- BERT | |||
| BERT urk | 85.95 | 85.33 | |
| T abi BERT | 83.84 | ||
| mm BERT |
| Model | SNLI-TR | MultiNLI-TR | Weighted avg. |
| Test size | |||
| T urkish BERT weet | |||
| ytu-c osmos- BERT | |||
| BERT urk | 87.21 | ||
| T abi BERT | 80.60 | 84.51 | |
| mm BERT |
| Model | TQuAD | XQuAD | Weighted avg. |
| Test size | |||
| T urkish BERT weet | 39.40 | ||
| ytu-c osmos- BERT | |||
| BERT urk | |||
| T abi BERT | 72.34 | 69.71 | |
| mm BERT |
| Model | BiText | MsMarco-TR | Scifact-TR | Fiqa-TR | NFCorpus-TR | Quora-TR | Weighted avg. |
| Test size | |||||||
| T urkish BERT weet | |||||||
| ytu-c osmos- BERT | 79.88 | 51.10 | 92.87 | ||||
| BERT urk | 32.47 | ||||||
| T abi BERT | 99.42 | 83.31 | 75.44 | ||||
| mm BERT |
| Model | Apps-TR | CosQA-TR | StackOverflowQA-TR | CodeSearchNet-21K-TR | Weighted avg. |
| Test size | |||||
| T urkish BERT weet | |||||
| ytu-c osmos- BERT | |||||
| BERT urk | 18.18 | ||||
| T abi BERT | 88.11 | 82.85 | 84.12 | 56.95 | |
| mm BERT |
| Model | PubMedRCT-10K-TR | SciCite-TR | Thesis-Abstract-Classification-11K | Weighted avg. |
| Test size | ||||
| T urkish BERT weet | ||||
| ytu-c osmos- BERT | 84.09 | |||
| BERT urk | 75.61 | |||
| T abi BERT | 50.77 | 70.06 | ||
| mm BERT |
| TQuAD | |||
| Model | Train | Val | Test |
| Split Size | |||
| BERT urk | |||
| T urkish BERT weet | |||
| ytu-c osmos- BERT | |||
| T abi BERT | |||
| TQuAD | |||
| Model | Published | 512-cap | |
| T urkish BERT weet | |||
| ytu-c osmos- BERT | |||
| BERT urk | |||
| T abi BERT | |||
| mm BERT | |||
| Dataset | Provider | License / terms | Redistrib. | Processing |
|---|---|---|---|---|
| Text classification | ||||
| News Cat | ( Guney, 2024 ) mcemilg/news-cat | Not declared | Unclear | Used as-is; standardized split |
| Product Reviews | ( Barmanbay, 2023 ) fthbrmnby/turkish_product_reviews | CC BY-SA 4.0, stated in the card’s licensing section and in the provider’s GitHub LICENCE file (the card’s license tag is unknown ) | Yes | Used as-is; standardized split |
| Bil Tweet News | ( Toraman, 2017 ; Toraman, 2022 ) ctoraman/BilTweetNews-sentiment-analysis | CC BY-NC-SA 4.0 | Non-commercial only | Used as-is; standardized split |
| Gender Hate Speech Turkish | ( Toraman et al., 2022 ; Şahinuç et al., 2023 ) ctoraman/gender-hate-speech-turkish | CC BY-NC-SA 4.0 | Non-commercial only | Used as-is; standardized split |
| Token classification | ||||