TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish
Authors: Melikşah Türker, A. Ebrar Kızıloğlu, Onur Güngör, Susan Üsküdarlı
Organizations: Department of Computer Engineering, Boğaziçi University, Türkiye · VNGRS, Istanbul, Türkiye · Technical University of Munich (TUM), Munich, Germany
The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish lacks a monolingual encoder trained from scratch with the advances consolidated in ModernBERT (rotary positional embeddings, FlashAttention, refined normalization). We introduce TabiBERT, a monolingual Turkish encoder based on the ModernBERT architecture, pretrained from scratch for one trillion tokens sampled from an 86.58B-token multi-domain corpus of web text (72%), scientific publications (19%), source code (6%) and mathematical content (0.3%). The model supports a context length of 8,192 tokens, sixteen times that of existing Turkish BERT models, and inherits the ModernBERT architecture's efficiency at long context. For rigorous and reproducible evaluation we introduce TabiBench, a benchmark of 27 datasets across eight task categories with standardized splits and evaluation protocols, summarized as a GLUE-style macro-average on a 0-100 scale. TabiBERT leads the Turkish models in five of eight categories and BERTurk, the previous best, in six of eight; the gains concentrate on question answering (+9.55 F1) and code retrieval (+2.41 NDCG@10), while the four short-text categories are near saturation. Its average of 77.28 exceeds BERTurk's 75.66; the multilingual mmBERT reaches 78.98 with twice the parameters and three times the training tokens, at 41% more tokens per Turkish input. We release model weights, training configurations and evaluation code as a transparent and reproducible foundation for future Turkish encoder research.
Figures & tables
Model
# params (M)
Data size (GB)
Context length
Code and math
Flash Attention
T abi B ench score
T urkish BERT weet
163
110
512
✗
✗
67.16
ytu-c osmos- BERT
111
75
512
✗
✗
71.95
BERT urk
110
35
512
✗
✗
75.66
T abi BERT
149
426
8192
✓
✓
77.28
Table 1: Turkish BERT-based models across scale, training data, architectural features, and benchmark performance.
Corpus
Type
# docs
# tokens (B)
Sampling rate
FineWeb-2 Turkish
Web
88,769,907
56.041
1
FineWeb-2 English
Web
5,000,000
5.676
1
Wikipedia-TR
Web
604,716
0.169
2
Dergipark
Scientific
334,429
1.738
1
Yöktez
Scientific
475,817
14.887
1
Books
Literary
5,078
0.621
3
Table 2: Training Corpora Statistics.
Figure 1 : Relative tokenizer efficiency, in per cent, on Turkish text across established Turkish BERT models and mmBERT, normalized to T abi BERT (dashed baseline). X-axis labels indicate models with their vocabulary sizes, while the Y-axis reports relative token counts required for text representation. T abi BERT ’s fertility is slightly higher than BERT urk’s despite its larger vocabulary, which we attribute to the pre-tokenization rule described in Section 3.2 . The multilingual mmBERT, despite its 256K-token vocabulary, exhibits 41.0 per cent higher fertility, reflecting greater subword fragmentation on Turkish text.
Dataset
Provenance
Metric
#Labels
Train
Val
Test
Text classification
News Cat ( Guney, 2024 )
Native
Macro-F1
5
750
150
250
Product Reviews ( Barmanbay, 2023 )
Native
Macro-F1
2
164,615
35,275
35,275
Bil Tweet News ( Toraman, 2017 ; Toraman, 2022 )
Native
Macro-F1
4
696
149
150
Gender Hate Speech Turkish ( Toraman et al., 2022 ; Şahinuç et al., 2023 )
Native
Macro-F1
3
16,002
1,998
2,000
Token classification
Table 3 : T abi B ench per-dataset composition: provenance (origin and, where translated, translating agent, with source citation below each name) and evaluation setup (metric, label count, train/validation/test split sizes), for each of the 27 constituent datasets, grouped by the eight task categories of Section 4.3 . Counts: 10 native, 2 human-translated, 9 machine-translated by prior work, 6 machine-translated by us; 10+2+9+6=27.
Figure 2 : Task-specific performance of models on T abi B ench by category: Text&NLI&Academic (macro-F1), Token Clf (micro-F1), STS (Pearson), QA (F1), and Information Retrieval&Code Retrieval (NDCG@10). T abi BERT consistently performs strongly across all tasks.
Model
# of params
Text clf.
Token clf.
STS
NLI
QA
Academic understanding
Information retrieval
Code retrieval
Total avg.
M
macro-F1
micro-F1
Pearson
macro-F1
F1
macro-F1
NDCG@10
NDCG@10
( T abi B ench)
T urkish BERT weet
163
79.71
92.02
75.86
79.10
38.13
60.60
68.40
43.49
67.16
ytu-c osmos- BERT
111
84.25
93.60
84.68
84.16
31.50
69.29
74.29
53.80
71.95
BERT urk
110
83.42
93.67
85.33
84.33
60.16
69.00
74.84
54.54
75.66
T abi BERT
149
83.44
93.42
84.74
84.51
69.71
70.06
75.44
56.95
77.28
mm BERT
307
82.54
93.81
87.05
84.38
71.47
70.34
76.20
66.02
78.98
Table 5 : Comparison of downstream task performance across all evaluated models. For each column, the highest score among the Turkish models (excluding the multilingual mm BERT ) is shown in bold. The evaluation metric used for each task type is also displayed in the column headers. T abi BERT achieves the highest score in five of the eight task categories among Turkish models; BERT urk leads two (token classification, STS) and ytu-c osmos- BERT one (text classification). The last two rows give, for each category, the best score among the three Turkish baselines and T abi BERT ’s margin over it; the previous best is ytu-c osmos- BERT for text classification and academic understanding and BERT urk for the other six categories. In these two rows the Total Avg column holds the mean of the eight per-category previous bests and T abi BERT ’s margin over that mean.
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
Model
NewsCat
BilTweetNews
GenderHateSpeech
ProductReviews
Weighted avg.
Test size
250
150
2,000
35,275
T urkish BERT weet
91.98
53.16
68.58
80.37
79.71
ytu-c osmos- BERT
97.21
53.39
71.02
85.04
84.25
BERT urk
95.60
57.87
68.25
84.30
83.42
T abi BERT
95.20
50.11
69.01
84.32
83.44
mm BERT
94.80
49.06
66.45
83.51
82.54
Appendix
Table 6 : Downstream performance of models on Text classification (eval metric: macro-F1). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
WikiNER
WikiANN-TR
PosUD-BOUN
PosUD-IMST
Weighted avg.
Test size
1,000
10,000
979
1,100
T urkish BERT weet
72.6
94.05
89.74
93.21
92.02
ytu-c osmos- BERT
78.96
95.41
89.19
94.30
93.60
BERT urk
79.44
95.37
89.82
94.60
93.67
T abi BERT
76.41
95.34
90.40
94.07
93.42
mm BERT
75.97
95.85
90.72
94.26
93.81
Appendix
Table 7 : Downstream performance of models on Token classification (eval metric: micro-F1). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
SICK-TR
STSb-TR
Weighted avg.
Test size
4,927
1,379
T urkish BERT weet
78.35
66.96
75.86
ytu-c osmos- BERT
85.27
82.55
84.68
BERT urk
85.95
83.12
85.33
T abi BERT
85.00
83.84
84.74
mm BERT
87.81
84.31
87.05
Appendix
Table 8 : Downstream performance of models on Semantic textual similarity (eval metric: Pearson correlation). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
SNLI-TR
MultiNLI-TR
Weighted avg.
Test size
9,824
4,923
T urkish BERT weet
83.05
71.21
79.10
ytu-c osmos- BERT
87.18
78.15
84.16
BERT urk
87.21
78.57
84.33
T abi BERT
86.47
80.60
84.51
mm BERT
86.44
80.28
84.38
Appendix
Table 9 : Downstream performance of models on Natural language inference (eval metric: macro-F1). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
TQuAD
XQuAD
Weighted avg.
Test size
2,520
179
T urkish BERT weet
38.04
39.40
38.13
ytu-c osmos- BERT
32.01
24.25
31.50
BERT urk
63.30
15.96
60.16
T abi BERT
72.34
32.61
69.71
mm BERT
71.55
70.40
71.47
Appendix
Table 10 : Downstream performance of models on Question answering (eval metric: F1). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
BiText
MsMarco-TR
Scifact-TR
Fiqa-TR
NFCorpus-TR
Quora-TR
Weighted avg.
Test size
3,000
31,692
339
1,706
12,334
15,675
T urkish BERT weet
95.98
74.20
68.25
36.23
27.79
86.84
68.40
ytu-c osmos- BERT
97.25
81.99
79.88
51.10
28.36
92.87
74.29
BERT urk
96.77
81.73
78.47
49.46
32.47
92.72
74.84
T abi BERT
99.42
83.31
74.22
50.67
31.19
92.47
75.44
mm BERT
99.49
83.68
80.00
52.29
33.44
92.78
76.20
Appendix
Table 11 : Downstream performance of models on Retrieval tasks (eval metric: NDCG@10). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
Apps-TR
CosQA-TR
StackOverflowQA-TR
CodeSearchNet-21K-TR
Weighted avg.
Test size
3,770
500
1,994
3,000
T urkish BERT weet
04.48
82.32
67.01
70.41
43.49
ytu-c osmos- BERT
14.47
82.04
82.65
79.35
53.80
BERT urk
18.18
85.72
79.93
78.15
54.54
T abi BERT
17.49
88.11
82.85
84.12
56.95
mm BERT
32.24
89.31
90.43
88.36
66.02
Appendix
Table 12 : Downstream performance of models on Code retrieval tasks (eval metric: NDCG@10). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
Model
PubMedRCT-10K-TR
SciCite-TR
Thesis-Abstract-Classification-11K
Weighted avg.
Test size
1,500
1,859
1,683
T urkish BERT weet
70.07
79.59
31.18
60.60
ytu-c osmos- BERT
75.32
84.09
47.56
69.29
BERT urk
75.61
81.60
49.20
69.00
T abi BERT
75.32
83.29
50.77
70.06
mm BERT
74.37
83.27
52.47
70.34
Appendix
Table 13 : Downstream performance of models on Academic understanding tasks (eval metric: macro-F1). The number of test samples per task is also given. For each column, the highest score among the Turkish models (excluding mm BERT ) is shown in bold.
TQuAD
Model
Train
Val
Test
Split Size
11,803
2,418
2,520
BERT urk
11.98
13.15
10.87
T urkish BERT weet
11.07
12.12
10.95
ytu-c osmos- BERT
11.32
11.95
10.83
T abi BERT
13.17
13.98
10.83
Appendix
Table 14 : Share of TQuAD question answering examples exceeding 512 tokens under each model’s tokenizer, by dataset split. Input length is measured as question + context, special tokens included, exactly as fed to the model ( truncation=’only_second’ ); uncased models are lowercased first. Values are per cent of examples in each split.
TQuAD
Model
Published
512-cap
Δ
T urkish BERT weet
38.04
38.04
0.00
ytu-c osmos- BERT
32.01
32.01
0.00
BERT urk
63.30
63.30
0.00
T abi BERT
72.34
70.31
−2.03
mm BERT
71.55
69.26
−2.29
Appendix
Table 15 : Published TQuAD question answering F1 scores versus scores obtained when the context is capped at 512 tokens (eval metric: F1). Published scores use 1,024 tokens for T abi BERT and mm BERT and 512 tokens, the maximum length they support, for BERT urk, T urkish BERT weet, and ytu-c osmos- BERT . Since the three baseline models’ TQuAD scores are already 512 -capped, only T abi BERT and mm BERT required a new 512 -token TQuAD run; no other cell required a new fine-tuning run. These new runs reuse each model’s original best hyperparameter configuration rather than repeating hyperparameter tuning, due to compute limitations.
( Toraman et al., 2022 ; Şahinuç et al., 2023 ) ctoraman/gender-hate-speech-turkish
CC BY-NC-SA 4.0
Non-commercial only
Used as-is; standardized split
Token classification
Appendix
Table 16 : T abi B ench source inventory: original provider (with a direct link to the source repository, where one exists), located license/terms, resulting redistribution status, and processing applied by the T abi B ench authors, for each constituent dataset (Section 4.5 ). License information was compiled by consulting original dataset repositories, dataset cards, and associated publications where available; entries marked not declared indicate that no license could be located, not that a dataset is confirmed to be unrestricted. This table reflects a best-effort review rather than an independent legal audit; readers with license-sensitive use cases should verify terms directly against the cited original sources.
Encoder-only Transformers remain effective for discriminative and representation-learning tasks, yet Polish encoders still largely rely on BERT/RoBERTa-style architectures. We introduce \textbf{Polish ModernBERT}, a family of four Polish encoders available at Base and Large scales, each with 512-token and 8K context variants. We adapt the ModernBERT pretraining recipe through staged selection experiments and release a long-context benchmark covering legal topic classification, ideological decision-direction prediction, factual-consistency assessment over literary plot summaries, and human-rights violation assessment. Across 30 tasks, Polish ModernBERT achieves the best overall performance among the evaluated Polish encoders, reaching 83.99 and 85.11 for the Base-8K and Large-8K models, respectively. On long-context tasks, the 8K variants improve over matched Polish RoBERTa-8K baselines from 67.47 to 77.15 and from 75.88 to 78.49 at the Base and Large scales, respectively. The Base-8K model achieves this gain with 22% fewer parameters (149M vs.\ 190M). Efficiency measurements in representative inference setups show lower peak memory usage and latency than matched Polish RoBERTa baselines in both 512-token and 8K settings. Polish ModernBERT-8K-Base additionally achieves the best result on a Polish retrieval benchmark among the evaluated encoders below 300M parameters.
Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata +1
National Information Processing Institute, Warsaw, Poland
Encoder-only transformer models remain essential for production NLP pipelines. We introduce moBERTo, a Portuguese adaptation of ModernBERT obtained through continued pretraining of the ModernBERT-base checkpoint on 60 billion tokens (5 epochs over a 12-billion-token corpus curated from FineWeb2 and filtered with educational and STEM classifiers). We preserve the original architecture, including rotary positional embeddings, alternating local-global attention, flash attention, and unpadding. We evaluate moBERTo across information retrieval (including long-context retrieval at up to 8,192 tokens), document classification, named entity recognition, and natural language understanding. Our best variant, which combines a Portuguese tokenizer with subword-matching embedding transfer and long-context post-training, achieves the highest average reranking nDCG@10 across three Portuguese retrieval benchmarks and the best results on PLUE-PT. Through ablation studies, we show that (i) continued pretraining is strongly preferable to training from scratch, particularly for preserving long-context capabilities; (ii) tokenizer adaptation improves token-level tasks but degrades long-context retrieval; (iii) a dedicated long-context post-training phase at 8,192 tokens further improves reranking and NER; and (iv) encoder-only architectures remain competitive with larger decoder-only alternatives for discriminative tasks. We publicly release the model weights at https://huggingface.co/Tropic-AI/moBERTo and training data at https://huggingface.co/datasets/Tropic-AI/moberto-pretraining-dataset-c4-compatible on Hugging Face.
Thiago Laitz, Thales Sales Almeida, João Guilherme Alves Santos +1
UNICAMP, Campinas, Brazil · Tropic AI · Maritaca AI
Large language models have achieved strong performance across many NLP tasks, yet Urdu remains comparatively underexplored due to limited resources and fragmented evaluation settings. To address this gap, we introduce DunbaaBERT, a family of Urdu RoBERTa-base models trained from scratch with Byte-BPE vocabularies of 32k, 52k, and 96k tokens on a deduplicated 17GB Urdu corpus. We evaluate DunbaaBERT across intrinsic and downstream Urdu NLP benchmarks covering linguistic acceptability, news classification, offensive language detection, and sentiment analysis while analyzing vocabulary-size effects on performance and efficiency trade-offs. Across benchmarks, the DunbaaBERT variants achieve competitive performance against strong multilingual baselines while consistently maintaining favorable efficiency trade-offs. Interestingly, larger vocabularies do not consistently improve downstream effectiveness, with DunbaaBERT32k repeatedly providing the strongest overall efficiency profile. Overall, our results demonstrate that carefully curated Urdu-specific encoder models can remain highly competitive despite comparatively compact model and training scales. All models are released under the MIT license.
Iffat Maab, Waleed Jamil, Raphael Schmitt
Research and Development Center for Large Language Models (LLMC), National Institute of Informatics, Tokyo · Independent Researcher, Edinburgh, United Kingdom · School of Computation, Information and Technology, Technical University of Munich, Germany +1