cs.CLDec 28, 2025

TabiBERT: A Large-Scale ModernBERT Foundation Model and A Unified Benchmark for Turkish

Authors: Melikşah Türker, A. Ebrar Kızıloğlu, Onur Güngör, Susan Üsküdarlı

Organizations: Department of Computer Engineering, Boğaziçi University, Türkiye · VNGRS, Istanbul, Türkiye · Technical University of Munich (TUM), Munich, Germany

Abstract

The introduction of BERT established encoder-only transformer models as a foundational paradigm in natural language processing. Encoder-only models remain the standard tool for classification, tagging and retrieval, where contextual representations and low inference cost matter more than text generation, yet Turkish lacks a monolingual encoder trained from scratch with the advances consolidated in ModernBERT (rotary positional embeddings, FlashAttention, refined normalization). We introduce TabiBERT, a monolingual Turkish encoder based on the ModernBERT architecture, pretrained from scratch for one trillion tokens sampled from an 86.58B-token multi-domain corpus of web text (72%), scientific publications (19%), source code (6%) and mathematical content (0.3%). The model supports a context length of 8,192 tokens, sixteen times that of existing Turkish BERT models, and inherits the ModernBERT architecture's efficiency at long context. For rigorous and reproducible evaluation we introduce TabiBench, a benchmark of 27 datasets across eight task categories with standardized splits and evaluation protocols, summarized as a GLUE-style macro-average on a 0-100 scale. TabiBERT leads the Turkish models in five of eight categories and BERTurk, the previous best, in six of eight; the gains concentrate on question answering (+9.55 F1) and code retrieval (+2.41 NDCG@10), while the four short-text categories are near saturation. Its average of 77.28 exceeds BERTurk's 75.66; the multilingual mmBERT reaches 78.98 with twice the parameters and three times the training tokens, at 41% more tokens per Turkish input. We release model weights, training configurations and evaluation code as a transparent and reproducible foundation for future Turkish encoder research.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Polish ModernBERT: The Long and Short of Polish Language Understanding

    Sep 1, 2026Michał Perełkiewicz, Sławomir Dadas, Rafał Poświata +1Bert-Based ModelsQuestion-Answering Benchmarks

  2. moBERTo: A Modern Encoder for Portuguese via Continued Pretraining of ModernBERT

    Jun 21, 2026Thiago Laitz, Thales Sales Almeida, João Guilherme Alves Santos +1Bert-Based ModelsTransformer Encoder

  3. DunbaaBERT: From Sacrifice to Semantics

    May 26, 2026Iffat Maab, Waleed Jamil, Raphael SchmittUrduMultilingual Language Models