cs.CLApr 28, 2026

Language corpora for the Dutch medical domain

Authors: B. van Es

Organizations: Central Diagnostic Laboratory, University Medical Center Utrecht, Utrecht, The Netherlands. · R&D, B-lab, Castricum, The Netherlands.

Abstract

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. \ \textbf{Results:} The resulting corpus comprises ±\pm 35 billion tokens across the medical domain in about 100 million documents, freely available on Hugging Face. \ \textbf{Conclusion:} This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.

Explore similar work

CardsList
  1. TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling

    Sep 22, 2026Julien Knafou, Luc Mottin, Anaïs Mottaz +2