cs.CLJun 2, 2025

Common Corpus: The Largest Collection of Ethical Data for LLM Pre-Training

Authors: Pierre-Carl Langlais, Pavel Chizhov, Catherine Arnett, Carlos Rosas-Hinostroza, Mattia Nee, Eliot Krzystof Jones, Irène Girard, David Mach, +2 more

Organizations: PleIAs, Paris, France · CAIRO, Technical University of Applied Sciences Würzburg-Schweinfurt, Germany · EleutherAI

Abstract

Large Language Models (LLMs) are pre-trained on large amounts of data from different sources and domains. Such datasets often contain trillions of tokens, including large portions of copyrighted or proprietary content, which raises questions about the legal use of such models. This underscores the need for truly open pre-training data that complies with data security regulations. In this paper, we introduce Common Corpus, the largest open dataset for LLM pre-training. The data assembled in Common Corpus are either uncopyrighted or under open licenses, totaling about two trillion tokens. The dataset contains a wide variety of languages, ranging from the high-resource European languages to some low-resource languages rarely represented in pre-training datasets. In addition, it includes a large amount of code data. The diversity of data sources in terms of covered domains and time periods opens up the paths for both research and entrepreneurial needs across diverse areas of knowledge. In this paper, we present the detailed provenance of data assembling and the details of dataset filtering and curation. We train two small language models on Common Corpus and find that they perform comparably to other models of their size, indicating that our dataset is suitable for multilingual pretraining. Common Corpus represents a key contribution to the ecosystem for open science research on Large Language Models.

Figures & tables

Appendix figures & tables14 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

    Sep 29, 2026Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois +7Synthetic Training DataSynthetic Data

  2. MC-PDD: Masked Corpus-Level Pretraining Data Detection for Black-Box Large Language Models

    Jun 6, 2026Kaixin Lan, Mu You, Tao Fang +3PretrainingMachine-Generated Text Detection

  3. MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

    Jul 1, 2026Maximilian Idahl, Jörg Tiedemann, Sampo Pyysalo +19Multilingual Language ModelsLarge Language Model Pretraining