cs.CLSep 21, 2026

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

Authors: Yusser Al Ghussin, Eva Gavaller, Cristina España-Bonet, Josef van Genabith, Simon Ostermann

Organizations: Saarland University · German Research Center for Artificial Intelligence (DFKI) · Barcelona Supercomputing Center (BSC-CNS), Barcelona, Catalonia, Spain

Abstract

Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduce FineWeb-CLaR, a large-scale annotated dataset derived from FineWeb and FineWeb-2 that places web documents on a shared culture-language-region axis for corpus auditing and benchmark alignment. FineWeb-CLaR annotates the full 30.9B-document collection from FineWeb and FineWeb-2 with URL-derived region labels and cultural-topic provenance. Our region resolver assigns a non-empty region to 25.61% of documents (7.92B). For cultural-topic analysis, we induce locale-specific topics and project them onto the 14 leaves of the Cultural Taxonomy of Liu et al. (2025), producing Locale Topic Distributions (LTDs) for corpus-side comparison. We also annotate 277 cultural NLP benchmarks with the same taxonomy, language coverage, and region coverage. Together, these resources enable direct comparison between corpus-side pretraining evidence and benchmark-side evaluation coverage.

Figures & tables

Appendix figures & tables11 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. The Culture Funnel: You Can't Align What isn't in the Data

    Jun 11, 2026Ananya Sahu, Mehrnaz Mofakhami, Daniel D'Souza +3Cultural AwarenessLarge Language Model Alignment

  2. AlignCultura: Towards Culturally Aligned Large Language Models?

    Apr 21, 2026Gautam Siddharth Kashyap, Mark Dras, Usman NaseemCultural AwarenessLarge Language Model Alignment

  3. BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

    Aug 13, 2026Jophin John, Michael Hoffmann, Jan Fillies +2Multilingual BenchmarkDialects