FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing
Organizations: Saarland University · German Research Center for Artificial Intelligence (DFKI) · Barcelona Supercomputing Center (BSC-CNS), Barcelona, Catalonia, Spain
Abstract
Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely indexed with comparable metadata. Benchmarks increasingly target culturally situated phenomena at the level of languages, regions, and locale-specific practices, while web-scale corpora are usually organized only by language. A shared culture-language-region layer makes these resources comparable, enabling audits of whether a target cultural phenomenon is represented in pretraining data, evaluated by benchmarks or both. To this end, we introduce FineWeb-CLaR, a large-scale annotated dataset derived from FineWeb and FineWeb-2 that places web documents on a shared culture-language-region axis for corpus auditing and benchmark alignment. FineWeb-CLaR annotates the full 30.9B-document collection from FineWeb and FineWeb-2 with URL-derived region labels and cultural-topic provenance. Our region resolver assigns a non-empty region to 25.61% of documents (7.92B). For cultural-topic analysis, we induce locale-specific topics and project them onto the 14 leaves of the Cultural Taxonomy of Liu et al. (2025), producing Locale Topic Distributions (LTDs) for corpus-side comparison. We also annotate 277 cultural NLP benchmarks with the same taxonomy, language coverage, and region coverage. Together, these resources enable direct comparison between corpus-side pretraining evidence and benchmark-side evaluation coverage.
Figures & tables
| # | Source | Signal | Conf |
|---|---|---|---|
| 1 | query | explicit locale parameter | high |
| 2 | host_path | lookup table | inherited |
| 3 | host | lookup table | inherited |
| 4 | domain | lookup table | inherited |
| 5 | url_hint | curated URL token | url_hint |
| 6 | tld | non-branded ccTLD | ccTLD |
| Setting | Source / tier | Documents | % |
|---|---|---|---|
| query | 15,283,520 | 0.05 | |
| lookup-table | host_path | 222,919,273 | 0.72 |
| host | 3,458,247,367 | 11.19 | |
| domain | 946,609,661 | 3.06 | |
| fallback | url_hint | 342,035,324 | 1.11 |
| tld | 2,932,873,160 | 9.49 |
| Tier | IP | og:locale | All | All |
|---|---|---|---|---|
| high | 21.9 | 41.4 | 9.1 | 14.4 |
| medium | 19.8 | 16.1 | 3.6 | 19.0 |
| low | 17.1 | 9.6 | 3.8 | 22.6 |
| broad-only | 37.8 | 17.1 | 9.3 | 33.2 |
| Locale | Topic | CT leaf |
|---|---|---|
| en-US | Christian Religion | Values-General |
| en-US | Restaurant Dining | Concepts |
| en-US | US Politics | Demographics |
| ar-EG | Travel Destinations | Concepts |
| ar-EG | Sports Media | Context |
| zh-CN | Digital Media & Devices | Artifacts |
Appendix figures & tables11 assets
Supplementary material from the paper’s appendix.
Appendix
| Rule | Evidence pattern | Weight |
|---|---|---|
| A1 | ccTLD + HTML lang + Content-Language agree | 3.0 |
| A2 | ccTLD + one header agree | 2.5 |
| A3 | ccTLD beats high-resolution default | 2.0 |
| A4 | ccTLD overrides country-specific header | 1.0 |
| A5 | ccTLD alone (no header signal) | 2.0 |
| A6 | Path/subdomain hint + header | 2.0 |
| Column | Dtype | Meaning / values |
|---|---|---|
| key_type | str | domain , host , or host_path |
| key | str | URL signature (normalised, lowercased) |
| n_docs | int | raw record count contributing to the signature |
| top_region | str | winning region (ISO 3166-1 alpha-2) or empty |
| share | float | vote share of the winning region |
| n_distinct_regions | int | distinct candidate regions seen |
| Variant | Kernel | 266 region cells all covered | locale classes moved 2 pt (max) | all 5 pairs above floor | 5-pair order preserved | Spearman , language medians mean [95% interval] |
|---|---|---|---|---|---|---|
| swap-soft | uniform | 100% | 0% (0.52) | 100% | 44% | 0.77 [0.72, 0.81] |
| swap-soft | softmax | 100% | 0% (0.41) | 100% | 59% | 0.83 [0.79, 0.87] |
| hard | uniform | 100% | 100% (25.79) | 100% | 15% | 0.46 [0.42, 0.52] |
| hard | softmax | 100% | 100% (24.97) | 100% | 15% | 0.48 [0.44, 0.54] |
| Inside | Outside | |
|---|---|---|
| Languages | 94 | 203 |
| Locales | 5,419 | 535 |
| Merged topics | 728,950 | 72,394 |
| Sampled documents | 27,922,687 | 1,772,582 |
| Adjudicator change rate | 7.5% | 3.5% |
| 95% CI (cluster bootstrap) | [7.3, 7.6] | [3.1, 3.9] |
| Split-half JSD ( ) | ||||
|---|---|---|---|---|
| Locale size | Locales | Median | IQR | |
| 1–2k | 1,047 | 10.6 | [6.2, 16.3] | 28.2 |
| 2–5k | 1,106 | 6.6 | [3.7, 10.3] | 18.2 |
| 5–10k | 650 | 4.3 | [2.3, 6.7] | 11.6 |
| 10k (cap) | 1,912 | 3.2 | [1.6, 5.0] | 8.7 |
| English-locale mass | Non-English main-language mass (full) | Main-lang. locales | ||||
| Region | sample | full | growth | docs | share of region | w/ double-blind cells |
| Caribbean | 1.4M | 66.5M | 48.0 | 3.6M | 5% | 3 |
| Oceania | 13.8M | 719.0M | 52.1 | 309k | 1% | 0 |
| C. America | 372k | 18.8M | 50.5 | 33.9M | 60% | 0 |
| Middle Africa | 449k | 20.5M | 45.8 | 1.6M | 5% | 0 |
| C. Asia | 85k | 4.7M | 55.6 | 15.7M | 76% | 1 |
| Quantity | 100BT snap. | 350BT snap. | Full release |
|---|---|---|---|
| Total documents | 5,176,421,725 | 5,547,240,607 | 30,914,158,759 |
| Resolved (non-XX) | 2,960,753,780 | 3,030,341,027 | 7,917,968,305 |
| Resolution rate | 57.20% | 54.63% | 25.61% |
| Lookup-table coverage | 1,636,856,362 | 1,681,071,710 | 4,627,776,301 |
| Lookup-table rate | 31.62% | 30.30% | 14.97% |
| query (reported separately) | 7,206,682 | 7,314,090 | 15,283,520 |
| Artifact | Format | License |
|---|---|---|
| Region-annotated corpus | Parquet shards | ODC-By 1.0 |
| URL-signature lookup table | CSV | CC-BY 4.0 |
| Locale Topic Distributions | Parquet | CC-BY 4.0 |
| Per-document culture annotations | Parquet | ODC-By 1.0 |
| Benchmark survey | BibTeX + CSV | CC-BY 4.0 |
| Pipeline code | Git repository | Apache-2.0 |
| Element | Datasets |
|---|---|
| Ideational | |
| Concepts (21) | Hu et al. (2024) ; Cao et al. (2023) ; Nayak et al. (2024) ; Romero et al. (2024) ; Winata et al. (2025) ; Myung et al. (2024a) ; Nguyen et al. (2023b) ; Guo et al. (2025) ; Majewska et al. (2022) ; Li et al. (2024d) ; Li et al. (2024e) ; Yin et al. (2021) ; Koto et al. (2024b) ; Liu et al. (2021) ; Hu et al. (2023) ; Wang et al. (2021) ; Ayash et al. (2025) ; Becattini et al. (2023) ; Thapliyal et al. (2022) ; Li et al. (2023b) ; Rei et al. (2023) |
| Knowledge (32) | Romero et al. (2024) ; Koto et al. (2024a) ; Myung et al. (2024a) ; Joy (2025) ; Nguyen et al. (2023b) ; Li et al. (2024c) ; Wibowo et al. (2024) ; Chiu et al. (2024b) ; Li et al. (2024d) ; Shi et al. (2024) ; Keleg and Magdy (2023) ; Hardalov et al. (2020) ; Semenov and Sennrich (2025) ; Yin et al. (2022) ; Singh et al. (2025) ; Koto et al. (2024b) ; Koto et al. (2023) ; Son et al. (2024) ; FitzGerald et al. (2022) ; Sakai et al. (2024) ; Lin et al. (2021) ; Kassner et al. (2021) ; Xia and Monti (2021) ; Pramodya et al. (2025) ; Jiang et al. (2020) ; Ponti et al. (2020) ; Yuan et al. (2024) ; Liu et al. (2024) ; Singh et al. (2024) ; Röttger et al. (2024) ; Lewis et al. (2020) ; Clark et al. (2020) |
| Values - general (15) | Scaria et al. (2024) ; Li et al. (2024b) ; Kirk et al. (2024) ; Aakanksha et al. (2024) ; Pistilli et al. (2024) ; Karinshak et al. (2024) ; Zahraei and Asgari (2025) ; Xu et al. (2024) ; Cahyawijaya et al. (2025) ; Li et al. (2024a) ; Durmus et al. (2024) ; Zhao et al. (2024) ; Sorensen et al. (2024) ; Al Kautsar et al. (2025) ; Hanges and Dickson (2004) |
| Values - bias (28) | Zulaika and Saralegi (2025) ; Joshi et al. (2025) ; Ravikiran and Annamalai (2021) ; Palta and Rudinger (2023) ; Nozza et al. (2021) ; Neplenbroek et al. (2024) ; Wang et al. (2025) ; Bhutani et al. (2024) ; Mitchell et al. (2025) ; Nadeem et al. (2021) ; Öztürk et al. (2023) ; Troles and Schmid (2021) ; Stanovsky et al. (2019) ; Kappl (2025) ; Xu et al. (2024) ; Lauscher et al. (2020) ; Parrish et al. (2022a) ; Nangia et al. (2020a) ; Strazda and Spanakis (2025) ; Névéol et al. (2022) ; Sahoo et al. (2024) ; Jin et al. (2023) ; Mukherjee et al. (2023a) ; Zhao et al. (2018) ; Lauscher and Glavaš (2019) ; Tomar et al. (2025) ; Rudinger et al. (2018) ; Röttger et al. (2024) |
| Values - hate (13) | Joshi et al. (2025) ; Ravikiran and Annamalai (2021) ; Mandl et al. (2021) ; Bassignana et al. (2018) ; Trager et al. (2025) ; Luu et al. (2021) ; Lee et al. (2023) ; Vargas et al. (2022) ; Vargas et al. (2026) ; Mathew et al. (2020) ; Bui et al. (2025) ; Dementieva et al. (2025) ; Wulczyn et al. (2017) |