CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
Authors: Berke Arda, Ahmetcan Yavuz, Paul Gerry, Sebastian Lobentanzer, Nobin Sarwar, Joan Giner-Miguelez, Kongtao Chen, Luyao Zhang, +2 more
Organizations: ETH Zurich · CSAIL, MIT · Helmholtz Zentrum München · University of Maryland, Baltimore County · Barcelona Supercomputing Center · Google · Duke Kunshan University · ETH AI Center
Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.
Figures & tables
Figure 1: The CroissantMiner data-creation pipeline. (1) Corpus: ML dataset papers from top-downloaded Hugging Face datasets (vision, NLP, audio, etc.). (2) Extraction: a single-pass LLM generates 30-field Croissant metadata per paper, forming the silver split and the pre-fills for the gold split. (3) Annotation: annotators rate each pre-fill against the source paper on a 3-level rubric (Correct / Partially Correct / Not Correct) with failure-mode labels (e.g., hallucination, incomplete), producing 9,595 ratings over 3,060 cells. (4) Adjudication: majority vote resolves 95% of cells; senior-author review resolves the rest.
Figure 2: The four agentic architectures evaluated against single-pass extraction. All systems take a paper PDF as input and return a 30-field Croissant JSON-LD record (Section 4.1 ). They differ in how the extraction is split into LLM calls and code steps, and three of them can also fill some fields from outside metadata.
Candidate judge
cal.
val.
pooled
GLM-5
0.900
0.880
0.890
DeepSeek V3.2
0.828
0.882
0.855
GPT-5.5
0.784
0.904
0.847
Gemini 2.5 Pro †
0.900
0.743
0.823
Qwen 3 Max
0.784
0.857
0.823
Llama 4 Maverick
0.781
0.855
0.821
Table 1: Judge selection: Cohen’s quadratic-weighted κ against human consensus on 60 development cells (30 calibration, 30 validation).
Rank
System
Architecture
Core
RAI
Composite [95% CI]
1
Claude Sonnet 4.6 ∗
Single-Pass
0.752
0.687
0.709 [0.688, 0.729]
2
Claude Opus 4.7 ∗
Single-Pass
0.676
0.711
0.699 [0.665, 0.732]
3
GPT-5.4
Single-Pass
0.653
0.671
0.665 [0.648, 0.692]
4
Qwen 3.6 35B-A3B
Single-Pass
0.698
0.601
0.634 [0.615, 0.654]
5
GLM-5.1
Single-Pass
0.675
0.599
0.625 [0.604, 0.645]
6
Gemini 2.5 Flash
Single-Pass
0.615
0.616
0.616 [0.592, 0.639]
Table 2: Results on the test split , scored with the GLM-5 Tier-2 judge on audited gold samples (§ 5.1 ). Core averages the 10 core fields (rule-based), RAI the 20 RAI fields (LLM judge), and Composite weights all 30 fields equally (2,000-replicate bootstrap CIs). Ranks are reported per architecture type. Anthropic-family systems are marked ∗ . Claude Sonnet 4.5, which generated the gold pre-fills, is listed as an unranked diagnostic: its score largely reflects agreement with its own output and is not comparable with the other rows.
Figure 3: Failure-mode taxonomy across 1,592 error ratings. Cell counts per (field, failure-mode) bin; rows sorted by total descending within each panel. Incomplete extraction dominates (43%) for long-form RAI fields. The single counter-pattern is sc:publisher , where Wrong Section dominates: annotators disagree about whether the dataset host, venue, or author affiliation counts as the publisher.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Type
Metric
Definition
Gold
Silver
name
Core
Constrained
Dataset name
100%
100%
description
Core
Short-text
Brief dataset description
100%
100%
url
Core
Constrained
Access URL
96%
86%
license
Core
Constrained
Distribution license
24%
19%
creator
Core
Short-text
Dataset creator(s)
100%
100%
publisher
Core
Constrained
Publishing venue/org
80%
74%
Appendix
Table 3: Overview of the 30 metadata fields in the Croissant 1.1 schema (10 core + 20 RAI). Coverage shows the percentage of datasets where each field has a non-null value in the human gold (N=102) and in the silver extraction (N=500).
Figure 4: Annotation reliability for each metadata field. Bars show Gwet’s AC 1 with 95% bootstrap confidence intervals over 3,060 cells (102 papers × 30 fields) derived from 9,595 ratings from 22 annotators. AC 1 is reported alongside Krippendorff’s α to account for skewed label distributions, where α can underestimate agreement. Background colors indicate Landis–Koch thresholds (fair ≥0.4 , moderate ≥0.6 , substantial ≥0.8 ). Agreement is high overall (95.0% majority consensus) but varies by field, with constrained fields showing higher agreement than open-ended fields.
Rank
System
Architecture
Core
RAI
Composite, pilot gold [95% CI]
Composite, current gold
1
Claude Sonnet 4.6 ∗
Single-Pass
0.712
0.666
0.681 [0.627, 0.749]
0.695
2
GPT-5.4 Mini
Single-Pass
0.702
0.665
0.678 [0.613, 0.733]
0.608
3
Claude Opus 4.7 ∗
Single-Pass
0.695
0.619
0.645 [0.585, 0.704]
0.708
4
GLM-5.1
Single-Pass
0.671
0.511
0.566 [0.503, 0.632]
0.581
5
Gemini 3.1 Pro Preview
Single-Pass
0.622
0.524
0.559 [0.508, 0.605]
0.546
6
Qwen 3.6 35B-A3B
Single-Pass
0.619
0.516
0.553 [0.499, 0.617]
0.591
Appendix
Table 4: Re-annotation pilot on 10 test papers, laid out as Table 2 . Core, RAI and the first composite are scored against the GPT-5.4-seeded gold; the last column is the composite against the current gold on the same papers. Systems are ranked by the pilot-gold composite within each architecture group, and ∗ marks Anthropic-family systems. The two seed models, GPT-5.4 for the pilot gold and Sonnet 4.5 for the current gold, are inflated against the gold they seeded and are listed separately without a rank. The two Gemini-based Locator-Extractor runs are omitted because their pilot scores include failed API calls that were later re-run.
CardGen
MOLE
HF Auto
DataDoc
CroissantMiner
Input source
Paper + repo
Paper
Data files
Paper
Paper
Output format
Free text
Structured
JSON-LD
Free text
JSON-LD
Target schema
None
Masader
Croissant
None
Croissant 1.1
Core fields
Partial
30+
✓
×
✓ (10)
RAI fields
Partial
×
×
Partial
✓ (20)
Croissant-conformant
×
×
×
×
✓
Appendix
Table 5: Comparison of automated dataset documentation systems.
Figure 5: Prompt used to extract Croissant metadata fields.
Figure 6: Prompt used for structured LLM-based judging (final rubric, used for all reported scores).
Field
Agree
Field
Agree
annotatorDemographics
10/10
dataAnnotationPlatform
7/10
annotationsPerItem
9/10
dataImputationProtocol
6/10
dataCollectionTimeframe
9/10
dataCollection
6/10
machineAnnotationTools
9/10
dataAnnotationAnalysis
6/10
dataSocialImpact
9/10
personalSensitiveInformation
6/10
dataBiases
9/10
dataCollectionRawData
5/10
Appendix
Table 6: Exact agreement between the deployed judge and human consensus on the 200-cell audit. Each entry gives the number of matching ratings out of 10 audited cells per RAI field.
System
name
description
url
license
creator
publisher
datePublished
inLanguage
citeAs
isLiveDataset
Claude Sonnet 4.6
0.98
0.70
0.85
0.42
0.74
0.62
0.90
0.70
0.76
0.85
Claude Opus 4.7
0.92
0.65
0.83
0.59
0.63
0.66
0.81
0.60
0.53
0.53
GPT-5.4
0.91
0.63
0.83
0.57
0.75
0.30
0.83
0.72
0.77
0.24
ReAct (4.6)
0.94
0.64
0.84
0.27
0.81
0.62
0.88
0.73
0.72
0.90
PSpec (4.6)
0.97
0.65
0.86
0.48
0.81
0.06
0.92
0.74
0.78
0.73
Qwen 3.6 35B-A3B
0.81
0.61
0.80
0.39
0.78
0.59
0.90
0.72
0.45
0.94
Appendix
Table 7: Per-field accuracy on the 88-paper test split for the 10 core (sc:/cr:) fields. Scores in [0,1]; Tier 1 rule-based metric (exact-match for constrained, token-F1 for short-text). Higher is better. The last row gives the number of test papers (of 88) whose gold documents the field.
System
(1)
(2)
(3)
(4)
(5)
(6)
(7)
(8)
(9)
(10)
(11)
(12)
(13)
(14)
(15)
(16)
(17)
(18)
(19)
(20)
Claude Sonnet 4.6
0.55
0.73
0.62
0.69
0.84
0.66
0.94
0.25
0.96
0.74
0.76
0.12
0.87
0.58
0.79
0.61
0.70
0.98
0.82
0.53
Claude Opus 4.7
0.54
0.76
0.63
0.75
0.82
0.65
0.93
0.50
0.90
0.83
0.71
0.33
0.81
0.57
0.75
0.66
0.75
0.95
0.81
0.55
GPT-5.4
0.56
0.64
0.64
0.68
0.84
0.67
0.98
0.40
0.93
0.79
0.73
0.00
0.84
0.60
0.71
0.71
0.67
0.99
0.65
0.40
ReAct (4.6)
0.49
0.56
0.55
0.67
0.81
0.66
0.94
0.31
0.89
0.77
0.64
0.00
0.83
0.20
0.79
0.34
0.68
0.95
0.80
0.33
PSpec (4.6)
0.34
0.46
0.59
0.68
0.67
0.59
0.96
0.22
0.95
0.68
0.73
0.00
0.85
0.47
0.79
0.63
0.76
0.97
0.74
0.33
Qwen 3.6 35B-A3B
0.46
0.69
0.44
0.51
0.77
0.57
0.89
0.50
0.84
0.64
0.61
0.12
0.74
0.51
0.61
0.56
0.64
0.86
0.61
0.47
Appendix
Table 8: Per-field accuracy on the 88-paper test split for the 20 RAI (rai:) fields. Scores in [0,1] from the GLM-5 LLM judge (1=Correct, 0.5=Partial, 0=Wrong). Higher is better. The last row gives the number of test papers (of 88) whose gold documents the field; scores on fields with small support, such as (12) and (8), are unstable. Field key: (1) annotationsPerItem, (2) annotatorDemographics, (3) dataAnnotationAnalysis, (4) dataAnnotationPlatform, (5) dataAnnotationProtocol, (6) dataBiases, (7) dataCollection, (8) dataCollectionMissingData, (9) dataCollectionRawData, (10) dataCollectionTimeframe, (11) dataCollectionType, (12) dataImputationProtocol, (13) dataLimitations, (14) dataManipulationProtocol, (15) dataPreprocessingProtocol, (16) dataReleaseMaintenancePlan, (17) dataSocialImpact, (18) dataUseCases, (19) machineAnnotationTools, (20) personalSensitiveInformation.
Figure 7: Croissant field documentation rates across the 602-paper CroissantMiner benchmark. Each bar shows the percentage of papers with a non-null value for the field. Core fields are consistently reported (often higher than 99%), while RAI fields show substantial variability, ranging from widely documented (e.g., use cases, limitations) to rarely reported (e.g., imputation, missing data). This highlights a large gap between core metadata coverage and RAI documentation in current dataset papers.
Figure 8: Silver-split fill rates track gold-split fill rates. Each point is one of the 30 Croissant fields; the x -coordinate is the fraction of the 102 human-validated gold papers in which the field is populated, and the y -coordinate is the same fraction in the 500 silver papers (Sonnet 4.5 reference). Pearson r=0.839 , Spearman ρ=0.823 ( n=30 , p<10−7 ). Core fields (blue circles) and RAI fields (orange squares) lie mostly on or below the diagonal: the silver extraction documents most fields less often than the human gold. The labelled outlier is cr:isLiveDataset (gold 100%, silver 7.8%): human annotators are instructed to decide for every paper, while the silver model only emits this field when the text explicitly references live updates.
Model
Parser
Core
RAI
Overall
Sonnet 4.6
PyPDF2 (ours)
0.770
0.475
0.573
PyMuPDF
0.748
0.475
0.566
Docling, full export †
0.749
0.472
0.564
Docling, body only †
0.670
0.478
0.542
GPT-5.4
PyPDF2 (ours)
0.670
0.464
0.532
PyMuPDF
0.670
0.473
0.541
Appendix
Table 9: PDF parser ablation. Same prompt and scoring pipeline; only the text extractor varies. 30 test-split papers (seed 42). Docling uses the default converter with the full export (body, page headers/footers, figure text); grey rows use Docling’s default body-only export. RAI cells for all rows were judged in one session, so compare within this table only. † 29 papers: Docling’s default parser resolves one paper’s legacy Type 3 glyph names as ZapfDingbats symbols (its pdfium backend reads it correctly); Sonnet 4.6 refuses the resulting text.
Error type
ReAct (n=30)
Parallel Specialists (n=30)
Triage + Critique (n=27)
Locator-Extractor (n=30)
Total (n=117)
left empty although documented
15 (50%)
13 (43%)
13 (48%)
10 (33%)
51 (44%)
incomplete
8 (27%)
10 (33%)
9 (33%)
13 (43%)
40 (34%)
wrong detail
5 (17%)
5 (17%)
3 (11%)
5 (17%)
18 (15%)
judge noise
2 (7%)
2 (7%)
2 (7%)
0
6 (5%)
other (verbose drift, over-hedged paraphrase)
0
0
0
2 (7%)
2 (2%)
Appendix
Table 10: Error types in 117 sampled RAI cells where an agentic system using Sonnet 4.6 scored below single-pass Sonnet 4.6 on the same paper and field. Each cell receives one label, with the reference answer treated as authoritative. Percentages are calculated within each column and rounded to the nearest integer.
Model
Composite [95% CI]
Core
RAI
Qwen3-4B
0.376 [0.359, 0.394]
0.586
0.272
Qwen3-8B
0.431 [0.411, 0.452]
0.577
0.358
Qwen3-14B
0.482 [0.458, 0.506]
0.659
0.393
Qwen3-32B
0.528 [0.505, 0.554]
0.615
0.484
Appendix
Table 11: Qwen3 models of increasing size as single-pass extractors, scored on the same cells of 66 test papers.
System
Architecture
USD per paper
vs. single-pass
Claude Sonnet 4.6
Single-Pass
0.126
Claude Opus 4.7
Single-Pass
0.269
GPT-5.4
Single-Pass
0.093
Qwen 3.6 35B-A3B
Single-Pass
0 †
GLM-5.1
Single-Pass
0.063
Gemini 2.5 Flash
Single-Pass
0.013 ‡
Appendix
Table 12: Estimated mean per-paper inference costs for the systems in Table 2 , based on recorded token usage on the 88-paper test split and standard public list prices at the time of the experiments. Means exclude papers with unrecorded token usage. Ratios compare each system with single-pass extraction on the same backbone; the hybrid is compared with single-pass Gemini 3.1 Pro. † Self-hosted on an institutional cluster with no API charge; compute costs are not included. ‡ Unrecorded Gemini thinking tokens are excluded, understating output-token costs. Inputs are priced as uncached where cached-token counts were not recorded. § ReAct sends a growing prompt at each turn. If all prompt tokens after the first turn were cache hits, the estimated costs would be 0.146forGPT−5.4and0.103 for Gemini 3.1 Pro, still excluding Gemini thinking tokens.