CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
Authors: Berke Arda, Ahmetcan Yavuz, Paul Gerry, Sebastian Lobentanzer, Nobin Sarwar, Joan Giner-Miguelez, Kongtao Chen, Luyao Zhang, +2 more
Organizations: ETH Zurich · CSAIL, MIT · Helmholtz Zentrum München · University of Maryland, Baltimore County · Barcelona Supercomputing Center · Google · Duke Kunshan University · ETH AI Center
Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.
Figures & tables
Figure 1: The CroissantMiner data-creation pipeline. (1) Corpus: ML dataset papers from top-downloaded Hugging Face datasets (vision, NLP, audio, etc.). (2) Extraction: a single-pass LLM generates 30-field Croissant metadata per paper, forming the silver split and the pre-fills for the gold split. (3) Annotation: annotators rate each pre-fill against the source paper on a 3-level rubric (Correct / Partially Correct / Not Correct) with failure-mode labels (e.g., hallucination, incomplete), producing 9,595 ratings over 3,060 cells. (4) Adjudication: majority vote resolves 95% of cells; senior-author review resolves the rest.
Figure 2: The four agentic architectures evaluated against single-pass extraction. All systems take a paper PDF as input and return a 30-field Croissant JSON-LD record (Section 4.1 ). They differ in how the extraction is split into LLM calls and code steps, and three of them can also fill some fields from outside metadata.
Candidate judge
cal.
val.
pooled
GLM-5
0.900
0.880
0.890
DeepSeek V3.2
0.828
0.882
0.855
GPT-5.5
0.784
0.904
0.847
Gemini 2.5 Pro †
0.900
0.743
0.823
Qwen 3 Max
0.784
0.857
0.823
Llama 4 Maverick
0.781
0.855
0.821
Table 1: Judge selection: Cohen’s quadratic-weighted κ against human consensus on 60 development cells (30 calibration, 30 validation).
Rank
System
Architecture
Core
RAI
Composite [95% CI]
1
Claude Sonnet 4.6 ∗
Single-Pass
0.752
0.687
0.709 [0.688, 0.729]
2
Claude Opus 4.7 ∗
Single-Pass
0.676
0.711
0.699 [0.665, 0.732]
3
GPT-5.4
Single-Pass
0.653
0.671
0.665 [0.648, 0.692]
4
Qwen 3.6 35B-A3B
Single-Pass
0.698
0.601
0.634 [0.615, 0.654]
5
GLM-5.1
Single-Pass
0.675
0.599
0.625 [0.604, 0.645]
6
Gemini 2.5 Flash
Single-Pass
0.615
0.616
0.616 [0.592, 0.639]
Table 2: Results on the test split , scored with the GLM-5 Tier-2 judge on audited gold samples (§ 5.1 ). Core averages the 10 core fields (rule-based), RAI the 20 RAI fields (LLM judge), and Composite weights all 30 fields equally (2,000-replicate bootstrap CIs). Ranks are reported per architecture type. Anthropic-family systems are marked ∗ . Claude Sonnet 4.5, which generated the gold pre-fills, is listed as an unranked diagnostic: its score largely reflects agreement with its own output and is not comparable with the other rows.
Figure 3: Failure-mode taxonomy across 1,592 error ratings. Cell counts per (field, failure-mode) bin; rows sorted by total descending within each panel. Incomplete extraction dominates (43%) for long-form RAI fields. The single counter-pattern is sc:publisher , where Wrong Section dominates: annotators disagree about whether the dataset host, venue, or author affiliation counts as the publisher.
Appendix figures & tables15 assets
Supplementary material from the paper’s appendix.
Appendix
Field
Type
Metric
Definition
Gold
Silver
name
Core
Constrained
Dataset name
100%
100%
description
Core
Short-text
Brief dataset description
100%
100%
url
Core
Constrained
Access URL
96%
86%
license
Core
Constrained
Distribution license
24%
19%
creator
Core
Short-text
Dataset creator(s)
100%
100%
publisher
Core
Constrained
Publishing venue/org
80%
74%
Appendix
Table 3: Overview of the 30 metadata fields in the Croissant 1.1 schema (10 core + 20 RAI). Coverage shows the percentage of datasets where each field has a non-null value in the human gold (N=102) and in the silver extraction (N=500).
Figure 4: Annotation reliability for each metadata field. Bars show Gwet’s AC 1 with 95% bootstrap confidence intervals over 3,060 cells (102 papers × 30 fields) derived from 9,595 ratings from 22 annotators. AC 1 is reported alongside Krippendorff’s α to account for skewed label distributions, where α can underestimate agreement. Background colors indicate Landis–Koch thresholds (fair ≥0.4 , moderate ≥0.6 , substantial ≥0.8 ). Agreement is high overall (95.0% majority consensus) but varies by field, with constrained fields showing higher agreement than open-ended fields.
Rank
System
Architecture
Core
RAI
Composite, pilot gold [95% CI]
Composite, current gold
1
Claude Sonnet 4.6 ∗
Single-Pass
0.712
0.666
0.681 [0.627, 0.749]
0.695
2
GPT-5.4 Mini
Single-Pass
0.702
0.665
0.678 [0.613, 0.733]
0.608
3
Claude Opus 4.7 ∗
Single-Pass
0.695
0.619
0.645 [0.585, 0.704]
0.708
4
GLM-5.1
Single-Pass
0.671
0.511
0.566 [0.503, 0.632]
0.581
5
Gemini 3.1 Pro Preview
Single-Pass
0.622
0.524
0.559 [0.508, 0.605]
0.546
6
Qwen 3.6 35B-A3B
Single-Pass
0.619
0.516
0.553 [0.499, 0.617]
0.591
Appendix
Table 4: Re-annotation pilot on 10 test papers, laid out as Table 2 . Core, RAI and the first composite are scored against the GPT-5.4-seeded gold; the last column is the composite against the current gold on the same papers. Systems are ranked by the pilot-gold composite within each architecture group, and ∗ marks Anthropic-family systems. The two seed models, GPT-5.4 for the pilot gold and Sonnet 4.5 for the current gold, are inflated against the gold they seeded and are listed separately without a rank. The two Gemini-based Locator-Extractor runs are omitted because their pilot scores include failed API calls that were later re-run.
CardGen
MOLE
HF Auto
DataDoc
CroissantMiner
Input source
Paper + repo
Paper
Data files
Paper
Paper
Output format
Free text
Structured
JSON-LD
Free text
JSON-LD
Target schema
None
Masader
Croissant
None
Croissant 1.1
Core fields
Partial
30+
✓
×
✓ (10)
RAI fields
Partial
×
×
Partial
✓ (20)
Croissant-conformant
×
×
×
×
✓
Appendix
Table 5: Comparison of automated dataset documentation systems.
Figure 5: Prompt used to extract Croissant metadata fields.
Figure 6: Prompt used for structured LLM-based judging (final rubric, used for all reported scores).
Field
Agree
Field
Agree
annotatorDemographics
10/10
dataAnnotationPlatform
7/10
annotationsPerItem
9/10
dataImputationProtocol
6/10
dataCollectionTimeframe
9/10
dataCollection
6/10
machineAnnotationTools
9/10
dataAnnotationAnalysis
6/10
dataSocialImpact
9/10
personalSensitiveInformation
6/10
dataBiases
9/10
dataCollectionRawData
5/10
Appendix
Table 6: Exact agreement between the deployed judge and human consensus on the 200-cell audit. Each entry gives the number of matching ratings out of 10 audited cells per RAI field.
System
name
description
url
license
creator
publisher
datePublished
inLanguage
citeAs
isLiveDataset
Claude Sonnet 4.6
0.98
0.70
0.85
0.42
0.74
0.62
0.90
0.70
0.76
0.85
Claude Opus 4.7
0.92
0.65
0.83
0.59
0.63
0.66
0.81
0.60
0.53
0.53
GPT-5.4
0.91
0.63
0.83
0.57
0.75
0.30
0.83
0.72
0.77
0.24
ReAct (4.6)
0.94
0.64
0.84
0.27
0.81
0.62
0.88
0.73
0.72
0.90
PSpec (4.6)
0.97
0.65
0.86
0.48
0.81
0.06
0.92
0.74
0.78
0.73
Qwen 3.6 35B-A3B
0.81
0.61
0.80
0.39
0.78
0.59
0.90
0.72
0.45
0.94
Appendix
Table 7: Per-field accuracy on the 88-paper test split for the 10 core (sc:/cr:) fields. Scores in [0,1]; Tier 1 rule-based metric (exact-match for constrained, token-F1 for short-text). Higher is better. The last row gives the number of test papers (of 88) whose gold documents the field.
System
(1)
(2)
(3)
(4)
(5)
(6)
(7)
(8)
(9)
(10)
(11)
(12)
(13)
(14)
(15)
(16)
(17)
(18)
(19)
(20)
Claude Sonnet 4.6
0.55
0.73
0.62
0.69
0.84
0.66
0.94
0.25
0.96
0.74
0.76
0.12
0.87
0.58
0.79
0.61
0.70
0.98
0.82
0.53
Claude Opus 4.7
0.54
0.76
0.63
0.75
0.82
0.65
0.93
0.50
0.90
0.83
0.71
0.33
0.81
0.57
0.75
0.66
0.75
0.95
0.81
0.55
GPT-5.4
0.56
0.64
0.64
0.68
0.84
0.67
0.98
0.40
0.93
0.79
0.73
0.00
0.84
0.60
0.71
0.71
0.67
0.99
0.65
0.40
ReAct (4.6)
0.49
0.56
0.55
0.67
0.81
0.66
0.94
0.31
0.89
0.77
0.64
0.00
0.83
0.20
0.79
0.34
0.68
0.95
0.80
0.33
PSpec (4.6)
0.34
0.46
0.59
0.68
0.67
0.59
0.96
0.22
0.95
0.68
0.73
0.00
0.85
0.47
0.79
0.63
0.76
0.97
0.74
0.33
Qwen 3.6 35B-A3B
0.46
0.69
0.44
0.51
0.77
0.57
0.89
0.50
0.84
0.64
0.61
0.12
0.74
0.51
0.61
0.56
0.64
0.86
0.61
0.47
Appendix
Table 8: Per-field accuracy on the 88-paper test split for the 20 RAI (rai:) fields. Scores in [0,1] from the GLM-5 LLM judge (1=Correct, 0.5=Partial, 0=Wrong). Higher is better. The last row gives the number of test papers (of 88) whose gold documents the field; scores on fields with small support, such as (12) and (8), are unstable. Field key: (1) annotationsPerItem, (2) annotatorDemographics, (3) dataAnnotationAnalysis, (4) dataAnnotationPlatform, (5) dataAnnotationProtocol, (6) dataBiases, (7) dataCollection, (8) dataCollectionMissingData, (9) dataCollectionRawData, (10) dataCollectionTimeframe, (11) dataCollectionType, (12) dataImputationProtocol, (13) dataLimitations, (14) dataManipulationProtocol, (15) dataPreprocessingProtocol, (16) dataReleaseMaintenancePlan, (17) dataSocialImpact, (18) dataUseCases, (19) machineAnnotationTools, (20) personalSensitiveInformation.
Figure 7: Croissant field documentation rates across the 602-paper CroissantMiner benchmark. Each bar shows the percentage of papers with a non-null value for the field. Core fields are consistently reported (often higher than 99%), while RAI fields show substantial variability, ranging from widely documented (e.g., use cases, limitations) to rarely reported (e.g., imputation, missing data). This highlights a large gap between core metadata coverage and RAI documentation in current dataset papers.
Figure 8: Silver-split fill rates track gold-split fill rates. Each point is one of the 30 Croissant fields; the x -coordinate is the fraction of the 102 human-validated gold papers in which the field is populated, and the y -coordinate is the same fraction in the 500 silver papers (Sonnet 4.5 reference). Pearson r=0.839 , Spearman ρ=0.823 ( n=30 , p<10−7 ). Core fields (blue circles) and RAI fields (orange squares) lie mostly on or below the diagonal: the silver extraction documents most fields less often than the human gold. The labelled outlier is cr:isLiveDataset (gold 100%, silver 7.8%): human annotators are instructed to decide for every paper, while the silver model only emits this field when the text explicitly references live updates.
Model
Parser
Core
RAI
Overall
Sonnet 4.6
PyPDF2 (ours)
0.770
0.475
0.573
PyMuPDF
0.748
0.475
0.566
Docling, full export †
0.749
0.472
0.564
Docling, body only †
0.670
0.478
0.542
GPT-5.4
PyPDF2 (ours)
0.670
0.464
0.532
PyMuPDF
0.670
0.473
0.541
Appendix
Table 9: PDF parser ablation. Same prompt and scoring pipeline; only the text extractor varies. 30 test-split papers (seed 42). Docling uses the default converter with the full export (body, page headers/footers, figure text); grey rows use Docling’s default body-only export. RAI cells for all rows were judged in one session, so compare within this table only. † 29 papers: Docling’s default parser resolves one paper’s legacy Type 3 glyph names as ZapfDingbats symbols (its pdfium backend reads it correctly); Sonnet 4.6 refuses the resulting text.
Error type
ReAct (n=30)
Parallel Specialists (n=30)
Triage + Critique (n=27)
Locator-Extractor (n=30)
Total (n=117)
left empty although documented
15 (50%)
13 (43%)
13 (48%)
10 (33%)
51 (44%)
incomplete
8 (27%)
10 (33%)
9 (33%)
13 (43%)
40 (34%)
wrong detail
5 (17%)
5 (17%)
3 (11%)
5 (17%)
18 (15%)
judge noise
2 (7%)
2 (7%)
2 (7%)
0
6 (5%)
other (verbose drift, over-hedged paraphrase)
0
0
0
2 (7%)
2 (2%)
Appendix
Table 10: Error types in 117 sampled RAI cells where an agentic system using Sonnet 4.6 scored below single-pass Sonnet 4.6 on the same paper and field. Each cell receives one label, with the reference answer treated as authoritative. Percentages are calculated within each column and rounded to the nearest integer.
Model
Composite [95% CI]
Core
RAI
Qwen3-4B
0.376 [0.359, 0.394]
0.586
0.272
Qwen3-8B
0.431 [0.411, 0.452]
0.577
0.358
Qwen3-14B
0.482 [0.458, 0.506]
0.659
0.393
Qwen3-32B
0.528 [0.505, 0.554]
0.615
0.484
Appendix
Table 11: Qwen3 models of increasing size as single-pass extractors, scored on the same cells of 66 test papers.
System
Architecture
USD per paper
vs. single-pass
Claude Sonnet 4.6
Single-Pass
0.126
Claude Opus 4.7
Single-Pass
0.269
GPT-5.4
Single-Pass
0.093
Qwen 3.6 35B-A3B
Single-Pass
0 †
GLM-5.1
Single-Pass
0.063
Gemini 2.5 Flash
Single-Pass
0.013 ‡
Appendix
Table 12: Estimated mean per-paper inference costs for the systems in Table 2 , based on recorded token usage on the 88-paper test split and standard public list prices at the time of the experiments. Means exclude papers with unrecorded token usage. Ratios compare each system with single-pass extraction on the same backbone; the hybrid is compared with single-pass Gemini 3.1 Pro. † Self-hosted on an institutional cluster with no API charge; compute costs are not included. ‡ Unrecorded Gemini thinking tokens are excluded, understating output-token costs. Inputs are priced as uncached where cached-token counts were not recorded. § ReAct sends a growing prompt at each turn. If all prompt tokens after the first turn were cache hits, the estimated costs would be 0.146forGPT−5.4and0.103 for Gemini 3.1 Pro, still excluding Gemini thinking tokens.
Croissant has emerged as the metadata standard for machine learning datasets, providing a structured, JSON-LD-based format that makes dataset discovery, automated ingestion, and reproducible analysis machine-checkable across ML platforms. Adoption has accelerated, and NeurIPS now requires Croissant metadata in every submission to its dataset tracks. Yet in practice Croissant generation usually starts with uploading data to a public platform, a path infeasible for governed and large local repositories that hold much of the high-value data ML increasingly relies on. We release Croissant Baker, a local-first, open-source command-line tool that generates validated Croissant metadata directly from a dataset directory through a modular handler registry. We evaluate Croissant Baker on over 140 datasets, scaling to MIMIC-IV at 886 million rows and 374 Parquet files. On held-out comparisons against producer-authored or standards-derived ground truth, Croissant Baker reaches 97-100% agreement across multiple domains.
Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer +17
Technical University of Munich · Massachusetts Institute of Technology · Helmholtz Munich +15
Reproducibility is fundamental to the scientific method, yet remains a critical challenge in machine learning. Contributing factors include underspecified execution details and brittle software environments. Human-centric remedies, such as checklists and manual verification, help but require intensive effort and fail to scale. To address this, we introduce Croissant Tasks: a declarative, machine-actionable metadata format that abstracts low-level implementation details into high-level specifications. This format enables conceptual reproducibility: verifying claims via independent, agent-generated implementations rather than brittle source code replication. We contribute: (1) the Croissant Tasks specification, formally decoupling task problem from solution; (2) an automated LLM pipeline that retrofits existing benchmarks into this format; and (3) empirical validation showing autonomous agents can ingest these specifications to generate functional, accurate reproduction pipelines from scratch. We envision this format as a new foundation for automated and conceptual reproducibility in machine learning.
Croissant is the de facto machine-readable descriptor for ML datasets: JSON-LD over schema.org. Since version 1.1 it also carries data use conditions, recommending DUO and ODRL for them. What no version specifies is how any of them is evaluated: no decision procedure, no bound on evaluation cost, no outcome for a condition an implementation cannot evaluate, no record of what was checked, and nothing on composition with caller-side authority. We supply that half. An additive profile lets a dataset declare the operations it admits and the conditions under which it admits them, over a closed set of five operators whose decision procedure is given in full, so a gate decides from the descriptor alone and records what it checked. Two corpora evaluate it and their evidence is kept apart. Three descriptors that gated a real nf-core pipeline give the deployment result: decisions from a profile document match the gate's native descriptor record for record, stripping the layer leaves a valid Croissant document, and the added cost is 11.7 μs against a 119 μs decision. A corpus generated from the profile's grammar gives the breadth, covering every operator, refusal class and conformance clause. Across its valid cases, 552 complete decision records agree three ways -- native descriptor, profile terms, and the same policy as ODRL in usageInfo. The carrier is therefore not the contribution; the evaluation semantics is. Finally, caller-bound and data-bound policies range over non-overlapping state spaces, so neither permit set contains the other.