cs.CLOct 5, 2026

CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

Authors: Berke Arda, Ahmetcan Yavuz, Paul Gerry, Sebastian Lobentanzer, Nobin Sarwar, Joan Giner-Miguelez, Kongtao Chen, Luyao Zhang, +2 more

Organizations: ETH Zurich · CSAIL, MIT · Helmholtz Zentrum München · University of Maryland, Baltimore County · Barcelona Supercomputing Center · Google · Duke Kunshan University · ETH AI Center

Abstract

Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

Figures & tables

Appendix figures & tables15 assets

Supplementary material from the paper’s appendix.

Appendix

Explore similar work

CardsList
  1. Croissant Baker: Metadata Generation for Discoverable, Governable, and Reusable ML Datasets

    May 14, 2026Rafi Al Attrach, Rajna Fani, Sebastian Lobentanzer +17MetadataTaco

  2. Croissant Tasks: A Metadata Format for Reproducible Machine Learning Evaluations

    May 28, 2026Omar Benjelloun, Leonardo Martins Bianco, Isabelle Guyon +8ReproducibilityMle-Bench Lite

  3. A Policy Profile for Croissant: Refusal as a Property of the Dataset

    Sep 17, 2026Alexander ChernovGeneration ProvenanceTaco