Skip to main content

CroissantMiner

Extract Croissant metadata, including the 20 Responsible AI fields, from the paper that introduces an ML dataset.

Demo Dataset Tests Python 3.10 to 3.13 License: MIT

Code, benchmark and systems of CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets (NeurIPS 2026, Evaluations and Datasets Track). Paper: arXiv link follows.

News

  • 1 Oct 2026: on PyPI (pip install croissantminer), with --merge-into to add the fields to a dataset's Croissant file on Hugging Face, a guide for NeurIPS dataset submissions and a leaderboard open to new systems.
  • 30 Sep 2026: croissantminer extract, one command from a paper to a Croissant file.
  • 28 Sep 2026: code, dataset and demo released.
  • 24 Sep 2026: accepted at NeurIPS 2026 (Evaluations and Datasets Track).

Documenting a dataset in the Croissant format, including its Responsible AI (RAI) fields, takes time, and venues such as the NeurIPS Evaluations and Datasets Track ask for it. CroissantMiner reads the paper that introduces a dataset and drafts all 30 fields of the Croissant 1.1 schema: 10 core fields (name, license, creators and others) and 20 RAI fields (how the data was collected and annotated, known biases, limitations, intended uses and others). You check the draft and publish it.

  • For dataset authors: one command turns a paper into a Croissant file that passes the MLCommons validator.
  • For researchers: a benchmark of 602 dataset papers with human gold annotations for 102 of them, the outputs and scores of 24 extraction systems, and a leaderboard that scores new ones with the paper's scorer.

Try it in your browser

The demo runs the six systems below on a PDF you upload. It needs your own Anthropic or OpenAI API key, which is sent only to that provider and not stored.

The demo after extracting the GSM8K paper with Triage + Critique: 19 of 30 fields, each with a supporting quote from the paper

Quick start

Python 3.10 to 3.13.

pip install "croissantminer[validate]"   # the extraction tool and the Croissant validator
export ANTHROPIC_API_KEY=...              # or put it in a .env file in the folder you run from
croissantminer extract paper.pdf
Extracting with Single-pass · Claude Sonnet 4.6 (usually about 30 s)...
Found 22 of 30 fields (core 9 of 10, Responsible AI 13 of 20) with single-pass in 35 s, about $0.06.
Not found: license, rai:dataCollectionMissingData, rai:dataCollectionTimeframe, ...
Wrote:  paper.croissant.json (Croissant 1.1)
Check:  passes the mlcroissant validator (2 recommended properties missing)
These are drafts by a language model: check each value against the paper before publishing.
Example output for GSM8K (--hf-id openai/gsm8k), shortened
{
  "@context": {
    "@language": "en",
    "@vocab": "https://schema.org/",
    "sc": "https://schema.org/",
    "cr": "http://mlcommons.org/croissant/",
    "rai": "http://mlcommons.org/croissant/RAI/",
    "dct": "http://purl.org/dc/terms/",
    "conformsTo": "dct:conformsTo"
  },
  "@type": "sc:Dataset",
  "conformsTo": "http://mlcommons.org/croissant/1.1",
  "@id": "https://huggingface.co/datasets/openai/gsm8k",
  "name": "GSM8K",
  "url": "https://github.com/openai/grade-school-math",
  "publisher": {"@type": "Organization", "name": "OpenAI"},
  "datePublished": "2021-11-18",
  "rai:dataCollection": "Problems were initially collected by hiring freelance ...",
  "rai:dataCollectionType": "Manual Human Curator, Others",
  "rai:dataAnnotationPlatform": "Upwork (upwork.com) for initial collection; Surge AI ...",
  "rai:dataBiases": "Seed questions used to assist contractors were automatically ...",
  // and description, inLanguage, cr:citeAs, creator, cr:isLiveDataset and 9 more rai: fields
}
Option What it does
-o my_dataset.json choose the output file
--method react use another system (list them with croissantminer methods)
--hf-id org/name give the dataset's Hugging Face id, so the agentic systems can check its license and URL
--card README.md read the dataset card together with the paper
--fields values.json also save the extracted values with their supporting quotes
--merge-into org/name add the fields to the dataset's Croissant file on Hugging Face (see Using the file)

croissantminer validate my_dataset.json checks any Croissant file with the MLCommons validator.

From Python:

from croissantminer import extract

result = extract("paper.pdf")            # method="single-pass" by default
print(result.summary())                  # Found 22 of 30 fields (core 9 of 10, Responsible AI 13 of 20)
result.fields["rai:dataCollection"]      # one extracted value
result.croissant                         # the Croissant 1.1 file as a dict

To change the code, install from a clone instead: git clone https://github.com/berkearda/croissantminer, then pip install -e ".[validate]" in that folder.

Which method to choose

The five system designs: single-pass extraction, Parallel Specialists, Triage + Critique, Locator-Extractor and a ReAct agent

The five system designs, from the paper. Single-pass reads the whole paper in one model call; the four agentic systems split the work into steps.

All six methods are systems from the paper, with the same code and settings. Score: the composite over the 30 fields on the 88 test papers (see Results). Time and cost: one run on the 22-page GSM8K paper.

Method Model Score Time Cost API key
single-pass (default) Claude Sonnet 4.6 0.709 35 s $0.06 ANTHROPIC_API_KEY
single-pass-gpt GPT-5.4 0.665 about 30 s not measured OPENAI_API_KEY
react Claude Sonnet 4.6 0.652 80 s $0.17 ANTHROPIC_API_KEY
parallel-specialists Claude Sonnet 4.6 0.647 15 s $0.42 ANTHROPIC_API_KEY
triage-critique Claude Sonnet 4.6 0.624 35 s $0.08 ANTHROPIC_API_KEY
locator-extractor Claude Sonnet 4.6 0.566 40 s $0.09 ANTHROPIC_API_KEY

Start with single-pass: it is the most accurate and among the cheapest. triage-critique and locator-extractor return a supporting quote for most values (--fields), which makes checking faster, and react gives a reason for each field it leaves empty.

Using the file

The file describes the dataset: the core fields and the Responsible AI fields the paper supports. It does not list the data files and their columns (Croissant's distribution and recordSet), which a data host generates from the files themselves. Hosts such as Hugging Face, Kaggle and OpenML publish such a file for their datasets, but without the Responsible AI fields. CroissantMiner adds its fields to the host's file:

croissantminer extract paper.pdf --merge-into org/name     # extract, then merge into the file Hugging Face generates
croissantminer merge org/name paper.croissant.json         # merge a file you already extracted
croissantminer merge host_croissant.json paper.croissant.json   # any host's file, by path or URL

The merged file keeps everything the host wrote (name, URL, license, files and columns) and adds the Responsible AI fields and any core field the host lacks; a value the host already has is never replaced. On GSM8K, the merged file kept Hugging Face's 3 data files and 4 record sets, gained 12 Responsible AI fields and passed the validator. A private or gated Hugging Face dataset needs HF_TOKEN.

Submitting a dataset to NeurIPS? The step-by-step guide covers the Croissant file the Evaluations and Datasets Track requires, including the three Responsible AI items you add yourself.

Before you publish the file

  • Check every value against the paper. The fields are drafts. Typical mistakes are a value the paper does not state, a detail from a related dataset, or, for an anonymous submission, the page header taken as the publisher.
  • Empty fields are left out of the Croissant file, never filled with placeholders. Add what you know.
  • The validator checks the format, not the content. A file that passes can still contain wrong values.
  • Your paper is sent to the model provider (Anthropic or OpenAI) under your API key and their terms.
  • API keys are read from the environment or a .env file, never from the command line, so they do not end up in your shell history.

The benchmark

How the benchmark was built: corpus, extraction, human annotation, adjudication to gold

How the benchmark was built, from the paper: 602 dataset papers, drafts of all 30 fields by Claude Sonnet 4.5, 9,595 ratings by 22 annotators, and a majority vote or an expert decision for each of the 3,060 gold cells.

  • Papers: 602 dataset papers. 102 have human-validated gold annotations (3,060 cells, 22 annotators) and 500 have LLM-generated silver annotations. The 102 gold papers are split into 14 development and 88 test papers.
  • Systems: single-pass extraction and four agentic architectures (ReAct, Parallel Specialists, Triage + Critique, Locator-Extractor), each with several LLM backbones.
  • Evaluation: rule-based scoring for the 10 core fields, an LLM judge (GLM-5) for the 20 RAI fields, and tests that check the scorer against the numbers in the paper.
  • Data: the annotations, system outputs and judge verdicts are on Hugging Face and in data/ (see data/README.md).

Results

Test split (88 papers). Core averages the 10 core fields, RAI the 20 RAI fields, and Composite weights all 30 fields equally; 95% confidence intervals come from 2,000 bootstrap samples over papers. The gold annotations were first drafted by Claude Sonnet 4.5 and then checked and corrected by annotators, so Anthropic-family systems are marked with *. Claude Sonnet 4.5 itself is shown for reference and not ranked.

System Architecture Core RAI Composite [95% CI]
Claude Sonnet 4.6* Single-pass 0.752 0.687 0.709 [0.688, 0.729]
Claude Opus 4.7* Single-pass 0.676 0.711 0.699 [0.665, 0.732]
GPT-5.4 Single-pass 0.653 0.671 0.665 [0.648, 0.692]
Qwen 3.6 35B-A3B Single-pass 0.698 0.601 0.634 [0.615, 0.654]
GLM-5.1 Single-pass 0.675 0.599 0.625 [0.604, 0.645]
Gemini 2.5 Flash Single-pass 0.615 0.616 0.616 [0.592, 0.639]
GPT-5.4 Mini Single-pass 0.561 0.614 0.596 [0.573, 0.620]
Gemini 3.1 Pro Preview Single-pass 0.577 0.596 0.590 [0.571, 0.609]
DeepSeek V3.2 Single-pass 0.617 0.555 0.575 [0.547, 0.604]
Mistral Small 4 Single-pass 0.616 0.482 0.527 [0.510, 0.544]
Llama 4 Scout 17B Single-pass 0.506 0.334 0.391 [0.374, 0.409]
ReAct (Sonnet 4.6)* ReAct 0.734 0.610 0.652 [0.626, 0.687]
ReAct (GPT-5.4) ReAct 0.688 0.603 0.631 [0.601, 0.663]
ReAct (Gemini 3.1 Pro) ReAct 0.723 0.511 0.582 [0.556, 0.605]
Parallel Specialists (Sonnet 4.6)* Parallel Specialists 0.699 0.621 0.647 [0.626, 0.667]
Parallel Specialists (GPT-5.4) Parallel Specialists 0.627 0.570 0.589 [0.572, 0.611]
Parallel Specialists (Gemini 3.1 Pro) Parallel Specialists 0.607 0.505 0.539 [0.518, 0.559]
Triage + Critique (Sonnet 4.6)* Triage + Critique 0.675 0.599 0.624 [0.606, 0.653]
Triage + Critique (GPT-5.4) Triage + Critique 0.592 0.540 0.557 [0.537, 0.585]
Triage + Critique (Gemini 3.1 Pro) Triage + Critique 0.513 0.397 0.436 [0.415, 0.457]
Locator-Extractor (Sonnet 4.6)* Locator-Extractor 0.643 0.528 0.566 [0.543, 0.592]
Locator-Extractor (GPT-5.4) Locator-Extractor 0.523 0.492 0.502 [0.481, 0.528]
Locator-Extractor (Gemini 3.1 Pro + GPT-5.4 Mini) Locator-Extractor 0.560 0.436 0.478 [0.456, 0.502]
Locator-Extractor (Gemini 3.1 Pro) Locator-Extractor 0.500 0.422 0.448 [0.423, 0.471]
Claude Sonnet 4.5* (reference) Single-pass 0.903 0.840 0.861 [0.830, 0.893]

Reproducing the paper

Installation

The paper's environment uses Python 3.10 or 3.11 and the pinned versions in requirements.txt:

git clone https://github.com/berkearda/croissantminer
cd croissantminer
pip install -r requirements.txt
pip install -e .

API keys are only needed to run the systems or the judge. Copy .env.example to .env and fill in the keys for the providers you use.

Checking the numbers

No API keys and no cost: the scores are recomputed from the stored system outputs (data/extractions/), judge verdicts (data/judged/) and gold annotations (data/annotations/gold.parquet), which are included in this repository.

  1. Check Table 2 and the per-field tables in the appendix against the published numbers:
    make reproduce
    
  2. Print Table 2 (Core, RAI and Composite with 95% confidence intervals, grouped as in the paper):
    make table2
    
  3. Run the pairwise significance tests (paired bootstrap, Wilcoxon and McNemar with BH-FDR correction):
    make significance
    

Running the systems and the scoring rules

docs/reproducing.md explains how to re-run each system on the benchmark (this needs API keys and the benchmark PDFs) and gives the exact scoring rules: rule-based scores for the 10 core fields, the GLM-5 judge for the 20 RAI fields, and how empty values are scored.

Tests

make test

About 120 tests, about 10 seconds, no API keys. They cover the scoring rules, the handling of model output, the command line and the Croissant output, and the check that Table 2 and Tables 5 and 6 are reproduced exactly. GitHub Actions runs them on every push, and the tool's tests on Python 3.10 to 3.13.

Extending CroissantMiner

  • A new method or model for the tool: methods are registered in croissantminer/methods.py (METHODS and run), model backbones in croissantminer/systems/helpers.py (MODELS), and the Croissant file is built in croissantminer/croissant.py.
  • A new system on the benchmark: make evaluate OUTPUTS=folder NAME=name scores it with the paper's scorer and judge model; the leaderboard explains the format and how to add your entry.
  • A wrong extraction, a bug or an idea: open an issue or a discussion.

Repository layout

Path Contents
croissantminer/ Package: command line and Python API (cli.py, api.py), the six methods (methods.py) and the systems' code (systems/), the Croissant file (croissant.py), the extraction prompt, PDF reading and the ReAct agent
scripts/ Benchmark runs of the systems, the judge, tables and figures (guide in scripts/README.md)
evaluation/ Field metrics and system registry used by the scorer
data/ Gold annotations, judge verdicts and system outputs
silver/ Selection and extraction of the 500 silver papers
tests/ Tests and the published numbers they check against
hf_space/ The Hugging Face Space demo
docs/ Reproduction guide (reproducing.md), README figures and the ReAct agent's prompts
legacy/ Early prototype and experiments from before the paper, kept for reference and not maintained

The scripts that read the named annotation sheets are not included, to protect the annotators' privacy; data/annotations/gold.parquet and data/annotations/iaa.parquet are their output. Comments that cite decisions.md or task numbers (T-###) refer to our internal project log, which is not included.

Community

Questions and ideas go to Discussions, bugs and wrong extractions to issues. Please read CONTRIBUTING.md before opening a pull request. Everyone taking part follows the code of conduct; security problems are reported as described in SECURITY.md. Changes are listed in CHANGELOG.md.

Citation

@inproceedings{arda2026croissantminer,
  title     = {CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets},
  author    = {Arda, Berke and Yavuz, Ahmetcan and Gerry, Paul and Lobentanzer, Sebastian and
               Sarwar, Nobin and Giner-Miguelez, Joan and Chen, Kongtao and Zhang, Luyao and
               Sachan, Mrinmaya and Akhtar, Mubashara},
  booktitle = {Advances in Neural Information Processing Systems (Evaluations and Datasets Track)},
  year      = {2026}
}

License

Code: MIT (see LICENSE). Annotations: CC BY 4.0 (see the dataset card). The papers remain under their authors' licenses.

Metadata

Release files for croissantminer 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for croissantminer 0.2.0
File Size Uploaded
croissantminer-0.2.0.tar.gz 168.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for croissantminer 0.2.0
File Interpreter ABI Platform
croissantminer-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 342.9 kB

Release files / croissantminer-0.2.0.tar.gz

Download URL croissantminer-0.2.0.tar.gz
Size 168.2 kB
Tags Source
SHA-256 checksum
How to use checksums
e204c9946301aefc39986a969f15fa2eaac846409840f9d81da42404a3ca353b
BLAKE2b-256 checksum
How to use checksums
e8a1267c99386b2fc07a0be79c8fb417ac63e813c21a96dba61fab38e84ac0e1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release files / croissantminer-0.2.0-py3-none-any.whl

Download URL croissantminer-0.2.0-py3-none-any.whl
Size 174.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
49792c216aac20ebbadaca7c0847550782c70ed26e0a9303dc83258acb63a52e
BLAKE2b-256 checksum
How to use checksums
f0371038caeee1a2d10026da6d1952fa07c7f51cb3f80815f5cc740d139ad9d6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 1, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page