Skip to content
AI.info

Research

CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets

Overview Research area: Machine-readable metadata standards for machine learning datasets (Croissant), automated information extraction from scientific papers, and benchmark construction for large lan

CroissantMiner: Automated Extraction and Validation of Croissant Metadata for ML Datasets
arXiv
2610.07132
Published
2026-10-05
Authors
Berke Arda, Ahmetcan Yavuz, Paul Gerry, Sebastian Lobentanzer, Nobin Sarwar, Joan Giner-Miguelez, Kongtao Chen, Luyao Zhang, Mrinmaya Sachan, Mubashara Akhtar

AI summary

Overview

  • Research area: Machine-readable metadata standards for machine learning datasets (Croissant), automated information extraction from scientific papers, and benchmark construction for large language model evaluation.
  • Technical level: Intermediate.
  • Scope: This paper introduces CroissantMiner, the first benchmark for end-to-end extraction of the full 30-field Croissant 1.1 metadata schema from ML dataset papers, and evaluates 24 extraction systems on it.

What This Paper Is About

Croissant is now a community standard for machine-readable dataset metadata — Hugging Face, Kaggle, and OpenML support it, and NeurIPS 2026 mandates Croissant Responsible AI (RAI) metadata files for dataset submissions — but filling in its fields is still manual, labor-intensive work that requires carefully reading dataset papers. Existing automated approaches cover only a small subset of fields and struggle with the variability of scientific papers.

The goal of this work is to build a benchmark and an open-source system that can extract the complete 30-field Croissant schema, including all 20 Responsible AI fields, directly from dataset papers, and to measure how well current models actually do at that task.

Key Contributions

  1. A benchmark for Croissant metadata extraction. 602 ML dataset papers — 102 with human-validated gold annotations and 500 with LLM-generated silver annotations — annotated against the full 30-field Croissant 1.1 schema (10 core and 20 RAI fields). Gold annotation involved 22 annotators from 18 affiliations and produced 9,595 ratings over 3,060 cells.
  2. A systematic evaluation. 24 extraction systems are compared, spanning seven proprietary frontier models, five open-weight models, and four agentic architectures, under a two-tier scoring framework that combines rule-based metrics with an LLM judge selected through human audit.
  3. Findings on extraction difficulty. Single-pass extraction consistently outperforms the four agentic architectures tested; long-form RAI fields requiring synthesis across a paper are the hardest to extract; the paper also reports annotation reliability and field-level difficulty analysis.
  4. Public resources. An open-source Python package, a Hugging Face Space demo, evaluation code, judge audit, and a leaderboard open to new systems.

Main Findings

  • Single-pass beats decomposition. With Claude Sonnet 4.6 fixed as backbone, all four agentic variants score below single-pass extraction (composite 0.709): ReAct 0.652 (Δ = −0.057), Parallel Specialists 0.647 (Δ = −0.061), Triage + Critique 0.624 (Δ = −0.085), and Locator-Extractor 0.566 (Δ = −0.143). All four differences are statistically significant under paired bootstrap and Wilcoxon signed-rank tests with BH-FDR correction (q < 0.001).
  • The pattern holds across backbones. Single-pass extraction again scored higher than every agentic architecture by 0.034 to 0.163 on GPT-5.4 and by 0.008 to 0.154 on Gemini 3.1 Pro.
  • The best system still leaves large headroom. Claude Sonnet 4.6 single-pass ranks first at 0.709 composite (95% CI [0.688, 0.729]), with core 0.752 and RAI 0.687. No configuration exceeds 0.71 composite.
  • Open-weight models are competitive. Qwen 3.6 35B-A3B scored 0.634 and GLM-5.1 scored 0.625, competing with several proprietary models. The full spread across evaluated systems ran from 0.391 (Llama 4 Scout 17B) to 0.709.
  • RAI fields are the hard part and dominate the score. RAI fields account for two-thirds of the composite score. Agentic variants show larger performance drops on RAI than on core fields. Claude Opus 4.7 is an exception, scoring slightly better on RAI (0.711) than core (0.676).
  • Agentic pipelines miss documented evidence. On test-split RAI cells where the gold value is documented, single-pass extraction returned empty only 2.4% of the time, while the agentic variants missed between 10.9% and 14.8% (Parallel Specialists 10.9%, Locator-Extractor 12.9%, ReAct 13.4%, Triage + Critique 14.8%) — a four- to six-fold increase.
  • Errors are misclassification more than fabrication. Among the 16.6% of pre-fill ratings annotators flagged, incomplete extraction accounted for 43.0%, wrong-section assignment 19.5%, and hallucination only 9.7%.
  • Annotation agreement tracks field openness. 61.5% of cells were unanimous. Constrained fields reached near-perfect agreement (cr:isLiveDataset, AC₁ = 0.96), while open-ended RAI fields were lowest (rai:dataAnnotationAnalysis, AC₁ = 0.45; rai:dataPreprocessingProtocol, AC₁ = 0.46). sc:publisher had AC₁ = −0.09 because annotators disagreed between dataset host, venue, and author affiliation — a schema ambiguity rather than annotator error.
  • Rarely documented fields are not the source of disagreement. rai:dataImputationProtocol, which is rarely documented, reached AC₁ = 0.93 because annotators agreed the field was absent.
  • Documentation coverage is uneven. Across all 602 papers an average of 19.2 of 30 fields (64.0%) are documented — 21.7 in gold, 18.7 in silver. Question-answering benchmarks document the most (21.0) and robotics the fewest (17.7). Core fields like name, description, and creator appear in over 99% of papers, while rai:dataImputationProtocol, rai:dataCollectionMissingData, and rai:dataCollectionTimeframe are all below 25%; rai:dataUseCases and rai:dataLimitations exceed 90%.
  • Leaving fields empty is a substantial part of the task. In the gold split, 27.7% of all cells have no documented value; the rate is 35.6% for RAI fields versus 11.9% for core fields.
  • Agentic approaches cost more without helping. Under the paper's token-accounting assumptions, agentic configurations cost 1.2× to 7.8× as much per paper as single-pass extraction at matched backbones.
  • Judge choice was audited. GLM-5 agreed best with human consensus (pooled κ = 0.890, stable at 0.900 and 0.880 across calibration and validation halves), while Gemini 2.5 Pro fell from 0.900 to 0.743 on validation. With the final rubric and audited gold, GLM-5 reaches κ = 0.708 on the same 60 cells; on 200 harder RAI cells it matched human consensus in 71.5% of cells (κ = 0.663) versus 72.0% agreement between the two human raters (κ = 0.739), and was more lenient than humans in 45 of 57 disagreement cases.
  • Paper length does not predict difficulty; sparsity does. For single-pass extraction, scores fall steadily from the best-documented to the most sparsely documented papers, while page count has no measurable effect.

Methodology in Plain English

The researchers built the benchmark from publicly available AI dataset papers, prioritizing highly downloaded Hugging Face datasets, spanning vision, NLP, multimodal learning, audio, code, robotics, and medical AI. The gold split was manually curated; the silver split was selected automatically from the Hugging Face Hub (500 papers linking to an arXiv paper, ranked by download count, deduplicated by arXiv identifier, and filtered to remove overlap with the gold split).

Annotation followed a pre-fill paradigm: an LLM (Claude Sonnet 4.5, at temperature 0 with a canonical prompt) first generated a complete 30-field metadata record for each paper, and human annotators verified and corrected those values rather than extracting from scratch. Sonnet 4.5 was chosen after pilot extractions on a 10-paper subset against GPT-4o-mini and Gemini 2.5 Pro. Each paper-field pair received at least three independent ratings on a three-level rubric (Correct / Partially Correct / Not Correct), plus failure-mode labels (incomplete, hallucination, wrong section, granularity mismatch, format error, other), free-text notes, and confidence scores. Majority voting resolved 95.0% of cells; a senior author adjudicated 140 of the remainder and the lead author resolved the other 14. Gold values record only what the paper itself states; information that exists only on a dataset card is marked [NULL - not found in paper].

For evaluation, the gold split was divided into a 14-paper development set and an 88-paper held-out test set. Scoring has two tiers: rule-based exact-match for constrained fields and token-F1 for short text, and a 1–3 ordinal LLM judge (GLM-5) for the twenty long-form RAI fields, with null predictions handled explicitly (correct-null skipped, false positives penalized). Composite scores average over all 30 fields and come with 2,000-replicate paper-clustered bootstrap confidence intervals; comparisons use paired bootstrap, Wilcoxon signed-rank, and McNemar's tests with BH-FDR correction.

The extraction pipeline itself has three stages: text extraction from PDF using PyPDF2 with reference sections removed by regex (appendices retained), LLM extraction via a canonical prompt or an architecture-specific procedure, and validation against the 30-field schema, where schema violations occurred in fewer than 1% of extractions. The four agentic architectures were designed to cover distinct decomposition patterns: Parallel Specialists (five specialist calls over field groups, then merging), Triage + Critique (a planning call, a full-paper extraction with supporting quotes, a self-critique pass, and a verification step), Locator-Extractor (section location, then per-group extraction over selected passages, then external enrichment and confidence scoring), and ReAct (a dynamic reasoning-and-tool-use loop with tools for reading the paper, retrieving paragraphs, looking up Hugging Face metadata, verifying URLs, and normalizing licenses). Three of the agentic architectures can also fill some fields from Hugging Face or Semantic Scholar metadata; single-pass and Parallel Specialists use only the paper.

Why This Matters

The work matters because manual metadata creation cannot scale to a landscape with over 500,000 public datasets on Hugging Face and more than 32,000 datasets created monthly in 2025, while the Croissant standard is being mandated by major venues. NeurIPS 2025 Datasets and Benchmarks track post-acceptance surveys report that 16% of authors encountered submission difficulties, primarily due to dataset hosting constraints and Croissant metadata generation challenges. The paper shows that this bottleneck is real and that current systems are far from solving it — and that adding agentic scaffolding makes things worse, not better, while costing more.

Real-world applications:

  • Dataset authors can use the released open-source system and Hugging Face Space to generate Croissant metadata drafts, then inspect and refine them before submission.
  • Venue and track reviewers can use generated drafts to check the completeness of submitted RAI metadata against what a paper actually documents.
  • Dataset platforms and repositories can integrate schema-complete extraction to populate contextual and RAI fields that cannot be inferred from the data files themselves.
  • Dataset users and consumers can rely on standardized metadata papers for discoverability, licensing, provenance, and responsible-use information.

Industry relevance: the benchmark's results cut against the common assumption that decomposing a task into multi-step agentic pipelines improves quality. Here decomposition degraded performance substantially on schema-complete extraction and cost 1.2× to 7.8× more per paper, a directly relevant finding for teams deciding how to architect document-understanding systems for metadata, documentation, or compliance workflows. The finding that open-weight models such as Qwen 3.6 35B-A3B and GLM-5.1 are competitive with several proprietary models also matters for organizations weighing self-hosted deployments. Finally, the released leaderboard, judge audit, and evaluation code provide shared infrastructure for measuring progress on this task.

Future Directions

  • Improvement must go beyond stronger models. The authors argue that better handling of structure and uncertainty is needed — encouraging systems to leave a field empty when evidence is unclear, validating outputs against field definitions, and refining ambiguous schema fields.
  • Schema-level ambiguities need resolution. The definition of fields such as publisher must be clarified for consistent evaluation and further progress; the low inter-annotator agreement on long-form RAI fields points to the same need for more precise field definitions in future schema versions.
  • Extending beyond the current scope. The benchmark covers English-language ML dataset papers and the 30-field Croissant 1.1 schema; extension to other languages, domains, and future schema versions is left to future work.
  • Closing the evaluation gap. Because Tier-2 scores rely on an LLM judge with imperfect agreement, especially on difficult fields, judge-based scores do not replace a full human evaluation. Several RAI fields also have very few documented cells in the test split — as few as two for rai:dataImputationProtocol — so more evidence is needed there. Pretraining contamination cannot be ruled out, and re-annotating 10 papers from a different model's pre-fills left the ranking stable but favoured the seed model and, slightly, its family, so residual anchoring to the pre-fill model cannot be excluded.

Target Audience

Researchers and engineers working on dataset documentation, metadata standards, and document-level information extraction from scientific papers will benefit most. The paper is also directly useful to dataset authors and venue organizers responsible for Croissant and RAI metadata compliance, to practitioners deciding between single-pass and agentic LLM architectures for structured extraction tasks, and to benchmark designers interested in the paper's two-tier scoring setup, human-audited judge selection, and pre-fill annotation protocol. Readers should be comfortable with standard LLM evaluation terminology such as inter-annotator agreement, bootstrap confidence intervals, and weighted kappa, though the paper explains each in context.

Authors’ abstract

Croissant has emerged as a standard for machine-readable dataset metadata, yet populating its fields remains labor-intensive and requires careful reading of accompanying dataset documentation. We present the first benchmark enabling end-to-end evaluation of metadata extraction aligned with a community-standard schema. The benchmark comprises 602 papers, including 102 with human-validated gold annotations and 500 with LLM-generated silver annotations, covering the full Croissant schema with both core and Responsible AI (RAI) fields. Using this benchmark, we evaluate a range of extraction systems spanning frontier models, open-weight models, and agentic architectures, under a two-tier evaluation framework that combines rule-based scoring with an LLM judge selected via human audit. We find that single-pass extraction consistently outperforms the four agentic architectures we evaluate: across backbones, these decomposed variants achieve lower accuracy than a single full-context pass. The largest gap appears on long-form RAI fields, which require synthesizing and interpreting information scattered across a paper rather than copying it from a single location, a setting where current systems remain far from reliable. We release the benchmark, evaluation code, judge audit, a live demo, and a leaderboard open to new systems.

Read the original paper