datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-software-repo-links-datacite-enrichment-format
arXiv Software Repository Links - DataCite Enrichment Format
A collection of metadata enrichments, formatted for DataCite's enrichment API, that add links between arXiv papers (via DOI) and the software repositories they reference or are supplemented by.
Quick Start
from datasets import load_dataset
ds = load_dataset("cometadata/arxiv-software-repo-links-datacite-enrichment-format")
Dataset Description
Each record is a DataCite-style enrichment instruction… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-software-repo-links-datacite-enrichment-format.datacite-affiliations-matched-ror-ids-datacite-enrichment-format
DataCite Author Affiliations Matched to ROR IDs - DataCite Enrichment Format
This dataset contains 20,512,320 enrichment records mapping author affiliation strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 5,812,774 unique DOIs from the DataCite Public Data File.
Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the creators field… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-affiliations-matched-ror-ids-datacite-enrichment-format.comet-datacite-enrichment-layer-ingest-format
COMET Enrichment Files - DataCite Enrichment Layer Ingest Format
This dataset contains metadata enrichment records in the DataCite enrichment ingest format, produced by the COMET (Collaborative Metadata) initiative. Each record represents a single metadata enrichment to be applied to an existing DataCite record, identified by DOI.
The enrichments were generated against the February 2026 version of the DataCite monthly data file using the datacite-enrichment tool.
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/comet-datacite-enrichment-layer-ingest-format.CISA_Enrichment
CISA Known Exploited Vulnerabilities Catalog Enrichment
The CISA recently started to publish the Known Exploited Vulnerabilities Catalog Enrichment to help federal agencies keep up with exploited vulnerabilities.
The data they provide is minimal, so I have built this jupyter notebook to enrich the data using the CIRCL public CVE API to add the following data points:
CWE
CVE Published Date
CVE Modified Date
Reference URLs
CPE 2.3 Data
A Github Action runs every 6 hours and updates… See the full description on the dataset page: https://huggingface.co/datasets/cvelist/CISA_Enrichment.datacite-funders-matched-ror-ids-datacite-enrichment-format
DataCite Funders Matched to ROR IDs - DataCite Enrichment Format
This dataset contains 1,008,697 enrichment records mapping funder name strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 697,791 unique DOIs from the DataCite Public Data File.
Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the fundingReferences field, providing a… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-funders-matched-ror-ids-datacite-enrichment-format.arxiv-author-affiliation-parsing-sample-datacite-enrichment-formatdatacite-procedural-resource-type-general-reclassifications-datacite-enrichment-format
