datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-author-affiliations-matched-ror-ids
arXiv Author Affiliations
This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers.
Dataset Description
This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.my-voxtral-datasetror-matching-train-validation-test
ROR Affiliation Matching (DataCite + Crossref + AffRoDB + Synthetic OpenAlex)
Raw author-affiliation strings paired with the ROR (Research Organization
Registry) identifiers they should resolve to, prepared for training and
evaluating affiliation matching and entity-linking systems.
The dataset ships six subsets (loadable as Hugging Face configs), each
split into train/validation/test:
Subset
Records
Source
Empty-label rows
crossref (default)
3,000
Crossref-derived… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/ror-matching-train-validation-test.datacite-affiliations-matched-ror-ids-datacite-enrichment-format
DataCite Author Affiliations Matched to ROR IDs - DataCite Enrichment Format
This dataset contains 20,512,320 enrichment records mapping author affiliation strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 5,812,774 unique DOIs from the DataCite Public Data File.
Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the creators field… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-affiliations-matched-ror-ids-datacite-enrichment-format.2025-08-datacite-ror-identifier-dois
ROR Identifier DOI Distribution
Summary
ror_identifier_doi_distribution.json provides an inventory of every ROR (or ROR-formatted) identifier in the August 2025 DataCite data file. Each entry describes how many affiliation rows inlcude the identifier, the list
of unique DOIs, and which providers/clients asserted it.
Structure
{
"identifier": "https://ror.org/01abc1234",
"occurrences": 512,
"dois": ["10.1234/abc", "10.1234/xyz"],
"providers":… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/2025-08-datacite-ror-identifier-dois.datacite-funders-matched-ror-ids-datacite-enrichment-format
DataCite Funders Matched to ROR IDs - DataCite Enrichment Format
This dataset contains 1,008,697 enrichment records mapping funder name strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 697,791 unique DOIs from the DataCite Public Data File.
Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the fundingReferences field, providing a… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-funders-matched-ror-ids-datacite-enrichment-format.RoRuDiRoRuDi - Romanian Rules for Dialects
1-roleplayfinance_data_v4.1attack_on_titan_wikifinance_data_v2act-configsRoRuDi
Romanian Rules for Dialects Dataset (RoRuDi)
You can also find the dataset here: https://huggingface.co/datasets/fmi-unibuc/RoRuDi
dsaa6000q_q3finance_data_v1attack_on_titan_wiki_chinesefinance_data_v3finance_data_v4
