datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crossref-unique-affiliations
Crossref unique affiliations
Exact organization strings extracted from Crossref snapshot 2026-07.
For the nine exact-string splits, no trimming, case folding, or Unicode normalization is performed. Repeated leaf
occurrences are counted, and empty decoded strings alone are excluded.
Normalized split
The normalized split groups the exact strings in all after applying this
normalization contract:
Transliterate Unicode text to ASCII with Unidecode.
Convert letters… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-unique-affiliations.crossref_metadata_2025
Dataset Overview
This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying.
Total size: 196.94 GB (parquet files)
Number of records: 34,308,730
Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025.crossref_metadata_2025_split
Dataset Overview
This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying.
Total size: 196.94 GB (parquet files)
Number of records: 34,308,730
Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.crossref-arxiv-citations
Crossref arXiv Citations
A dataset of arXiv preprints and their citations extracted from Crossref
metadata, validated against DataCite records.
Dataset Description
This dataset maps arXiv works to the works in Crossref that cite them. Each
record represents an arXiv preprint with all known citations from
Crossref-registered works.
Built from the Crossref Metadata Plus monthly snapshot 2026-07 and the
DataCite monthly data file 2026-07 with… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-arxiv-citations.crossref_metadata_embeddings_split_2025Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1
psm-rmp-chemical-threshold-crossref
OSHA PSM vs EPA RMP: chemical threshold quantities side by side
Canonical, always-current version: https://referencesource.org/psm-rmp-chemical-threshold-crossref/
Machine-readable: https://referencesource.org/psm-rmp-chemical-threshold-crossref/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-15
Stale after: 2027-08-15 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 225
Side-by-side comparison of… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/psm-rmp-chemical-threshold-crossref.crossref_metadata_embeddings_split_2025_binaryCreated vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using:
# Function to binarise float embeddings
def binarise(row):
# Make it a numpy array, since batching sends it as list
float_vector = np.array(row['vector'], dtype=np.float32)
# Binarise
binary_vector = np.where(float_vector >= 0, 1, 0)
# Pack it to make it milvus compatible
row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.crossref-scraper
Crossref Scraper · DOI Metadata, Authors, Journals & Citations
Scrape scholarly DOI metadata, works, journal articles, authors, citations, funding, and licenses from the Crossref REST API. Fast HTTP scraper with pay-per-event pricing.
Rows in this dataset
2,492
Fields
29
Collector runs behind it
50
Most recent observation
2026-08-04
Browsable presentation
https://reapx.dev/data/crossref-scraper/ — 2,492 entity pages
Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/crossref-scraper.crossref-preprint-article-relationships-matched-work-has-funding
Overview
This dataset contains preprint-article relationships from Crossref, filtered to entries where the matched work has a funding (grant) entry in OpenAlex.
Data Structure
preprint_doi (string): Preprint DOI from various repositories (e.g., 10.21203/rs.3.rs-, 10.1101/, 10.31234/osf.io/*)
article_doi (string): Published article DOI
deposited_by_article_publisher (boolean): Whether article publisher deposited the relationship
deposited_by_preprint_publisher (boolean):… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-preprint-article-relationships-matched-work-has-funding.theomachy-crossrefs
Wheel of Heaven Theomachy Cross-References
The combat myth (chaoskampf) across eight traditions — champion, adversary, chaos form, weapon, outcome, and reference — with links to the digitized texts.
Records: 8 traditions
Columns / structure: tradition, source_text, reference, champion, adversary, chaos_form, weapon, outcome, woh_library
Formats: CSV, JSON
License: CC0-1.0 (public domain)
Version: 2026.07
Provenance
Extracted from the live Wheel of Heaven corpus… See the full description on the dataset page: https://huggingface.co/datasets/wheelofheaven/theomachy-crossrefs.crossref-affiliations-ror-benchmark
