CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cometadata /crossref-unique-affiliations Crossref unique affiliations Exact organization strings extracted from Crossref snapshot 2026-07. For the nine exact-string splits, no trimming, case folding, or Unicode normalization is performed. Repeated leaf occurrences are counted, and empty decoded strings alone are excluded. Normalized split The normalized split groups the exact strings in all after applying this normalization contract: Transliterate Unicode text to ASCII with Unidecode. Convert letters… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-unique-affiliations.text100M<n<1B0 likes857 downloads27d agoHugging Face02bluuebunny /crossref_metadata_2025 Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025.textsentence-similarity10M<n<100M0 likes586 downloads1y agoHugging Face03bluuebunny /crossref_metadata_2025_split Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.textsentence-similarity10M<n<100M0 likes569 downloads1y agoHugging Face04cometadata /crossref-arxiv-citations Crossref arXiv Citations A dataset of arXiv preprints and their citations extracted from Crossref metadata, validated against DataCite records. Dataset Description This dataset maps arXiv works to the works in Crossref that cite them. Each record represents an arXiv preprint with all known citations from Crossref-registered works. Built from the Crossref Metadata Plus monthly snapshot 2026-07 and the DataCite monthly data file 2026-07 with… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-arxiv-citations.tabulartext-classification1M<n<10M0 likes105 downloads5d agoHugging Face05bluuebunny /crossref_metadata_embeddings_split_2025Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 textsentence-similarity10M<n<100M0 likes72 downloads1y agoHugging Face06referencesource /psm-rmp-chemical-threshold-crossref OSHA PSM vs EPA RMP: chemical threshold quantities side by side Canonical, always-current version: https://referencesource.org/psm-rmp-chemical-threshold-crossref/ Machine-readable: https://referencesource.org/psm-rmp-chemical-threshold-crossref/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-15 Stale after: 2027-08-15 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 225 Side-by-side comparison of… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/psm-rmp-chemical-threshold-crossref.textn<1K0 likes45 downloads27d agoHugging Face07bluuebunny /crossref_metadata_embeddings_split_2025_binaryCreated vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.textsentence-similarity10M<n<100M0 likes44 downloads1y agoHugging Face08reapxdev /crossref-scraper Crossref Scraper · DOI Metadata, Authors, Journals & Citations Scrape scholarly DOI metadata, works, journal articles, authors, citations, funding, and licenses from the Crossref REST API. Fast HTTP scraper with pay-per-event pricing. Rows in this dataset 2,492 Fields 29 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/crossref-scraper/ — 2,492 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/crossref-scraper.tabular1K<n<10K0 likes34 downloads2mo agoHugging Face09cometadata /crossref-preprint-article-relationships-matched-work-has-funding Overview This dataset contains preprint-article relationships from Crossref, filtered to entries where the matched work has a funding (grant) entry in OpenAlex. Data Structure preprint_doi (string): Preprint DOI from various repositories (e.g., 10.21203/rs.3.rs-, 10.1101/, 10.31234/osf.io/*) article_doi (string): Published article DOI deposited_by_article_publisher (boolean): Whether article publisher deposited the relationship deposited_by_preprint_publisher (boolean):… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-preprint-article-relationships-matched-work-has-funding.text100K<n<1M0 likes24 downloads1y agoHugging Face10wheelofheaven /theomachy-crossrefs Wheel of Heaven Theomachy Cross-References The combat myth (chaoskampf) across eight traditions — champion, adversary, chaos form, weapon, outcome, and reference — with links to the digitized texts. Records: 8 traditions Columns / structure: tradition, source_text, reference, champion, adversary, chaos_form, weapon, outcome, woh_library Formats: CSV, JSON License: CC0-1.0 (public domain) Version: 2026.07 Provenance Extracted from the live Wheel of Heaven corpus… See the full description on the dataset page: https://huggingface.co/datasets/wheelofheaven/theomachy-crossrefs.textn<1K1 likes17 downloads2mo agoHugging Face11cometadata /crossref-affiliations-ror-benchmarktext1K<n<10K0 likes5 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.