CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cometadata /crossref-unique-affiliations Crossref unique affiliations Exact organization strings extracted from Crossref snapshot 2026-07. For the nine exact-string splits, no trimming, case folding, or Unicode normalization is performed. Repeated leaf occurrences are counted, and empty decoded strings alone are excluded. Normalized split The normalized split groups the exact strings in all after applying this normalization contract: Transliterate Unicode text to ASCII with Unidecode. Convert letters… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-unique-affiliations.text100M<n<1B0 likes756 downloads26d agoHugging Face02RichardErkhov /April_2023_Public_Data_File_from_Crossref0 likes445 downloads2y agoHugging Face03bluuebunny /crossref_metadata_2025_split Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025_split.textsentence-similarity10M<n<100M0 likes161 downloads1y agoHugging Face04cometadata /crossref-datacite-citations Crossref DataCite Citations A dataset of DataCite-registered works and the Crossref-registered works that cite them, extracted from Crossref reference metadata and confirmed against the DataCite monthly data file. Dataset Description Each record in the citation configurations is one DataCite DOI together with every confirmed citing work found in Crossref reference metadata. A reference is confirmed when it carries a DOI registered in DataCite or an arXiv… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-datacite-citations.feature-extraction1M<n<10M0 likes110 downloads4d agoHugging Face05cometadata /crossref-arxiv-citations Crossref arXiv Citations A dataset of arXiv preprints and their citations extracted from Crossref metadata, validated against DataCite records. Dataset Description This dataset maps arXiv works to the works in Crossref that cite them. Each record represents an arXiv preprint with all known citations from Crossref-registered works. Built from the Crossref Metadata Plus monthly snapshot 2026-07 and the DataCite monthly data file 2026-07 with… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-arxiv-citations.tabulartext-classification1M<n<10M0 likes102 downloads4d agoHugging Face06bluuebunny /crossref_metadata_embeddings_split_2025Created vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 textsentence-similarity10M<n<100M0 likes71 downloads1y agoHugging Face07RobPT /tractor-crossref-assetsdocumentn<1K0 likes58 downloads3mo agoHugging Face08bluuebunny /crossref_metadata_2025 Dataset Overview This dataset contains bibliographic metadata from the public Crossref snapshot released in 2025. It provides core fields for scholarly documents, including DOI, title, abstract, authorship, publication month and year, and URLs. The entire public dump (~196.94 GB) was filtered and extracted into a parquet format for efficient loading and querying. Total size: 196.94 GB (parquet files) Number of records: 34,308,730 Use this dataset for large-scale text mining… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_2025.textsentence-similarity10M<n<100M0 likes51 downloads1y agoHugging Face09bluuebunny /crossref_metadata_embeddings_split_2025_binaryCreated vector embeddings for the abstract field for the dataset: bluuebunny/crossref_metadata_2025_split using mixedbread-ai/mxbai-embed-large-v1 and binarised it using: # Function to binarise float embeddings def binarise(row): # Make it a numpy array, since batching sends it as list float_vector = np.array(row['vector'], dtype=np.float32) # Binarise binary_vector = np.where(float_vector >= 0, 1, 0) # Pack it to make it milvus compatible row['vector'] =… See the full description on the dataset page: https://huggingface.co/datasets/bluuebunny/crossref_metadata_embeddings_split_2025_binary.textsentence-similarity10M<n<100M0 likes44 downloads1y agoHugging Face10referencesource /psm-rmp-chemical-threshold-crossref OSHA PSM vs EPA RMP: chemical threshold quantities side by side Canonical, always-current version: https://referencesource.org/psm-rmp-chemical-threshold-crossref/ Machine-readable: https://referencesource.org/psm-rmp-chemical-threshold-crossref/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-15 Stale after: 2027-08-15 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 225 Side-by-side comparison of… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/psm-rmp-chemical-threshold-crossref.textn<1K0 likes43 downloads26d agoHugging Face11reapxdev /crossref-scraper Crossref Scraper · DOI Metadata, Authors, Journals & Citations Scrape scholarly DOI metadata, works, journal articles, authors, citations, funding, and licenses from the Crossref REST API. Fast HTTP scraper with pay-per-event pricing. Rows in this dataset 2,492 Fields 29 Collector runs behind it 50 Most recent observation 2026-08-04 Browsable presentation https://reapx.dev/data/crossref-scraper/ — 2,492 entity pages Run the collector yourself… See the full description on the dataset page: https://huggingface.co/datasets/reapxdev/crossref-scraper.tabular1K<n<10K0 likes29 downloads2mo agoHugging Face12cometadata /crossref-preprint-article-relationships-matched-work-has-funding Overview This dataset contains preprint-article relationships from Crossref, filtered to entries where the matched work has a funding (grant) entry in OpenAlex. Data Structure preprint_doi (string): Preprint DOI from various repositories (e.g., 10.21203/rs.3.rs-, 10.1101/, 10.31234/osf.io/*) article_doi (string): Published article DOI deposited_by_article_publisher (boolean): Whether article publisher deposited the relationship deposited_by_preprint_publisher (boolean):… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/crossref-preprint-article-relationships-matched-work-has-funding.text100K<n<1M0 likes22 downloads1y agoHugging Face13wheelofheaven /theomachy-crossrefs Wheel of Heaven Theomachy Cross-References The combat myth (chaoskampf) across eight traditions — champion, adversary, chaos form, weapon, outcome, and reference — with links to the digitized texts. Records: 8 traditions Columns / structure: tradition, source_text, reference, champion, adversary, chaos_form, weapon, outcome, woh_library Formats: CSV, JSON License: CC0-1.0 (public domain) Version: 2026.07 Provenance Extracted from the live Wheel of Heaven corpus… See the full description on the dataset page: https://huggingface.co/datasets/wheelofheaven/theomachy-crossrefs.textn<1K1 likes17 downloads2mo agoHugging Face14davanstrien /crossreftest0 likes12 downloads3y agoHugging Face15cometadata /crossref-affiliations-ror-benchmarktext1K<n<10K0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.