CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01cometadata /arxiv-author-affiliations-matched-ror-ids arXiv Author Affiliations This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers. Dataset Description This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.texttext-classification1M<n<10M1 likes426 downloads8mo agoHugging Face02Wikidepia /id_scholar_huridocstext1M<n<10M0 likes219 downloads2y agoHugging Face03rl-research /researchqa_official_subset_idstextn<1K0 likes150 downloads10mo agoHugging Face04orionweller /cc-ids-to-title Common Crawl IDs to Titles JSON manifests mapping document IDs to their extracted titles. Usage from datasets import load_dataset titles = load_dataset("orionweller/cc-ids-to-title", "fw-edu") text100M<n<1B0 likes133 downloads1y agoHugging Face05netgoat-ai /Koda-IDS-CyberReasoning Koda IDS CyberReasoning Dataset Details Dataset Description Koda IDS CyberReasoning is a multi-domain reasoning dataset designed to support the development and evaluation of AI systems for cybersecurity incident detection, investigation, analysis, and response. The dataset contains conversational reasoning examples covering cybersecurity triage, attack-pattern analysis, incident recovery, false-positive suppression, security-related coding, tool… See the full description on the dataset page: https://huggingface.co/datasets/netgoat-ai/Koda-IDS-CyberReasoning.text10K<n<100K0 likes77 downloads5d agoHugging Face06ONESTRUCTION /Ishigaki-IDS-Bench Ishigaki-IDS-Bench Ishigaki-IDS-Bench is a bilingual benchmark for generating buildingSMART IDS 1.0 XML from user-facing building requirements. Data Format Each row uses the common conversational messages format: { "id": "row-0001", "messages": [ {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."} ], "language": "ja", "input_format": "natural_language", "turn_type": "single_turn", "domain": "architectural","ifc_versions":… See the full description on the dataset page: https://huggingface.co/datasets/ONESTRUCTION/Ishigaki-IDS-Bench.texttext-generationn<1K2 likes65 downloads4mo agoHugging Face07cometadata /datacite-affiliations-matched-ror-ids-datacite-enrichment-format DataCite Author Affiliations Matched to ROR IDs - DataCite Enrichment Format This dataset contains 20,512,320 enrichment records mapping author affiliation strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 5,812,774 unique DOIs from the DataCite Public Data File. Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the creators field… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-affiliations-matched-ror-ids-datacite-enrichment-format.text10M<n<100M0 likes34 downloads5mo agoHugging Face08ucrelnlp /wikipedia-ga-fa-ids Wikipedia Good and Featured Articles Contains the Wikipedia Article IDs and page titles for Good and Featured articles on Wikipedia for a given timestamped data dump, whereby the data was extracted from Wikipedia/Wikimedia SQL table dumps. This dataset covers 9 Language Wikipedia sites. For more information on how the dataset was generated see https://github.com/UCREL/wikipedia-ga-fa-extraction. The data is specific to a given data dump timestamp, the main tag of the repository… See the full description on the dataset page: https://huggingface.co/datasets/ucrelnlp/wikipedia-ga-fa-ids.text10K<n<100K0 likes27 downloads2mo agoHugging Face09e17do /id-slang-synthetic-nlp Indonesian Synthetic Informal Language NLP Dataset A 20,000-record synthetic dataset of Indonesian informal language and slang designed for NLP research, language classification, sentiment and emotion analysis, conversational AI, and LLM-related research. The dataset contains simulated Indonesian conversational examples with structured metadata describing informal expressions, emotional characteristics, and dominant regional context. Dataset Summary Property… See the full description on the dataset page: https://huggingface.co/datasets/e17do/id-slang-synthetic-nlp.texttext-classification10K<n<100K0 likes18 downloads1mo agoHugging Face10cometadata /datacite-funders-matched-ror-ids-datacite-enrichment-format DataCite Funders Matched to ROR IDs - DataCite Enrichment Format This dataset contains 1,008,697 enrichment records mapping funder name strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 697,791 unique DOIs from the DataCite Public Data File. Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the fundingReferences field, providing a… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-funders-matched-ror-ids-datacite-enrichment-format.text1M<n<10M0 likes17 downloads5mo agoHugging Face11joduor /protein-diffusion-ids This dataset is a remastered version prepared using Adaption's Adaptive Data platform. protein_diffusion_ids This dataset contains identifier strings for various protein design and structure prediction methods, including RFdiffusion, PepMLM, JointDiff, and AlphaFold. The entries consist of specific method names often appended with unique alphanumeric hashes, alongside null values indicating missing data. It appears to catalog unpublished affinity maturation techniques and… See the full description on the dataset page: https://huggingface.co/datasets/joduor/protein-diffusion-ids.textn<1K0 likes11 downloads6mo agoHugging Face12maxidl /duplicate_ids_spa_Latntabularn<1K0 likes10 downloads7mo agoHugging Face13aios-k2chj /aios-k2chj-pci-usb-ids pci-usb-ids Part of AIOS K\u00b2CHJ. See GitHub. textn<1K0 likes9 downloads2mo agoHugging Face14ONESTRUCTION /IDS-Benchgated [!NOTE] This result was produced as part of the "GENIAC (Generative AI Accelerator Challenge) Project" which is promoted by the Ministry of Economy, Trade and Industry and the New Energy and Industrial Technology Development Organization (NEDO), with the aim of strengthening Japan’s capabilities in generative AI development. This dataset is an evaluation dataset for the CSV-to-IDS task used to evaluate the Ishigaki-IDS model. It evaluates whether an LLM can generate appropriate IDS from CSV… See the full description on the dataset page: https://huggingface.co/datasets/ONESTRUCTION/IDS-Bench.textn<1K4 likes7 downloads6mo agoHugging Face15AntBri /vatex-idstext10K<n<100K0 likes5 downloads4mo agoHugging Face16msrovani /aios-k2chj-pci-usb-ids 💾 SDIO HWIDs — The Largest Public Hardware ID Collection 171.003 unique PCI/USB/ACPI hardware identifiers extracted from 65 Windows DriverPacks. 🎯 What is this? This dataset contains 171.003 unique hardware identifiers (HWIDs) extracted from 65 SDIO DriverPacks — the largest collection of Windows hardware IDs assembled for AI training. Each entry is a raw HWID string like: PCI\VEN_10DE&DEV_1E81&SUBSYS_8597174B USB\VID_0A5C&PID_5848&REV_0102… See the full description on the dataset page: https://huggingface.co/datasets/msrovani/aios-k2chj-pci-usb-ids.textn<1K0 likes5 downloads2mo agoHugging Face17mujtabagulzarsoomro /Mujtaba-idsBlind Spots of Frontier Models Model Tested Model name: TinyLlama-1.1B-intermediate-step-1431k-3T Model link: https://huggingface.co/TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T This model is a base pretrained language model released on Hugging Face and was not specifically fine tuned for a particular downstream task. How the Model Was Loaded The model was tested using Python in Google Colab with the Transformers library. Code used to load the model from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/mujtabagulzarsoomro/Mujtaba-ids.texttext-generationn<1K0 likes2 downloads7mo agoHugging Face18stdt1 /match-ids-3to2 Matching ids of entities in 3rd compared to 2nd edition of Nordisk familjebok Manually created dataset for matching ids between 3rd and 2nd edition of Nordisk familjebok of 100 entries. The ids are refering to entities in the 2 editions. textn<1K0 likes1 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.