datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-author-affiliations-matched-ror-ids
arXiv Author Affiliations
This dataset contains author affiliation data extracted from arXiv works, matched to Research Organization Registry (ROR) identifiers.
Dataset Description
This dataset was generated from all arXiv works as of 2025/12. The source PDFs were converted to markdown using markitdown, and author affiliations were then extracted using cometadata/affiliation-parsing-lora-Qwen3-8B-distil-GLM_4.5_Air. The extracted affiliations were matched to ROR IDs using… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/arxiv-author-affiliations-matched-ror-ids.id_scholar_huridocsresearchqa_official_subset_idscc-ids-to-title
Common Crawl IDs to Titles
JSON manifests mapping document IDs to their extracted titles.
Usage
from datasets import load_dataset
titles = load_dataset("orionweller/cc-ids-to-title", "fw-edu")
Koda-IDS-CyberReasoning
Koda IDS CyberReasoning
Dataset Details
Dataset Description
Koda IDS CyberReasoning is a multi-domain reasoning dataset designed to support the development and evaluation of AI systems for cybersecurity incident detection, investigation, analysis, and response.
The dataset contains conversational reasoning examples covering cybersecurity triage, attack-pattern analysis, incident recovery, false-positive suppression, security-related coding, tool… See the full description on the dataset page: https://huggingface.co/datasets/netgoat-ai/Koda-IDS-CyberReasoning.Ishigaki-IDS-Bench
Ishigaki-IDS-Bench
Ishigaki-IDS-Bench is a bilingual benchmark for generating buildingSMART IDS 1.0 XML from user-facing building requirements.
Data Format
Each row uses the common conversational messages format:
{
"id": "row-0001",
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "..."}
],
"language": "ja",
"input_format": "natural_language",
"turn_type": "single_turn",
"domain": "architectural","ifc_versions":… See the full description on the dataset page: https://huggingface.co/datasets/ONESTRUCTION/Ishigaki-IDS-Bench.datacite-affiliations-matched-ror-ids-datacite-enrichment-format
DataCite Author Affiliations Matched to ROR IDs - DataCite Enrichment Format
This dataset contains 20,512,320 enrichment records mapping author affiliation strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 5,812,774 unique DOIs from the DataCite Public Data File.
Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the creators field… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-affiliations-matched-ror-ids-datacite-enrichment-format.wikipedia-ga-fa-ids
Wikipedia Good and Featured Articles
Contains the Wikipedia Article IDs and page titles for Good and Featured articles on Wikipedia for a given timestamped data dump, whereby the data was extracted from Wikipedia/Wikimedia SQL table dumps. This dataset covers 9 Language Wikipedia sites. For more information on how the dataset was generated see https://github.com/UCREL/wikipedia-ga-fa-extraction.
The data is specific to a given data dump timestamp, the main tag of the repository… See the full description on the dataset page: https://huggingface.co/datasets/ucrelnlp/wikipedia-ga-fa-ids.id-slang-synthetic-nlp
Indonesian Synthetic Informal Language NLP Dataset
A 20,000-record synthetic dataset of Indonesian informal language and slang designed for NLP research, language classification, sentiment and emotion analysis, conversational AI, and LLM-related research.
The dataset contains simulated Indonesian conversational examples with structured metadata describing informal expressions, emotional characteristics, and dominant regional context.
Dataset Summary
Property… See the full description on the dataset page: https://huggingface.co/datasets/e17do/id-slang-synthetic-nlp.datacite-funders-matched-ror-ids-datacite-enrichment-format
DataCite Funders Matched to ROR IDs - DataCite Enrichment Format
This dataset contains 1,008,697 enrichment records mapping funder name strings from DataCite metadata to Research Organization Registry (ROR) identifiers. It covers 697,791 unique DOIs from the DataCite Public Data File.
Each record is formatted as a DataCite enrichment input record, designed for use with the DataCite enrichment pipeline. Records use the updateChild action on the fundingReferences field, providing a… See the full description on the dataset page: https://huggingface.co/datasets/cometadata/datacite-funders-matched-ror-ids-datacite-enrichment-format.protein-diffusion-ids
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
protein_diffusion_ids
This dataset contains identifier strings for various protein design and structure prediction methods, including RFdiffusion, PepMLM, JointDiff, and AlphaFold. The entries consist of specific method names often appended with unique alphanumeric hashes, alongside null values indicating missing data. It appears to catalog unpublished affinity maturation techniques and… See the full description on the dataset page: https://huggingface.co/datasets/joduor/protein-diffusion-ids.duplicate_ids_spa_Latnaios-k2chj-pci-usb-ids
pci-usb-ids
Part of AIOS K\u00b2CHJ. See GitHub.
IDS-Bench
[!NOTE]
This result was produced as part of the "GENIAC (Generative AI Accelerator Challenge) Project" which is promoted by the Ministry of Economy, Trade and Industry and the New Energy and Industrial Technology Development Organization (NEDO), with the aim of strengthening Japan’s capabilities in generative AI development.
This dataset is an evaluation dataset for the CSV-to-IDS task used to evaluate the Ishigaki-IDS model. It evaluates whether an LLM can generate appropriate IDS from CSV… See the full description on the dataset page: https://huggingface.co/datasets/ONESTRUCTION/IDS-Bench.vatex-idsaios-k2chj-pci-usb-ids
💾 SDIO HWIDs — The Largest Public Hardware ID Collection
171.003 unique PCI/USB/ACPI hardware identifiers extracted from 65 Windows DriverPacks.
🎯 What is this?
This dataset contains 171.003 unique hardware identifiers (HWIDs) extracted from 65 SDIO DriverPacks — the largest collection of Windows hardware IDs assembled for AI training.
Each entry is a raw HWID string like:
PCI\VEN_10DE&DEV_1E81&SUBSYS_8597174B
USB\VID_0A5C&PID_5848&REV_0102… See the full description on the dataset page: https://huggingface.co/datasets/msrovani/aios-k2chj-pci-usb-ids.Mujtaba-idsBlind Spots of Frontier Models
Model Tested
Model name: TinyLlama-1.1B-intermediate-step-1431k-3T
Model link:
https://huggingface.co/TinyLlama/TinyLlama-1.1B-intermediate-step-1431k-3T
This model is a base pretrained language model released on Hugging Face and was not specifically fine tuned for a particular downstream task.
How the Model Was Loaded
The model was tested using Python in Google Colab with the Transformers library.
Code used to load the model
from transformers import… See the full description on the dataset page: https://huggingface.co/datasets/mujtabagulzarsoomro/Mujtaba-ids.match-ids-3to2
Matching ids of entities in 3rd compared to 2nd edition of Nordisk familjebok
Manually created dataset for matching ids between 3rd and 2nd edition of Nordisk familjebok of 100 entries. The ids are refering to entities in the 2 editions.
