datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nist-publications-raw
NIST Publications - Raw PDFs
596 NIST cybersecurity publications in original PDF format - Complete source data for the nist-cybersecurity-training dataset.
Dataset Description
This dataset contains the raw, unprocessed PDF files downloaded from the NIST Computer Security Resource Center (CSRC). These are the exact source documents used to create the NIST cybersecurity training dataset and fine-tune the HackIDLE-NIST-Coder model.
Contents
596 PDF documents (2.0… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-publications-raw.NIST-In-Situ-IN625-LPBF-Overhangsnist-cybersecurity-training
NIST Cybersecurity Training Dataset v1.1
The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs
Version 1.1 Highlights
What's New in v1.1:
✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents
✅ Fixed 6,150 broken DOI links via format normalization
✅ Removed 202 malformed DOIs (double URL prefixes)
✅ Validated and fixed 124,946 total links
✅ Cataloged 72,698 broken links for future recovery
✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.opus-doctor-patient-conversations-all-human-diseases
Opus-4.8-High-Thinking generated Doctor-Patient Conversations for All Human Diseases
Covers every human disease listed on my previous work here: nisten/all-human-diseases
The dataset strictly used Opus 4.8 - High and was cleaned over 3 times via Opus 4.8, 4.7 and 4.6. Minor corrections were needed upon each pass mainly to bypass single word safety filters like i.e. monkeypox.
The main hallucination noticed during generation was that Opus would make up wrong PMID ( PubMed ID )… See the full description on the dataset page: https://huggingface.co/datasets/nisten/opus-doctor-patient-conversations-all-human-diseases.NIST
NIST Publications - Raw PDFs
596 NIST cybersecurity publications in original PDF format - Complete source data for the nist-cybersecurity-training dataset.
Dataset Description
This dataset contains the raw, unprocessed PDF files downloaded from the NIST Computer Security Resource Center (CSRC). These are the exact source documents used to create the NIST cybersecurity training dataset and fine-tune the HackIDLE-NIST-Coder model.
Contents
596 PDF… See the full description on the dataset page: https://huggingface.co/datasets/roundrock/NIST.nist-gdt-pmi-vlm-benchmark
NIST GD&T/PMI VLM Benchmark
This benchmark measures exact-match transcription of geometric dimensioning and tolerancing (GD&T) and product and manufacturing information (PMI) from rendered NIST Fully-Toleranced Test Case drawing pages. A row supplies the page image and target element_id; the expected output is one engineering-significant specification string.
The reported evaluation uses open transcription: image plus element_id.
page_answer_choices is included for anyone who… See the full description on the dataset page: https://huggingface.co/datasets/CLARKBENHAM/nist-gdt-pmi-vlm-benchmark.details_nisten__smaugzilla-77b
Dataset Card for Evaluation run of nisten/smaugzilla-77b
Dataset automatically created during the evaluation run of model nisten/smaugzilla-77b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nisten__smaugzilla-77b.details_nisten__shqiponja-15b-v1
Dataset Card for Evaluation run of nisten/shqiponja-15b-v1
Dataset automatically created during the evaluation run of model nisten/shqiponja-15b-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nisten__shqiponja-15b-v1.battlefield-medic-sharegpt
🏥⚔️ Synthetic Battlefield Medical Conversations
For the multilingual version (non-sharegpt foormat) that includes the title columns go here https://huggingface.co/datasets/nisten/battlefield-medic-multilingual
Over 3000 conversations incorporating 2000+ human diseases and over 1000 battlefield injuries from various scenarios
Author: Nisten Tahiraj
License: MIT
This dataset consists of highly detailed synthetic conversations… See the full description on the dataset page: https://huggingface.co/datasets/nisten/battlefield-medic-sharegpt.all-human-diseases
DRAFT LIST OF ALL HUMAN DISEASES
Way more data coming.
License is AGPL to avoid predatory players, I don't care if you or your startup use it.
If you see an issue comment on huggingface or this google doc.
https://docs.google.com/spreadsheets/d/14MYgMc9CZOZEdp63utIxvATTtTIVzj1RSjf1jAUtWiM/edit?usp=sharing
Scientific evidence-based medicine only, no opinions please. I repeat, ZERO OPINIONS please.
Sources:
https://www.cdc.gov/health-topics.html… See the full description on the dataset page: https://huggingface.co/datasets/nisten/all-human-diseases.nist_800_53battlefield-medic-multilingual
🏥⚔️ Synthetic Battlefield Medical Conversations
Over 3000 conversations incorporating 2000+ human diseases and over 1000 battlefield injuries from various scenarios
Author: Nisten Tahiraj
License: MIT
Note that unlike the other english dataset I posted here these are NOT in sharegpt format but include a compatible conversation inside the json which makes it EASY to convert to sharegpt or chatml.I will post the conversetion script to… See the full description on the dataset page: https://huggingface.co/datasets/nisten/battlefield-medic-multilingual.NIST-CyberSecurity-Framework
# NIST Cybersecurity Framework 2.0 Question Answering Dataset
Dataset Summary
The NIST Cybersecurity Framework 2.0 Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Cybersecurity
Framework (CSF) 2.0.
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to cybersecurity risk management,
cybersecurity governance, enterprise risk management… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-CyberSecurity-Framework.nist_mugshotsmotor-fuel-dispenser-measurement-tolerances-nist
How far off a US fuel pump, LPG meter or EV charger is allowed to be (NIST Handbook 44)
Canonical, always-current version: https://referencesource.org/motor-fuel-dispenser-measurement-tolerances-nist/
Machine-readable: https://referencesource.org/motor-fuel-dispenser-measurement-tolerances-nist/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-20
Stale after: 2027-08-20 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/motor-fuel-dispenser-measurement-tolerances-nist.NIST-GenAI-Profile
NIST Generative AI Profile Question Answering Dataset
Dataset Summary
This dataset contains question-and-answer records derived from the National Institute of Standards and Technology publication NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. The source publication is a cross-sectoral companion resource to the NIST AI Risk Management Framework and focuses on risks that are unique to, or exacerbated by… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-GenAI-Profile.deepwiki-public-repo-reviews
DeepWiki Repository Summaries
A single-turn instruction dataset of full technical repository summaries, sourced from
DeepWiki — an AI-generated wiki platform for GitHub repositories.
Each record is one question → one long-form answer covering the architecture, components,
data flows, APIs, and implementation details of a GitHub repository.
Archive Status - LAST UPDATED APRIL 15 2026
Stat
Value
Repos archived (done)
6,920
Repos seeded (total probed)
7,852… See the full description on the dataset page: https://huggingface.co/datasets/nisten/deepwiki-public-repo-reviews.nist-cryptographic-algorithm-deprecation-schedule
NIST cryptographic algorithm deprecation schedule: what is approved, deprecated, and disallowed, and when
Canonical, always-current version: https://referencesource.org/nist-cryptographic-algorithm-deprecation-schedule/
Machine-readable: https://referencesource.org/nist-cryptographic-algorithm-deprecation-schedule/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-16
Stale after: 2027-02-12 (past this date, prefer the canonical copy —
it re-verifies on a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nist-cryptographic-algorithm-deprecation-schedule.nist-csf-en
🛡️ NIST Cybersecurity Framework 2.0 - English Dataset
Comprehensive bilingual NIST Cybersecurity Framework 2.0 dataset (February 2024) in French and English, prepared for HuggingFace.
📋 Dataset Content
This dataset includes complete coverage of NIST CSF 2.0 framework:
1. Functions (6 entries)
🏛️ Govern (GV)
🔍 Identify (ID)
🛡️ Protect (PR)
🔎 Detect (DE)
⚡ Respond (RS)
🔄 Recover (RC)
2. Categories & Subcategories (35+ entries)
Detailed… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/nist-csf-en.NIST-AI-Risk-Management-Framework
# NIST AI Risk Management Framework Question Answering Dataset
Dataset Summary
The NIST AI Risk Management Framework Question Answering Dataset is a synthetic
instruction-style question-answering dataset derived from the NIST Artificial
Intelligence Risk Management Framework (AI RMF 1.0).
The dataset is designed to support training, fine-tuning, retrieval evaluation, and
domain-specific question-answering use cases related to AI risk management,
trustworthy AI, responsible AI… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-AI-Risk-Management-Framework.kl3m-filter-data-dotgov-www.nist.govdetails_nisten__shqiponja-59b-v1
Dataset Card for Evaluation run of nisten/shqiponja-59b-v1
Dataset automatically created during the evaluation run of model nisten/shqiponja-59b-v1 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nisten__shqiponja-59b-v1.ATLAS-NIST-Dataset-v2
ATLAS-NIST-Dataset-v2
Dataset Overview
This dataset contains 3,000 synthetic samples designed for training Risk Assessment Small Language Models (SLMs) in the Welfare/Public Service domain, specifically focusing on unemployment benefit scenarios.
Attribution
This release marks the Anna Ko Milestone, engineered following specific requirements provided by Anna Ko to ensure rigorous validation and regulatory compliance.
History of Improvements
V1… See the full description on the dataset page: https://huggingface.co/datasets/nislam-mics/ATLAS-NIST-Dataset-v2.details_nisten__bigdoc-c34b-instruct-tf32
Dataset Card for Evaluation run of nisten/bigdoc-c34b-instruct-tf32
Dataset automatically created during the evaluation run of model nisten/bigdoc-c34b-instruct-tf32 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nisten__bigdoc-c34b-instruct-tf32.details_nisten__BigCodeLlama-92b
Dataset Card for Evaluation run of nisten/BigCodeLlama-92b
Dataset automatically created during the evaluation run of model nisten/BigCodeLlama-92b on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_nisten__BigCodeLlama-92b.nist-coa-pdfThis is a set of chemical levels measured in NIST SRMs as reported in their Certificates of Analysis (COA) PDF documents. Data was manually extracted by Dr. Lane C. Sander of NIST.NIST-Privacy-Framework
NIST Privacy Framework Dataset
Dataset Summary
This dataset contains question-and-answer records derived from NIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0, published by the National Institute of Standards and Technology on January 16, 2020.
The source document provides a voluntary, risk-based framework for helping organizations improve privacy through enterprise risk management. It is designed to support privacy… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-Privacy-Framework.NISTSP800NIST-Managing-AI-Misuse-Risk
NIST Managing Misuse Risk for Dual-Use Foundation Models Question Answering Dataset
Dataset Summary
This dataset contains question-and-answer records derived from NIST AI 800-1 2pd, Managing Misuse Risk for Dual-Use Foundation Models, a second public draft issued by the U.S. AI Safety Institute at the National Institute of Standards and Technology in January 2025.
The source document provides voluntary guidance for improving the safety, security, and trustworthiness of… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/NIST-Managing-AI-Misuse-Risk.msms_nist_disjoint_probe_retrieval_20260622
NIST MS/MS Murcko-Disjoint Probe and Retrieval Benchmark
This dataset packages one fixed online-probe task and two fixed retrieval-task
pair tables derived from the raw NIST high-resolution MS/MS MGF staged in
raw/hr_msms_nist.mgf.
Layout
nist_100k_online_probe/: train/val/test Parquets built from an exact
100,000-spectrum sample without
replacement, then processed by the existing online-probe Murcko split logic.
Produced rows: train=68,870, val=12,862,
test=16… See the full description on the dataset page: https://huggingface.co/datasets/wchen99998/msms_nist_disjoint_probe_retrieval_20260622.
