datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
battlefield-medic-sharegpt
🏥⚔️ Synthetic Battlefield Medical Conversations
For the multilingual version (non-sharegpt foormat) that includes the title columns go here https://huggingface.co/datasets/nisten/battlefield-medic-multilingual
Over 3000 conversations incorporating 2000+ human diseases and over 1000 battlefield injuries from various scenarios
Author: Nisten Tahiraj
License: MIT
This dataset consists of highly detailed synthetic conversations… See the full description on the dataset page: https://huggingface.co/datasets/nisten/battlefield-medic-sharegpt.nist_800_53motor-fuel-dispenser-measurement-tolerances-nist
How far off a US fuel pump, LPG meter or EV charger is allowed to be (NIST Handbook 44)
Canonical, always-current version: https://referencesource.org/motor-fuel-dispenser-measurement-tolerances-nist/
Machine-readable: https://referencesource.org/motor-fuel-dispenser-measurement-tolerances-nist/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-20
Stale after: 2027-08-20 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/motor-fuel-dispenser-measurement-tolerances-nist.battlefield-medic-multilingual
🏥⚔️ Synthetic Battlefield Medical Conversations
Over 3000 conversations incorporating 2000+ human diseases and over 1000 battlefield injuries from various scenarios
Author: Nisten Tahiraj
License: MIT
Note that unlike the other english dataset I posted here these are NOT in sharegpt format but include a compatible conversation inside the json which makes it EASY to convert to sharegpt or chatml.I will post the conversetion script to… See the full description on the dataset page: https://huggingface.co/datasets/nisten/battlefield-medic-multilingual.nist-cryptographic-algorithm-deprecation-schedule
NIST cryptographic algorithm deprecation schedule: what is approved, deprecated, and disallowed, and when
Canonical, always-current version: https://referencesource.org/nist-cryptographic-algorithm-deprecation-schedule/
Machine-readable: https://referencesource.org/nist-cryptographic-algorithm-deprecation-schedule/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-16
Stale after: 2027-02-12 (past this date, prefer the canonical copy —
it re-verifies on a… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nist-cryptographic-algorithm-deprecation-schedule.nist-juliet-cnist-csf-2.0-sft
NIST CSF 2.0 — Policy-Following SFT Dataset
A supervised fine-tuning (SFT) dataset of 266 chat-format records grounded verbatim in the
NIST Cybersecurity Framework (CSF) 2.0 (NIST CSWP 29, 2024-02-26).
Built to teach a model to answer cybersecurity-governance questions faithfully to the framework — no invented controls or requirements.
Format
Each record is OpenAI chat format with a system / user / assistant turn, plus metadata:
{
"id": "cat-GV.RR"… See the full description on the dataset page: https://huggingface.co/datasets/SeanJIE250/nist-csf-2.0-sft.NIST-CVE-DATASETalcapybaraAlbanian translation pipeline + tree-of-thought enhancement of https://huggingface.co/datasets/LDJnr/Capybara shareGPT formated dataset for efficient multiturn training of any AI to learn Albanian.
Work in progress. This 8k subset should be good to try training with. The pipline filtering and cleanup leaves out a lot of data, therefor expect this to improve.
###Sample code that trains (still under construction, expect to change)
#Author: Nisten Tahiraj
#License: Apache 2.0
##… See the full description on the dataset page: https://huggingface.co/datasets/nisten/alcapybara.NIST-SP-800-171r2This dataset contains the requirements found in Chapter 3 of the NIST SP 800-171 document titled Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations.
adaption-nist-biosafety-seq-bench
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-nist_biosafety_seq_bench
This dataset comprises structured biological sequence records designed as a strategic benchmark for training AI systems in biosafety, biosecurity, and synthetic biology governance. Each sample includes genomic data, organism identifiers, risk classification labels, and review status metadata to support pathogen detection and function prediction tasks. The… See the full description on the dataset page: https://huggingface.co/datasets/joduor/adaption-nist-biosafety-seq-bench.nisten__franqwenstein-35b-details
Dataset Card for Evaluation run of nisten/franqwenstein-35b
Dataset automatically created during the evaluation run of model nisten/franqwenstein-35b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nisten__franqwenstein-35b-details.defendable-pain-nist-airmf-pain-v0.1
NIST AI RMF Pain Receipt
"the framework" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 9 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
9 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-nist-airmf-pain-v0.1.nisten__tqwendo-36b-details
Dataset Card for Evaluation run of nisten/tqwendo-36b
Dataset automatically created during the evaluation run of model nisten/tqwendo-36b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/nisten__tqwendo-36b-details.NIST_CSF_2.0_Fine_Tune
