datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.prompt-injections-benchmark
Dataset: Qualifire Benchmark Prompt Injection(Jailbreak vs. Benign) Datasets
Overview
This dataset contains 5,000 prompts, each labeled as either jailbreak or benign. The dataset is designed for evaluating AI models' robustness against adversarial prompts and their ability to distinguish between safe and unsafe inputs.
Dataset Structure
Total Samples: 5,000
Labels: jailbreak, benign
Columns:
text: The input text
label: The classification (jailbreak or benign)… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/prompt-injections-benchmark.cvefixes-security-ir-graphrag
CVEfixes Security IR GraphRAG
This is a content-addressed, remotely routable Security IR release of the exact pinned CVEfixes snapshot. It packages the byte-identical original Parquet data together with a searchable corpus, BM25 postings, CUDA-generated vectors, typed graph nodes and edges, bounded adjacency indexes, and a source-CID-to-original-row lookup.
All entries are inert, non-authoritative evidence. Candidate and formal-logic rows cannot grant execution authority; exact… See the full description on the dataset page: https://huggingface.co/datasets/Publicus/cvefixes-security-ir-graphrag.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.security-kg
Security Knowledge Graph Triples
Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis.
Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV
Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/s0u9ata/security-kg.cyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/ruchit11111/coding-agent-security-benchmark.coding-agent-security-benchmark
Coding Agent Security Benchmark
A benchmark for evaluating whether an LLM can correctly identify security
violations in the behavior of an autonomous coding agent - spanning
dangerous shell commands, credential leakage, prompt injection, supply-chain
risk, privacy leaks, and more.
Each row is a single message sampled from a coding-agent session (a user
instruction, a tool call the agent issued, a tool's response, or the agent's
own output) paired with a ground-truth security… See the full description on the dataset page: https://huggingface.co/datasets/rogue-security/coding-agent-security-benchmark.security-paper-datasets
Dataset Card for "security-paper-datasets"
More Information needed
bitcoin-wallet-security-qa
Bitcoin Wallet Security Dataset
A high-quality question–answer dataset of 500 records focused on Bitcoin wallet
security, self-custody, backup and recovery planning, and common attack vectors. It is
built to train and evaluate AI systems that help people secure their Bitcoin — fine-tuning
LLMs, powering retrieval-augmented generation (RAG), security-focused assistants, and
educational chatbots.
Every record pairs a realistic security question with a detailed, self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-security-qa.security
Security Knowledge Graph Triples
Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis.
Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV
Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/security.security-master
US Security Master
Which company a ticker belonged to, on a date. Which filer a CUSIP points at.
What a company used to be called.
23 529 filers · 143 932 identifier spans · 10 056 033 dated observations ·
74 392 CUSIPs · 1960 to 2026
This is the glue for the rest of the family. Every other ZipLime dataset keys
on the SEC's CIK, which never changes — and every price series, broker feed and
research note keys on a ticker, which changes all the time. Joining the two
with today's… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/security-master.nemotron-terminal-security
nemotron-terminal-security
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "security". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-security.mmlu-computer_security
Dataset Card for "mmlu-computer_security"
More Information needed
llm-security-leaderboard-contentslinux-security-meanfield
Linux Security Meanfield Corpus
A commit-keyed corpus of security-relevant commits across 22 Linux
base-system repositories, unified on a single schema that carries both
CVE-dossiered fixes and non-CVE security-signal commits. This is the
Phase-1 release artifact for the mean-field survey paper.
Splits
cve_dossiered (2,254 rows): one row per
(fix_commit, CVE) pair from the scope-audited CVE dossier corpus.
non_cve_signal (21,609 rows): commit-anchored security
signal… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/linux-security-meanfield.security
Security Knowledge Graph Triples
Security data from 24 sources represented as Subject-Predicate-Object (SPO) triples in Parquet format, ready for knowledge-graph construction, graph-ML, RAG pipelines, and threat-intelligence analysis.
Sources: ATT&CK · CAPEC · CWE · CVE · CPE · D3FEND · ATLAS · CAR · ENGAGE · F3 · EPSS · KEV · Vulnrichment · GHSA · Sigma · ExploitDB · MISP Galaxies · LOLBAS · LOLDrivers · Atomic Red Team · NIST 800-53 · Nuclei · EUVD · OSV
Last updated:… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/security.HARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.task692_mmmlu_answer_generation_computer_security
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task692_mmmlu_answer_generation_computer_security
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task692_mmmlu_answer_generation_computer_security.ShellRisk-Bench
ShellRisk-Bench
ShellRisk-Bench is a reproducible benchmark for context-free binary risk
classification of individual shell-command submissions. It asks whether a
command poses meaningful cyber or system risk when evaluated without task,
user, or session context.
Release: The v0.1 Parquet train and test splits are publicly available
through Dataset Viewer and load_dataset().
The benchmark contains a deterministic train split of 16,772 rows and test
split of 4,194 rows. The test… See the full description on the dataset page: https://huggingface.co/datasets/kontext-security/ShellRisk-Bench.SecKnowledge-Eval
SecKnowledge 2.0 Evaluation Benchmark
The official evaluation benchmark suite from Toward Cybersecurity-Expert Small Language Models (ICML 2026), where we introduce the CyberPal 2.0 model family alongside these benchmarks. This repository releases the internal evaluation datasets developed to assess LLMs on core cybersecurity capabilities that existing public benchmarks do not adequately cover: adversarial robustness on CTI knowledge, cross-taxonomy reasoning, consequence-centric… See the full description on the dataset page: https://huggingface.co/datasets/cyber-pal-security/SecKnowledge-Eval.stackexchange_securitydatapoints_round1_dpsk_security_shard1_daytona_n100k1task733_mmmlu_answer_generation_security_studies
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task733_mmmlu_answer_generation_security_studies
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task733_mmmlu_answer_generation_security_studies.ai-agent-security-policy-decisions
AI Agent Security Policy Decisions
ai-agent-security-policy-decisions is a 2,400-record synthetic dataset for classifying proposed AI-agent tool actions as allow, deny, require_human_approval, or allow_with_restrictions. Each scenario includes identity and permission context, sensitivity, risk factors, required controls, a concise rationale, and a safer alternative.
The dataset addresses the decision point between an agent proposing an action and a tool or policy gateway… See the full description on the dataset page: https://huggingface.co/datasets/rksharma1947/ai-agent-security-policy-decisions.Organized_PreTrain_Cyber_Security_640kIndustryCorpus2_other_information_services_information_security
IndustryCorpus2: Information Services
This repository contains the IndustryCorpus2: Information Services domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao},
year… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_other_information_services_information_security.datapoints_round1_dpsk_security_shard2_daytona_n50k1terminal_bench_2_nemotron_terminal_security__Qwen3_8B_20260414_004927
