datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.White-Hat-Security-Agent-Prompts-600K
White Hat Security Agent Prompts 600K
Overview
The White-Hat-Security-Agent-Prompts-600K dataset is a practitioner-perspective security prompts corpus of 596,295 richly contextualized queries, designed to represent how real-world defensive security professionals communicate, interrogate, and reason through active threat scenarios.
Where most security datasets catalogue CVEs, malware signatures, or CTF write-ups, this collection teaches models to operate from inside the… See the full description on the dataset page: https://huggingface.co/datasets/yatin-superintelligence/White-Hat-Security-Agent-Prompts-600K.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.cyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.bitcoin-wallet-security-qa
Bitcoin Wallet Security Dataset
A high-quality question–answer dataset of 500 records focused on Bitcoin wallet
security, self-custody, backup and recovery planning, and common attack vectors. It is
built to train and evaluate AI systems that help people secure their Bitcoin — fine-tuning
LLMs, powering retrieval-augmented generation (RAG), security-focused assistants, and
educational chatbots.
Every record pairs a realistic security question with a detailed, self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-security-qa.Security-TTP-Mapping
The Security Attack Pattern (TTP) Recognition or Mapping Task
We share in this repo the MITRE ATT&CK mapping datasets, with training, validation and test splits.
The datasets can be considered as an emerging and challenging multilabel classification NLP task, with over 600 hierarchical classes.
NOTE: due to their security nature, these datasets contain textual information about malware and other security aspects.
Datasets
TRAM
This dataset belongs to CTID… See the full description on the dataset page: https://huggingface.co/datasets/tumeteor/Security-TTP-Mapping.cairo-security-audits
Cairo Security Audits
A source-traceable corpus of public Cairo and Starknet security-audit metadata and normalized finding annotations.
Version 0.3.0 packages every entry in the audit inventory frozen at keep-starknet-strange/starknet-skills@17a76e8. It covers 32 accessible reports from 10 auditing firms and 286 normalized finding annotations. Eleven records are checked against rendered reports and two link to exact vulnerable/fixed commits. The release does not redistribute… See the full description on the dataset page: https://huggingface.co/datasets/starknet-ai/cairo-security-audits.nemotron-terminal-security
nemotron-terminal-security
Per-source partition of nvidia/Nemotron-Terminal-Corpus,
filtered to source == "security". The difficulty column preserves the original
easy / medium / mixed split (na for the dataset_adapters/* files, which
did not carry a difficulty label).
Partitioning scheme:
adapters_{code,math,swe} — rows from dataset_adapters/{code,math,swe}.parquet
{skill} (e.g. debugging, security, …) — rows from
synthetic_tasks/skill_based/{easy,medium… See the full description on the dataset page: https://huggingface.co/datasets/laion/nemotron-terminal-security.turkish_cyber_security_controls_benchmark
Turkish Cyber Security Controls Benchmark
Türkçe siber güvenlik kontrol seçimi ve kontrol denetimi yeteneğini ölçmek için
hazırlanmış, senaryo tabanlı çoktan seçmeli değerlendirme kümesidir.
v0.1.0, uzman incelemesine açık ilk sürümdür ve NIST SP 800-53 Rev. 5,
Release 5.2.0 kontrol kataloğunu hedefler.
Kapsam
100 Türkçe senaryo
NIST SP 800-53'ün 20 kontrol ailesinin her birinden 5 soru
64 kontrol seçimi sorusu
17 denetim kanıtı sorusu
19 denetim yargısı sorusu… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/turkish_cyber_security_controls_benchmark.oauth-api-security-en
OAuth & API Security Dataset (EN)
Comprehensive English dataset covering OAuth 2.0 vulnerabilities, API attacks (OWASP API Top 10 2023), security controls, and Q&A pairs for training cybersecurity-specialized language models.
Dataset Contents
Category
Entries
Description
OAuth 2.0 Vulnerabilities
20
Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage
API Attacks
25
BOLA, BFLA, BOPLA, SSRF, GraphQL DoS, gRPC injection, CORS… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-en.dspy-security-bench-v01-results
dspy-security-bench: v0.1 + v0.1.1 results
Raw evaluation outputs from
dspy-security-bench.
Cite or audit these numbers without needing to clone the repo or re-run
the benchmark.
What's in here
File
Contents
Rows
workspace_v01_results.csv
Original v0.1 launch run. Workspace suite, 3 optimizers (unoptimized, BootstrapFewShot, MIPROv2 light), 2 attacks (direct, important_instructions), N=5 user × 1 injection × 1 seed.
30
workspace_v01_summary.csv
v0.1… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-v01-results.oauth-api-security-fr
Dataset OAuth & Securite API (FR)
Dataset francophone complet sur les vulnerabilites OAuth 2.0, les attaques API (OWASP API Top 10 2023), les controles de securite, et les questions-reponses pour l'entrainement de modeles de langage specialises en cybersecurite.
Contenu du Dataset
Categorie
Nombre d'entrees
Description
Vulnerabilites OAuth 2.0
20
Authorization Code Interception, CSRF, PKCE bypass, JWT attacks, token leakage
Attaques API
25
BOLA, BFLA, BOPLA… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/oauth-api-security-fr.SecKnowledge-Eval
SecKnowledge 2.0 Evaluation Benchmark
The official evaluation benchmark suite from Toward Cybersecurity-Expert Small Language Models (ICML 2026), where we introduce the CyberPal 2.0 model family alongside these benchmarks. This repository releases the internal evaluation datasets developed to assess LLMs on core cybersecurity capabilities that existing public benchmarks do not adequately cover: adversarial robustness on CTI knowledge, cross-taxonomy reasoning, consequence-centric… See the full description on the dataset page: https://huggingface.co/datasets/cyber-pal-security/SecKnowledge-Eval.ai-agent-security-policy-decisions
AI Agent Security Policy Decisions
ai-agent-security-policy-decisions is a 2,400-record synthetic dataset for classifying proposed AI-agent tool actions as allow, deny, require_human_approval, or allow_with_restrictions. Each scenario includes identity and permission context, sensitivity, risk factors, required controls, a concise rationale, and a safer alternative.
The dataset addresses the decision point between an agent proposing an action and a tool or policy gateway… See the full description on the dataset page: https://huggingface.co/datasets/rksharma1947/ai-agent-security-policy-decisions.security-attacks-MITREcloud-security-en
Bilingual Cloud Security Dataset
🔒 Description
Complete bilingual (English/French) cloud security dataset covering AWS, Azure, and GCP. This dataset contains detailed information about common misconfigurations, cloud security best practices, and educational Q&A.
Total Data: 80 misconfigurations + 50 best practices + 100 Q&A pairs (FR + EN)
📊 Content
1. Cloud Misconfigurations (80 entries)
Comprehensive catalog of common cloud configuration errors:… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/cloud-security-en.llm-security-en
LLM Security & Prompt Injection Dataset (EN)
Comprehensive English dataset on Large Language Model (LLM) security, covering attack techniques, defense patterns, OWASP LLM Top 10 (2025) and AI Act compliance.
Description
This dataset provides a structured knowledge base in English for training, fine-tuning and awareness on the security of LLM-based applications. It covers the full landscape of threats and defenses for generative AI systems.
Dataset Content… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/llm-security-en.omnimcp_supabase_row_level_security_ai_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_supabase_row_level_security_ai_teaser.bitcoin-security-reasoning-100k
Dataset Card for Bitcoin Security Reasoning 100K
100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code.
Dataset Details
Dataset Description
This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.security-instructions
Security Insutrctions 2.5K
A set of Cybersecurity questions pertaining to different areas of security.
flaws-cloudtrail-security-qa
CloudTrail Security Q&A Dataset
A comprehensive dataset of security-focused questions and answers based on AWS CloudTrail logs, designed for training and evaluating AI agents on cloud security analysis tasks.
Dataset Overview
This dataset contains:
~150 questions across 16 CloudTrail database partitions
Time period: February 2017 - August 2020
4 different AI models used for question generation
DuckDB databases with actual CloudTrail data
Mixed answerable/unanswerable… See the full description on the dataset page: https://huggingface.co/datasets/odemzkolo/flaws-cloudtrail-security-qa.k8s-security-en
Kubernetes Security - Complete Dataset
A comprehensive bilingual (FR/EN) dataset covering misconfigurations, attack paths, and Q&A about Kubernetes security.
📊 Dataset Content
1. Kubernetes Misconfigurations (60+ entries)
Categories: RBAC, PodSecurity, NetworkPolicy, Secrets, ImageSecurity, RuntimeSecurity, Admission, Logging, APIServer, etcd
Fields: ID, name, category, descriptions FR/EN, risk level (Critical/High/Medium/Low)
Additional Details: remediation… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/k8s-security-en.dspy-security-bench-trainset-workspace
dspy-security-bench: workspace trainset (v0.1)
This is the synthetic, environment-grounded query-only trainset used to
optimize DSPy programs in v0.1 of
dspy-security-bench,
a benchmark that measures whether DSPy prompt optimization affects the
prompt-injection robustness of agentic LLM programs.
What's in here
192 query / ground-truth pairs grounded in the
AgentDojo workspace suite's
default environment (calendar, inbox, files).
{"prompt": "What is the… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-trainset-workspace.supply-chain-security
Supply Chain Security Dataset
A comprehensive bilingual (French/English) dataset focused on software supply chain security, covering modern frameworks, best practices, and notable security incidents.
Dataset Description
This dataset provides instruction-response pairs for training AI models on supply chain security topics. It covers critical areas including SBOM management, dependency security, code signing, vulnerability management, and real-world attack case studies.… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/supply-chain-security.quantum-cryptography-and-post-quantum-security
Neura Parse — Quantum Cryptography & Post-Quantum Security
A deep vertical on cryptography that uses quantum mechanics and on classical cryptography built to resist quantum attack. It covers quantum key distribution (BB84, B92, six-state, SARG04, E91, BBM92, decoy-state, MDI-QKD, TF-QKD, CV-QKD), device-independent protocols, composable and finite-key security proofs, quantum hacking with countermeasures, classical post-processing (reconciliation, privacy amplification… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-cryptography-and-post-quantum-security.AIForge-1K-Security
AIForge-03-Security
Security Dataset for AI and Programming Tasks
Overview
AIForge-03-Security is a curated English dataset designed for AI systems working on security tasks in software engineering and programming.
Contents
data.jsonl
data.json
metadata.json
Use Cases
AI agent training
Supervised fine-tuning
Evaluation and benchmarking
Software engineering research
Example Record
{
"id": "AISEC_00001"… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/AIForge-1K-Security.ethereum-smart-contract-security-qa
Ethereum Smart Contract Security Dataset
A high-quality question–answer dataset of 100 records focused exclusively on Ethereum
smart contract security. It is built to train and evaluate AI systems that explain, identify,
classify, and mitigate the most common Ethereum smart contract vulnerabilities — LLM
fine-tuning, retrieval-augmented generation (RAG), AI security assistants, and secure
Solidity education.
Every record explains one vulnerability, attack pattern, secure coding… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/ethereum-smart-contract-security-qa.DoD-Instruction-5200-01-Information-Security-Program
DoD Information Security and SCI Protection Question-Answer Dataset
Maintainer: Terry Eppler
Owner: US Federal Government
Dataset Summary
This dataset contains document-grounded question-and-answer records based on DoD Instruction 5200.01, “DoD Information Security Program and Protection of Sensitive Compartmented Information (SCI),” dated April 21, 2016, and incorporating Change 2 effective October 1, 2020.
The source establishes the overarching Department… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-5200-01-Information-Security-Program.combine-llm-security-benchmark
Combined LLM Security Benchmark 🔐
A comprehensive, unified benchmark dataset for evaluating Large Language Models (LLMs) on cybersecurity tasks. This dataset combines 10 security benchmarks into a standardized format with 18,059 examples across 5 task types.
📊 Dataset Summary
This dataset consolidates multiple security-focused benchmarks into a single, easy-to-use format for comprehensive LLM evaluation across various cybersecurity domains:
Total Examples: 18,059
Total… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/combine-llm-security-benchmark.oak-security-sft
OAK (On-chain Attack Knowledge) Security SFT Corpus
A bilingual (EN+RU) instruction-tuning dataset for training language models to analyze on-chain attacks, DeFi exploits, blockchain forensics, and crypto security incidents. Every answer is grounded in the OAK taxonomy v0.1 — a structured knowledge base of adversary tactics, techniques, mitigations, software tools, threat groups, and real-world on-chain incidents.
Dataset Overview
Metric
Value
Total… See the full description on the dataset page: https://huggingface.co/datasets/z0n3x/oak-security-sft.
