datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum computing risks… See the full description on the dataset page: https://huggingface.co/datasets/Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Cybersecurity-Dataset-Fenrir-v2.1
Cybersecurity Defense Instruction-Tuning Dataset (v2.1)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.
1 What’s… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1.cybersecurity-qa-v2
Cybersecurity Q&A Dataset v2 — 2.6M Examples
A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics.
2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies.
Statistics
Source
Examples
Description
NIST NVD CVE Database
~1,954,225
All CVEs (2002–2025): overview, severity, detection, remediation
AlicanKiraz0/All-CVE-Records-Training-Dataset
~297,441
Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/rezaduty/cybersecurity-qa-v2.cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.nist-cybersecurity-training
NIST Cybersecurity Training Dataset v1.1
The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs
Version 1.1 Highlights
What's New in v1.1:
✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents
✅ Fixed 6,150 broken DOI links via format normalization
✅ Removed 202 malformed DOIs (double URL prefixes)
✅ Validated and fixed 124,946 total links
✅ Cataloged 72,698 broken links for future recovery
✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.Cybersecurity_Reasoning_Dataset
Cybersecurity Reasoning Dataset (Model-Agnostic)
A model-agnostic re-architecture of the Cybersecurity Reasoning Dataset. The original
corpus was format-bound to the Mistral/Llama ### Instruction: / ### Response: template;
this dataset losslessly separates reasoning content from format, providing one
neutral canonical corpus plus four per-family rendered training variants
(Mistral/Llama, DeepSeek, ChatML, Gemma).
Why this exists. Identical content scored 88.1 on a… See the full description on the dataset page: https://huggingface.co/datasets/dpevzner/Cybersecurity_Reasoning_Dataset.cybersecurity-master-dataset
Cybersecurity Master Dataset
Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions.
CyberSecurity_OWASP-sft-dataset
CyberSecurity_OWASP SFT Dataset
This dataset contains verifier-gated supervised fine-tuning examples for the
CyberSecurity_OWASP OpenEnv environment. Each row teaches one step of the
defensive local AppSec workflow: inspect policy/code, reproduce a local
authorization failure, submit a policy-tied diagnosis, patch the generated app,
run visible tests, and submit the fix.
Every kept trajectory is executed against the real local environment and must
pass the deterministic reward… See the full description on the dataset page: https://huggingface.co/datasets/Humanlearning/CyberSecurity_OWASP-sft-dataset.CyberSecurity-1M
CyberSecurity-1M
A large-scale, multi-source cybersecurity knowledge dataset containing 1.19M records across 16 categories, collected exclusively for academic, non-commercial research purposes. Last updated: 2026-05-27.
Disclaimer: This dataset is provided for academic research only. All content is aggregated from publicly available sources. The views, opinions, and information expressed in the dataset content do not represent the views or positions of the research team. The… See the full description on the dataset page: https://huggingface.co/datasets/WhitzardAgent/CyberSecurity-1M.Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.omnimcp_cybersecurity_secops_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/omnimcp_cybersecurity_secops_teaser.cyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.cybersecurity-theory-sft-gemma12b
Cybersecurity Theory SFT (Gemma 12B pack)
Curated 21,265-row cybersecurity theory instruction pack for LoRA supervised fine-tuning. Each example is a single-turn user → assistant pair covering offensive/defensive concepts, frameworks, CTF reasoning, vulnerability catalogs, and security tooling literacy — without agent tool traces or multi-turn harness data.
Paired MLX LoRA adapter trained on this pack (Nemotron 3 Super… See the full description on the dataset page: https://huggingface.co/datasets/True2456/cybersecurity-theory-sft-gemma12b.Cybersecurity-Dataset-Heimdall-v1.1
Cybersecurity Defense Instruction-Tuning Dataset (v1.1)
TL;DR
21 258 high‑quality system / user / assistant triples for training alignment‑safe, defensive‑cybersecurity LLMs. Curated from 100 000 + technical sources, rigorously cleaned and filtered to enforce strict ethical boundaries. Apache‑2.0 licensed.
1 What’s new in v1.1 (2025‑06‑21)
Change
v1.0
v1.1
Rows
2 500
21 258 (+760 %)
Covered frameworks
OWASP Top 10, NIST CSF
+ MITRE ATT&CK, ASD… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-Heimdall-v1.1.cybersecurity-soc-threat-hunting-sft-dpo-2026
🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects.
📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format)
A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles.
Dataset Description
This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.cybersecurity-questionaire
Dataset Card for cybersecurity-questionaire
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/MichaelPrimez/cybersecurity-questionaire/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MichaelPrimez/cybersecurity-questionaire.cybersecurity-sft-100k
Cybersecurity SFT 100K
A synthetic supervised fine-tuning dataset of 100,000 high-quality cybersecurity conversations designed to train AI assistants for security operations, threat analysis, incident response, and defensive security engineering.
Dataset Description
This dataset covers real-world security scenarios across 9 cybersecurity domains. Each record follows the ShareGPT conversation format with a practitioner-level query and a detailed, structured… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/cybersecurity-sft-100k.cybersecurity-sft-dataset
Cybersecurity SFT Dataset
A curated dataset for training cybersecurity-focused code models with structured JSON output capability.
Dataset Composition
Source
Count
Percentage
Description
CVE Records
10,000
50.0%
Multi-turn CVE vulnerability analysis
OpenCodeReasoning (NVIDIA)
5,000
25.0%
Chain-of-thought code reasoning
Code-Feedback
5,000
25.0%
Multi-turn code debugging and refinement
Synthetic Security (JSON)
5
<0.1%
JSON-structured CVE, MITRE ATT&CK… See the full description on the dataset page: https://huggingface.co/datasets/moro72842/cybersecurity-sft-dataset.cybersecurity-dataset-fenrir
Cybersecurity Defense Instruction-Tuning Dataset (v2.1)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/cybersecurity-dataset-fenrir.Cybersecurity-Dataset-Fenrir-v2.1
Cybersecurity Defense Instruction-Tuning Dataset (v2.1)
Created by Alican Kiraz
TL;DR
A ready-to-train dataset of 99,870 high-quality system / user / assistant triples for defensive, alignment-safe cybersecurity SFT training.
Apache-2.0 licensed and production-ready.
Scope: OWASP Top 10, MITRE ATT&CK, NIST CSF, CIS Controls, ASD Essential 8, modern authentication (OAuth 2 / OIDC / SAML), SSL / TLS, Cloud & DevSecOps, Cryptography, and AI Security.… See the full description on the dataset page: https://huggingface.co/datasets/Machivelli/Cybersecurity-Dataset-Fenrir-v2.1.ai-cybersecurity-en
AI in Offensive and Defensive Cybersecurity - English Dataset
Description
Comprehensive bilingual dataset covering the use of Artificial Intelligence in cybersecurity, from both the offensive (attackers) and defensive (defenders) perspectives. This is the English version.
Articles Covered
This dataset synthesizes knowledge from the following articles:
Offensive AI: How Attackers Use LLMs - LLM-based attack techniques
AI Threat Detection - AI-augmented SIEM… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-cybersecurity-en.cybersecurity-qa-v2
Cybersecurity Q&A Dataset v2 — 2.6M Examples
A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics.
2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies.
Statistics
Source
Examples
Description
NIST NVD CVE Database
~1,954,225
All CVEs (2002–2025): overview, severity, detection, remediation
AlicanKiraz0/All-CVE-Records-Training-Dataset
~297,441
Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/cybersecurity-qa-v2.cybersecurity-reasoning-cot-v1
🛡️ Expert Cybersecurity Reasoning Dataset (CoT)
This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic.
💎 Key Highlights
Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets.
Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.cybersecurity-controls-instructions
Cybersecurity Controls Instructions
Security control, incident response and risk management guidance from NIST Special Publications, turned into instruction-following examples.
Splits
split
rows
source documents
train
13,106
56
validation
4,840
18
test
5,697
18
Splits are held out by source document. Every chunk yields several
instruction rows, so a random row-level split would place the same passage in
train and test; whole documents are held… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/cybersecurity-controls-instructions.ai-cybersecurity-fr
IA en Cybersecurite Offensive et Defensive - Dataset Francais
Description
Dataset complet et bilingue couvrant l'utilisation de l'Intelligence Artificielle en cybersecurite, tant du cote offensif (attaquants) que defensif (defenseurs). Ce dataset est la version francaise.
Articles couverts
Ce dataset synthetise les connaissances des articles suivants :
IA Offensive : Comment les Attaquants Utilisent les LLM - Techniques d'attaque basees sur les LLM
Detection… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-cybersecurity-fr.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/Soban1234/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive
Trendyol Cybersecurity Defense Instruction-Tuning Dataset (v2.0)
🚀 TL;DR
53,202 meticulously curated system/user/assistant instruction-tuning examples covering 200+ specialized cybersecurity domains. Built by the Trendyol Security Team for training state-of-the-art defensive security AI assistants. Expanded from 21K to 53K rows with comprehensive coverage of modern security challenges including cloud-native threats, AI/ML security, quantum… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/Trendyol-Cybersecurity-Instruction-Tuning-Dataset-archive.deer_sec-japanese-cybersecurity-chatml-v2
deer_sec-japanese-cybersecurity-chatml-v2.0
📊 Dataset Details
Total Rows: 83,562 件
File Size: 約 492.5 MB
Format: JSONL (ChatML形式)
Language: 日本語 (Japanese)
概要 (Overview)
本データセットは、高度なサイバー防御と脅威インテリジェンスに特化したインストラクション・チューニング用のデータセットです。AlicanKiraz0様によって公開された Cybersecurity-Dataset-Fenrir-v2.0 を元に構築されています。
元の膨大な英語データの中から約 99.5% (83,562件) を抽出し、翻訳特化モデルである translategemma:12b を用いて高品質な日本語へ翻訳しました。その後、LLMのファインチューニング(LoRA等)にそのまま利用できるよう、厳格なデータクレンジングと整形を行っています。
特徴… See the full description on the dataset page: https://huggingface.co/datasets/deer-sec/deer_sec-japanese-cybersecurity-chatml-v2.DoD-Instruction-8500-01-Cybersecurity
🛡️ DoD Cybersecurity Question-Answer Dataset
Source: DoD Instruction 8500.01
Source Effective Date: March 14, 2014
Change Incorporated: Change 1, effective October 7, 2019
Source Organization: Office of the DoD Chief Information Officer
Source Ownership: United States Department of Defense
📋 Overview
Dataset Summary
The DoD Cybersecurity Question-Answer Dataset is a structured, document-grounded natural-language dataset derived from DoD… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/DoD-Instruction-8500-01-Cybersecurity.
