datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cybersecurity-qa-v2
Cybersecurity Q&A Dataset v2 — 2.6M Examples
A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics.
2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies.
Statistics
Source
Examples
Description
NIST NVD CVE Database
~1,954,225
All CVEs (2002–2025): overview, severity, detection, remediation
AlicanKiraz0/All-CVE-Records-Training-Dataset
~297,441
Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/rezaduty/cybersecurity-qa-v2.nist-cybersecurity-training
NIST Cybersecurity Training Dataset v1.1
The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs
Version 1.1 Highlights
What's New in v1.1:
✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents
✅ Fixed 6,150 broken DOI links via format normalization
✅ Removed 202 malformed DOIs (double URL prefixes)
✅ Validated and fixed 124,946 total links
✅ Cataloged 72,698 broken links for future recovery
✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.cybersecurity_32k_instruction_input_output
Dataset Card
The dataset Q&As are focused on identification of cyber threats, and text classification under the NIST taxonomy and ITC EBA IT risk classes
Dataset Details
Dataset Description
This dataset includes a mix of public reports and news and aims to be used for cyber security risk model training.
It includes 32k examples with instruction, input and output. The latter is the output from GPT.
Curated by: [Vanessa Lopes]
Language [EN]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vanessasml/cybersecurity_32k_instruction_input_output.cybersecurity-master-dataset
Cybersecurity Master Dataset
Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions.
cybersecurity-instruction-datasetcyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.cybersecurity-sft-datasetcybersecurity-soc-threat-hunting-sft-dpo-2026
🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects.
📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format)
A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles.
Dataset Description
This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.Cybersecurity-Dataset-v1
Cybersecurity Defense Training Dataset
Dataset Description
This dataset contains 2,500 high-quality instruction-response pairs focused on defensive cybersecurity education. The dataset is designed to train AI models to provide accurate, detailed, and ethically-aligned guidance on information security principles while refusing to assist with malicious activities.
Dataset Summary
Language: English
License: Apache 2.0
Format: Parquet
Size: 2,500 rows
Domain:… See the full description on the dataset page: https://huggingface.co/datasets/AlicanKiraz0/Cybersecurity-Dataset-v1.HARMLESS_Synthetic_Injected_PDFs_EDA
Injected PDFs - EDA and Evaluation Corpus
This repository holds the exploratory data analysis for a project on detecting harmless-but-real
attack payloads injected into PDF files, together with the dataset that analysis produced.
The project has two halves, both in the notebook Final_project_V7_EDA.ipynb:
Question
Input
Part 1
Is our synthetic corpus a stand-in for real malware, or is it something else?
The published CIC feature table (11,126 x 34)
Part 2
Is our… See the full description on the dataset page: https://huggingface.co/datasets/Cyber-security-final-project/HARMLESS_Synthetic_Injected_PDFs_EDA.cybersecurity-questionaire
Dataset Card for cybersecurity-questionaire
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/MichaelPrimez/cybersecurity-questionaire/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MichaelPrimez/cybersecurity-questionaire.kovidore-v2-cybersecurity-beirKoViDoRe v2 : Cybersecurity
This dataset, Cybersecurity, is a corpus of technical reports on cyber threat trends and security incident responses in Korea, intended for complex-document understanding tasks. It is one of the 4 corpora comprising the KoViDoRe v2 Benchmark.
Links
Github: https://github.com/whybe-choi/kovidore-benchmark
Collection: https://huggingface.co/collections/whybe-choi/kovidore-benchmark-beir-v2
Data Generation Pipeline:… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/kovidore-v2-cybersecurity-beir.cybersecurity-sft-dataset
Cybersecurity SFT Dataset
A curated dataset for training cybersecurity-focused code models with structured JSON output capability.
Dataset Composition
Source
Count
Percentage
Description
CVE Records
10,000
50.0%
Multi-turn CVE vulnerability analysis
OpenCodeReasoning (NVIDIA)
5,000
25.0%
Chain-of-thought code reasoning
Code-Feedback
5,000
25.0%
Multi-turn code debugging and refinement
Synthetic Security (JSON)
5
<0.1%
JSON-structured CVE, MITRE ATT&CK… See the full description on the dataset page: https://huggingface.co/datasets/moro72842/cybersecurity-sft-dataset.Trendyol-Cybersecurity-Instruction-Tuning-Dataset-Convertedrag_eval_cybersecuritycybersecurity-sft-datasetcybersecurity-nercybersecurity-qa-datasetai-cybersecurity-en
AI in Offensive and Defensive Cybersecurity - English Dataset
Description
Comprehensive bilingual dataset covering the use of Artificial Intelligence in cybersecurity, from both the offensive (attackers) and defensive (defenders) perspectives. This is the English version.
Articles Covered
This dataset synthesizes knowledge from the following articles:
Offensive AI: How Attackers Use LLMs - LLM-based attack techniques
AI Threat Detection - AI-augmented SIEM… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-cybersecurity-en.CyberSecurity-bigcybersecurity-reasoning-cot-v1
🛡️ Expert Cybersecurity Reasoning Dataset (CoT)
This dataset contains 89 high-fidelity, expert-verified reasoning records focusing on complex cybersecurity attack vectors. It is designed specifically for fine-tuning Large Language Models (LLMs) on sophisticated security analysis and threat logic.
💎 Key Highlights
Niche Rarity 1.0: Covers rare and emerging threats with zero prior representation in open-source datasets.
Advanced Vectors: Includes detailed reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/expertdata-factory/cybersecurity-reasoning-cot-v1.cybersecurity-controls-instructions
Cybersecurity Controls Instructions
Security control, incident response and risk management guidance from NIST Special Publications, turned into instruction-following examples.
Splits
split
rows
source documents
train
13,106
56
validation
4,840
18
test
5,697
18
Splits are held out by source document. Every chunk yields several
instruction rows, so a random row-level split would place the same passage in
train and test; whole documents are held… See the full description on the dataset page: https://huggingface.co/datasets/nuhmanpk/cybersecurity-controls-instructions.cybersecurity-qa-v2
Cybersecurity Q&A Dataset v2 — 2.6M Examples
A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics.
2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies.
Statistics
Source
Examples
Description
NIST NVD CVE Database
~1,954,225
All CVEs (2002–2025): overview, severity, detection, remediation
AlicanKiraz0/All-CVE-Records-Training-Dataset
~297,441
Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/Nobody05/cybersecurity-qa-v2.cyber-security-events-full
cyber-security-events-full
Dataset Description
This dataset contains cybersecurity events collected from honeypot infrastructure.
The data has been processed and feature-engineered for machine learning applications in threat detection and security analytics.
Feature Categories
Network Features
Connection flow statistics (bytes, packets, duration)
Protocol-specific metrics
Geographic information
IP reputation data
Behavioral Features… See the full description on the dataset page: https://huggingface.co/datasets/pyToshka/cyber-security-events-full.kovidore-v2-cybersecurity-mteb
KoVidore2CybersecurityRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions. This dataset, Cybersecurity, is a corpus of technical reports on cyber threat trends and security incident responses in Korea, intended for complex-document understanding tasks.
Task category
t2i
Domains
Social
Reference
https://github.com/whybe-choi/kovidore-data-generator
Source datasets:
whybe-choi/kovidore-v2-cybersecurity-beir… See the full description on the dataset page: https://huggingface.co/datasets/whybe-choi/kovidore-v2-cybersecurity-mteb.ai-cybersecurity-fr
IA en Cybersecurite Offensive et Defensive - Dataset Francais
Description
Dataset complet et bilingue couvrant l'utilisation de l'Intelligence Artificielle en cybersecurite, tant du cote offensif (attaquants) que defensif (defenseurs). Ce dataset est la version francaise.
Articles couverts
Ce dataset synthetise les connaissances des articles suivants :
IA Offensive : Comment les Attaquants Utilisent les LLM - Techniques d'attaque basees sur les LLM
Detection… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ai-cybersecurity-fr.cybersecurity-qa-v2
Cybersecurity Q&A Dataset v2 — 2.6M Examples
A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics.
2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies.
Statistics
Source
Examples
Description
NIST NVD CVE Database
~1,954,225
All CVEs (2002–2025): overview, severity, detection, remediation
AlicanKiraz0/All-CVE-Records-Training-Dataset
~297,441
Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/cybersecurity-qa-v2.africa-synth-telecom-cybersecurity-incident-logs-nigeria
Africa Synth Telecom Cybersecurity Incident Logs Nigeria | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-synth-telecom-cybersecurity-incident-logs-nigeria.cybersecurity-ner
