datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/reloading0101/threat-intelligence-dataset.cybersecurity-soc-threat-hunting-sft-dpo-2026
🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects.
📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.task322_jigsaw_classification_threat
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task322_jigsaw_classification_threat
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task322_jigsaw_classification_threat.task1722_civil_comments_threat_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1722_civil_comments_threat_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1722_civil_comments_threat_classification.threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/tonygarg/threat-intelligence-dataset.threat-intelligence-dataset-archive
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/ChipHolmes/threat-intelligence-dataset-archive.threat-intelligence
Comprehensive Threat Intelligence Dataset
Dataset Description
This comprehensive bilingual (French/English) threat intelligence dataset contains detailed information about Indicators of Compromise (IoCs), Tactics, Techniques, and Procedures (TTPs), APT groups, malware families, and threat hunting queries. The dataset is designed for training security analysts, threat hunters, and AI models focused on cybersecurity.
Dataset Summary
Languages: English (en)… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/threat-intelligence.threat-intelligence
THREAT_INTELLIGENCE
A preference dataset for THREAT_INTELLIGENCE, harvested from real, human-labelled sources and curated by an automated harvesting harness with an LLM quality gate.
Format
Standard preference / DPO schema — each row:
column
meaning
prompt
the request (originally prompt)
chosen
the human-preferred response
rejected
a worse response to the same prompt
source
the dataset/URL the row was harvested from
Splits… See the full description on the dataset page: https://huggingface.co/datasets/316usman/threat-intelligence.cyber-threat-intelligence
Cyber Threat Intelligence Dataset
A comprehensive cybersecurity dataset combining CVE vulnerability data, MITRE ATT&CK techniques, and CISA Known Exploited Vulnerabilities — structured for AI/ML training and security research.
Author: Soham DahivalkarLicense: MITCreated: 2026
Dataset Description
This dataset provides structured cybersecurity intelligence data collected from three authoritative public sources:
NVD (National Vulnerability Database) — CVE… See the full description on the dataset page: https://huggingface.co/datasets/Shomi28/cyber-threat-intelligence.Modern_Cyber_Threat_Simulation_DatasetModern Cyber Threat Simulation Dataset
Overview
The Modern Cyber Threat Simulation Dataset is a comprehensive collection of 200 simulated cyber threats, vulnerabilities, and exploits tailored for 2025's advanced technological landscape. Covering AI/ML, Blockchain, Cloud, and IoT domains, this dataset provides vulnerable code/configurations, fuzzing-based exploit scripts, mitigations, and AI training prompts to support cybersecurity research, red teaming, and defensive tool development. Each… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Modern_Cyber_Threat_Simulation_Dataset.cyber-threat-intelligence-custom-datamirror-threat-intelligence-dataset
Cyber Threat Intelligence Dataset for LLM Fine-Tuning
An instruction-tuning dataset for teaching language models to do cyber threat intelligence work: reading a CVE and explaining what the risk actually is, profiling a threat actor from its ATT&CK techniques, turning a Sigma rule into alert-triage guidance, mapping a campaign's kill chain, writing detection logic for a technique, and so on.
The four splits live under data/; the rest of this card documents how the set was built… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-threat-intelligence-dataset.cyber-threat-intelligence-custom-data-ko
Dataset Card for "cyber-threat-intelligence-custom-data"
Translated swaption2009/cyber-threat-intelligence-custom-data using nayohan/llama3-instrucTrans-enko-8b.
Modern_Cyber_Threat_Simulation_DatasetModern Cyber Threat Simulation Dataset
Overview
The Modern Cyber Threat Simulation Dataset is a comprehensive collection of 200 simulated cyber threats, vulnerabilities, and exploits tailored for 2025's advanced technological landscape. Covering AI/ML, Blockchain, Cloud, and IoT domains, this dataset provides vulnerable code/configurations, fuzzing-based exploit scripts, mitigations, and AI training prompts to support cybersecurity research, red teaming, and defensive tool development. Each… See the full description on the dataset page: https://huggingface.co/datasets/cyrusvirus949/Modern_Cyber_Threat_Simulation_Dataset.cyber-threat-intel-dataset
🛡️ CyberThreat Intel Dataset
A dataset containing 471 instruction-tuning pairs designed to teach LLMs how to generate automated, structured cybersecurity threat intelligence reports from raw CVE vulnerability data.
💻 Project GitHub: vanshkamra12/CyberThreat-Intel-LLM
🧠 Fine-Tuned Model: vanshkamra12/CyberSecurity-Model
Dataset Structure
Each line is a JSON object with three fields (Alpaca format):
instruction: The prompt asking the model to act as a threat… See the full description on the dataset page: https://huggingface.co/datasets/vanshkamra12/cyber-threat-intel-dataset.cybersecurity-threat-intelligence
🛡️ Cybersecurity threat intelligence dataset (FREE SAMPLE)
🚀 Looking for the full dataset? https://deniks.gumroad.com/l/svgbfp
📌 Overview
This repository contains a free preview (100 high-quality records) of a professionally curated instruction-tuning dataset. It features cleaned cybersecurity threat reports, vulnerability disclosures, and attack summaries formatted explicitly for training Large Language Models (LLMs) on InfoSec summarization and analysis.
🎯… See the full description on the dataset page: https://huggingface.co/datasets/deniks315/cybersecurity-threat-intelligence.Modern_Cyber_Threat_Simulation_DatasetModern Cyber Threat Simulation Dataset
Overview
The Modern Cyber Threat Simulation Dataset is a comprehensive collection of 200 simulated cyber threats, vulnerabilities, and exploits tailored for 2025's advanced technological landscape. Covering AI/ML, Blockchain, Cloud, and IoT domains, this dataset provides vulnerable code/configurations, fuzzing-based exploit scripts, mitigations, and AI training prompts to support cybersecurity research, red teaming, and defensive tool development. Each… See the full description on the dataset page: https://huggingface.co/datasets/SujaaViswanathan/Modern_Cyber_Threat_Simulation_Dataset.
