datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
crello
Dataset Card for Crello
Dataset Description
The Crello dataset is a collection of raster graphic designs originally compiled for the study of vector graphic documents. It contains document meta-data such as canvas size and pre-rendered elements such as images or text boxes. The original templates were collected from crello.com (now create.vista.com) and converted to a low-resolution format suitable for machine learning analysis. More recently, it has been used for… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/crello.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/SHSLab/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.cybersecurity-qa-v2
Cybersecurity Q&A Dataset v2 — 2.6M Examples
A large-scale cybersecurity Q&A dataset for fine-tuning LLMs on security topics.
2,621,468 examples covering vulnerabilities, attack techniques, weaknesses, and defensive strategies.
Statistics
Source
Examples
Description
NIST NVD CVE Database
~1,954,225
All CVEs (2002–2025): overview, severity, detection, remediation
AlicanKiraz0/All-CVE-Records-Training-Dataset
~297,441
Detailed CVE analysis with markdown… See the full description on the dataset page: https://huggingface.co/datasets/rezaduty/cybersecurity-qa-v2.nist-cybersecurity-training
NIST Cybersecurity Training Dataset v1.1
The largest open-source NIST cybersecurity training dataset for fine-tuning LLMs
Version 1.1 Highlights
What's New in v1.1:
✅ Added CSWP (Cybersecurity White Papers) series - 23 new documents
✅ Fixed 6,150 broken DOI links via format normalization
✅ Removed 202 malformed DOIs (double URL prefixes)
✅ Validated and fixed 124,946 total links
✅ Cataloged 72,698 broken links for future recovery
✅ 0 broken link markers remaining in… See the full description on the dataset page: https://huggingface.co/datasets/ethanolivertroy/nist-cybersecurity-training.Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
🌌 Omni-Frontier Distillation SFT
The Definitive Evolution of Open-Source Distillation & Human-Crafted Expertise
Repository: Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection
"The most comprehensive multi‑domain SFT corpus ever assembled — fusing 6.86 million cleaned distillation samples with 9.14 million human‑crafted expert examples across medical, cybersecurity, chemical, robotics, humanities, and more. 16 million… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-Coding-Med-dataset-collection.cybersecurity_32k_instruction_input_output
Dataset Card
The dataset Q&As are focused on identification of cyber threats, and text classification under the NIST taxonomy and ITC EBA IT risk classes
Dataset Details
Dataset Description
This dataset includes a mix of public reports and news and aims to be used for cyber security risk model training.
It includes 32k examples with instruction, input and output. The latter is the output from GPT.
Curated by: [Vanessa Lopes]
Language [EN]
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Vanessasml/cybersecurity_32k_instruction_input_output.OTR
OTR: Overlay Text Removal Dataset
OTR (Overlay Text Removal) is a synthetic benchmark dataset designed to advance research of text removal from images.It features complex, object-aware text overlays with clean, artifact-free ground truth images, enabling more challenging evaluation scenarios beyond traditional scene text datasets.
📦 Dataset Overview
Subset
Source Dataset
Content Type
# SamplesNotes
OTR-easy (test set)
MS-COCO
Simple backgrounds (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/OTR.cybersecurity-master-dataset
Cybersecurity Master Dataset
Unified and deduplicated cybersecurity SFT dataset containing CTF solutions, CVE analyses, vulnerability patches, and Python coding instructions.
cybergym-tasks
CyberGym task ID splits
Task ID splits used in our work on CyberGym vulnerability-reproduction
benchmarks. Each row is a single task identifier (no inputs / no outputs);
this dataset is intended as a task-list manifest for downstream evaluation
scripts that fetch the actual task workspaces from the CyberGym distribution.
Configs
Config
Split
Rows
Description
full
train
1507
Every task in the CyberGym Level-1 release
train
train
300
Training pool used in our… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/cybergym-tasks.camera
Dataset Card for CAMERA📷:
Table of Contents:
Dataset Card for Camera
Table of Contents
Dataset Details
Dataset Description
Dataset Sources
Uses
Direct Use
Dataset Information
Data Example
Dataset Structure
Citation
Dataset Details
Dataset Description
CAMERA (CyberAgent Multimodal Evaluation for Ad Text GeneRAtion) is the Japanese ad text generation dataset, which comprises actual data sourced from Japanese search ads and incorporates… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/camera.cybermetric-10000
Dataset Card for "cybermetric-10000"
More Information needed
Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2
🧬 Omni-Frontier Collection
Cybersecurity · Coding · Math · Science · RSI Reasoning — one unified SFT package
A unified, deduplicated, fully-browsable distillation & SFT corpus — every row real, every row visible.
📖 Jump to
What's inside · 🔁 Aggregation audit · 🛡 Cybersecurity · 💻 Coding · 🏭 Distillation deep-dive · 🔁 RSI · 🧮 Math/Science/More · 🎓 Training guide · 🔎 Browsing · 🧹 Quality · 🗺 Roadmap · 📄 License… See the full description on the dataset page: https://huggingface.co/datasets/Manusagents/Omni-Frontier-Distillation-SFT-Cyber-security-Coding-dataset-collection-v2.cybersecurity-instruction-datasetcyber-security
Cybersecurity Instruction-Tuning Dataset
A large, cleaned, multi-domain cybersecurity chat dataset for LLM finetuning,
built from 198 distinct sources spanning offensive security, blue-team
operations, vulnerability intelligence, cloud/AWS security, malware analysis,
digital forensics, and more. Every record is normalized to the standard
messages chat format and deduplicated at both file and record level.
⚠️ Research use only. This dataset is provided exclusively for… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/cyber-security.cybermetric_10000_v1
CyberMetric: Cybersecurity Multiple Choice Questions (10000_V1)
Dataset Description
This dataset contains 10,180 cybersecurity multiple choice questions from the CyberMetric benchmark suite. It focuses on cybersecurity knowledge evaluation, particularly covering topics related to:
🔐 Cryptography: Random Bit Generators, Key Derivation Functions, Encryption
💳 PCI DSS: Payment Card Industry Data Security Standards
🛡️ Security Controls: Access controls, privilege… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/cybermetric_10000_v1.wmdp-cyber-forget-corpuscybersecurity-sft-datasetCyberSecEval
CyberSecEval
The dataset source can be found here.
(CyberSecEval2 Version)
Abstract
Large language models (LLMs) introduce new security risks, but there are few comprehensive evaluation suites to measure and reduce these risks. We present CYBERSECEVAL 2, a novel benchmark to quantify LLM security risks and capabilities. We introduce two new areas for testing: prompt injection and code interpreter abuse. We evaluated multiple state of the art (SOTA) LLMs, including GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CyberSecEval.cyber-sft-agent-qwen38
Cyber-SFT-Agent (Qwen3.8-27B Native) — v3.1 production (8.4K)
Agentic tool-calling SFT dataset for Qwen3.8-27B (abliterated or base).
Teaches the model to call real tools — masscan → nmap → curl → whatweb → nikto → ffuf → gobuster → dirb → wpscan → hydra → sqlmap → john → hashcat → searchsploit → metasploit → linpeas → winpeas → crackmapexec → enum4linux → responder → impacket → netcat — read their output, re-plan from the evidence,
and synthesize a ranked attack. Companion to… See the full description on the dataset page: https://huggingface.co/datasets/hotdogs/cyber-sft-agent-qwen38.Benchmarks_CyberSec_RedSageMCQ
Dataset Card for RedSage-MCQ
Dataset Summary
RedSage-MCQ is a large-scale, high-quality multiple-choice question (MCQ) benchmark designed to evaluate the cybersecurity knowledge, skills, and tool proficiency of Large Language Models (LLMs). It is a component of the RedSage-Bench suite introduced in the paper "RedSage: A Cybersecurity Generalist LLM".
The dataset comprises 30,000 questions derived from RedSage-Seed, a curated collection of authoritative… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_RedSageMCQ.cybersecurity-soc-threat-hunting-sft-dpo-2026
🛡️ Enterprise Cybersecurity AI, SOC Tier-3 & Threat Hunting SFT/DPO Dataset (2026)
High-precision multi-turn instruction tuning and preference optimization dataset with step-by-step SOC Tier-3 Chain-of-Thought (<thought>) kill-chain diagnostic trees for fine-tuning LLMs (Llama-3.3, Qwen-2.5-Coder, DeepSeek-R1-Distill, Mistral) into Senior SOC Threat Hunters, Incident Responders, and Red-Team Defense Architects.
📊 Dataset Architecture & Highlights… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/cybersecurity-soc-threat-hunting-sft-dpo-2026.cyberQA
Dataset Card for "cyberQA"
More Information needed
PhishingEmailDetectionv2.0
Phishing Email Detection Dataset
A comprehensive dataset combining email messages and URLs for phishing detection.
Dataset Overview
Quick Facts
Task Type: Multi-class Classification
Languages: English
Total Samples: 200,000 entries
Size Split:
Email samples: 22,644
URL samples: 177,356
Label Distribution: Four classes (0, 1, 2, 3)
Format: Two columns - content and labels
Dataset Structure
Features
{
'content':… See the full description on the dataset page: https://huggingface.co/datasets/cybersectony/PhishingEmailDetectionv2.0.opencole
Dataset Card for OpenCOLE
The dataset contains additional synthetically generated annotations for each sample in crello v4.0.0 dataset.
Dataset Details
Dataset Description
Each instance has several attributes. id indicates the template id of crello v4.0.0 dataset.
{
'id': '5a1449f0d8141396fe997112',
'intention': "Create a Facebook post for a New Year celebration event at Templo Historic House. The post should invite viewers to join the celebration on… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/opencole.pick_and_place_unimanualcybermetric_2000_v1
CyberMetric: Cybersecurity Multiple Choice Questions (2000_V1)
Dataset Description
This dataset contains 2,000 cybersecurity multiple choice questions from the CyberMetric benchmark suite. It focuses on cybersecurity knowledge evaluation, particularly covering topics related to:
🔐 Cryptography: Random Bit Generators, Key Derivation Functions, Encryption
💳 PCI DSS: Payment Card Industry Data Security Standards
🛡️ Security Controls: Access controls, privilege management… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/cybermetric_2000_v1.cybermetric_500_v1
CyberMetric: Cybersecurity Multiple Choice Questions (500_V1)
Dataset Description
This dataset contains 500 cybersecurity multiple choice questions from the CyberMetric benchmark suite. It focuses on cybersecurity knowledge evaluation, particularly covering topics related to:
🔐 Cryptography: Random Bit Generators, Key Derivation Functions, Encryption
💳 PCI DSS: Payment Card Industry Data Security Standards
🛡️ Security Controls: Access controls, privilege management… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/cybermetric_500_v1.Benchmarks_CyberSec_SecBench
Dataset Card for SecBench (RISys-Lab Mirror)
⚠️ Disclaimer: > This repository is a mirror/re-host of the original SecBench dataset.RISys-Lab is not the author of this dataset. We are hosting this copy in Parquet format to ensure seamless integration and stability for our internal evaluation pipelines. All credit and rights belong to the original authors listed below.
Repository Intent
This Hugging Face dataset is a re-host of the original SecBench. It has been… See the full description on the dataset page: https://huggingface.co/datasets/RISys-Lab/Benchmarks_CyberSec_SecBench.Trendyol-Cybersecurity-Instruction-Tuning-Dataset
Trendyol Cybersecurity Instruction Tuning Dataset (GPT Format)
A conversational dataset in GPT/OpenAI messages format, converted from Trendyol/Trendyol-Cybersecurity-Instruction-Tuning-Dataset. Designed for training language models in advanced cyber-defense and security principles.
Dataset Description
This dataset contains 53,201 high-quality instruction-tuning examples focused on cybersecurity, converted to the standard GPT conversation format (messages) for… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/Trendyol-Cybersecurity-Instruction-Tuning-Dataset.cyber-sft-qa-qwen38
Cyber-SFT-QA (Qwen3.8-27B Native)
Cybersecurity SFT dataset for Qwen3.8-27B (abliterated or base).
Built from Trendyol/Cybersecurity-Instruction (42,807) + AlicanKiraz0/Cybersecurity-Dataset-Fenrir-v2.1 (29,711) —
merged, dedup'd, and converted to the exact thinking format Qwen3.8-27B expects.
Why this format (the important part)
Qwen3.8-27B's chat template wraps the assistant's reasoning in special tags.
Verified by rendering the real chat_template.jinja + the… See the full description on the dataset page: https://huggingface.co/datasets/hotdogs/cyber-sft-qa-qwen38.
