datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cve-sft-v5
CVE SFT Dataset v5
CVE SFT Dataset v5 is a structured instruction-following dataset for fine-tuning language models on cybersecurity vulnerability analysis. Built by Auren Research, it combines authoritative vulnerability metadata from the NIST National Vulnerability Database (NVD) with five generated fields that teach models to explain, reason about, and remediate real-world CVEs — including side-by-side vulnerable vs. safe code examples.Unlike most security… See the full description on the dataset page: https://huggingface.co/datasets/auren-research/cve-sft-v5.cve-top100-en
CVE Top 100 Dataset - English Edition
A comprehensive bilingual dataset of the 100 most critical and exploited CVEs from 2014 to 2024, with educational Q&A on vulnerability management.
Description
This dataset contains:
100 critical CVEs documented with precise technical details
50 French questions-answers on CVEs and vulnerability management
50 English questions-answers for bilingual learning
Detailed statistics on vulnerability trends
Vulnerability Types… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/cve-top100-en.cve-analysis
CVE & Vulnerability Analysis Dataset
A comprehensive vulnerability analysis and CVE research dataset. Each row is a detailed security analysis covering root cause, exploitation methodology, detection rules (Sigma/Splunk/Suricata), CVSS v3.1 scoring, MITRE ATT&CK mapping, and remediation guidance — verified by the same model in an independent review pass.
Overview
This dataset contains 9,999 structured vulnerability analyses across 20 security domains. Unlike simple… See the full description on the dataset page: https://huggingface.co/datasets/sh111111111111111/cve-analysis.cve-top100-fr
Dataset CVE Top 100 - Édition Française
Un dataset bilingue complet des 100 CVEs les plus critiques et exploitées de 2014 à 2024, avec Q&A éducatif sur la gestion des vulnérabilités.
Description
Ce dataset contient:
100 CVEs critiques documentées avec détails techniques précis
50 questions-réponses en français sur les CVEs et la gestion des vulnérabilités
50 questions-réponses en anglais pour un apprentissage bilingue
Statistiques détaillées sur les tendances des… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/cve-top100-fr.CVEs
CVEs — a full-coverage CVE chat dataset
1,625,017 chat conversations covering all 361,190 usable CVEs (1999–2026), built for fine-tuning cybersecurity assistants. Every known CVE in the official CVE List with severity enrichment from NVD (via the fkie-cad community feeds), rendered as English user/assistant conversations with varied phrasings, honest handling of missing data, and a per-CVE 99/1 train/validation split with zero leakage.
The schema matches oi-uae/cyber-security… See the full description on the dataset page: https://huggingface.co/datasets/oi-uae/CVEs.cve-intelligence-stream
CVE Security Intelligence Stream
Enterprise-grade CVE feed for Threat Intelligence and Risk Management teams.
Incremental CVE records from NVD API 2.0 enriched with CISA Known Exploited Vulnerabilities (KEV) metadata. Normalized for LLM SFT/RAG.
B2B Value Proposition
This dataset is built for security vendors, MSSPs, and enterprise SOC/GRC teams who need:
Use case
What you get
Threat Intel enrichment
CVSS severity, affected software, exploit… See the full description on the dataset page: https://huggingface.co/datasets/FXBIA/cve-intelligence-stream.cvetop100en
CVE Top 100 Dataset - English Edition
A comprehensive bilingual dataset of the 100 most critical and exploited CVEs from 2014 to 2024, with educational Q&A on vulnerability management.
Description
This dataset contains:
100 critical CVEs documented with precise technical details
50 French questions-answers on CVEs and vulnerability management
50 English questions-answers for bilingual learning
Detailed statistics on vulnerability trends
Vulnerability Types… See the full description on the dataset page: https://huggingface.co/datasets/Whoisjutanlee/cvetop100en.cve_dataset
TanDev CVE Dataset (NVD SlimPajama Corpus)
A SlimPajama‑style corpus of CVE entries derived from the NIST NVD (CVE 2.0) data feeds (2002→present). Each row is a cleaned, single‑document text representation of a CVE with structured metadata for CVSS, CWE(s), timestamps, and a canonical NVD link—ready for pretraining/continued‑pretraining, RAG, retrieval/evaluation, and downstream classifiers.
Token count: ~108.2M tokens.
License: "TanDev Proprietary License — All Rights Reserved"
⚠️… See the full description on the dataset page: https://huggingface.co/datasets/tandevllc/cve_dataset.
