datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet with… See the full description on the dataset page: https://huggingface.co/datasets/ayshajavd/code-security-vulnerability-dataset.Code-Vulnerability-FineTune
🔐 Code Vulnerability FineTome — CWE-Enriched Conversation Dataset
📌 Overview
This dataset converts raw security-labeled C/C++ code samples into instruction-following conversation pairs suitable for fine-tuning large language models (LLMs) on software vulnerability detection and analysis.
It is built by preprocessing and transforming the ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset (330k rows, sourced from DiverseVul + MITRE CWE enrichment) into… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune.Code-Vulnerability-Balanced
Code Vulnerability Balanced — CWE-Enriched Conversation Dataset
📌 Overview
This dataset is a balanced and shuffled version of
ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune,
which itself was derived from the original
ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset
(330k rows, sourced from DiverseVul + MITRE CWE enrichment).
The original fine-tuning dataset was imbalanced — the number of Vulnerable and Safe
samples were not equal — and the samples were not… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced.Code_Vulnerability_Dataset
🔐 Code Vulnerability Dataset (CWE-Enriched)
📌 Overview
This dataset is built from the bstee615/diversevul dataset and enhanced with structured vulnerability intelligence from the MITRE Common Weakness Enumeration (CWE) database.
It provides a rich, machine-readable representation of software vulnerabilities, mapping raw vulnerable code samples to standardized CWE classifications.
The dataset is designed for research and development in:
Vulnerability detection models… See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset.code_vulnerability_pythonCode_Vulnerability_Security_DPOcode-vulnerability-dpoCWE-Code_Vulnerability_Security_DPOcode_vulnerability_javaGeneric-Code-Vulnerability-Backdoorcode-vulnerability-evalmirror-code-security-vulnerability-dataset
Code Security Vulnerability Dataset
A curated multi-language dataset of 175,419 code samples labeled with 31 vulnerability classes (30 CWEs + safe) for training multi-label code vulnerability detection models. Labels are mapped to OWASP Top 10 2021 categories.
Dataset Details
Property
Value
Total Samples
175,419
Train / Val / Test
140,335 / 17,542 / 17,542
Languages
C, C++, Python, JavaScript, Java, PHP, Go
Labels
31 (multi-label)
Format
Parquet… See the full description on the dataset page: https://huggingface.co/datasets/alucent/mirror-code-security-vulnerability-dataset.cpp-vulnerability-dataset
