ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced
Code Vulnerability Balanced β CWE-Enriched Conversation Dataset π Overview This dataset is a balanced and shuffled version of ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune, which itself was derived from the original ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset (330k rows, sourced from DiverseVul + MITRE CWE enrichment). The original fine-tuning dataset was imbalanced β the number of Vulnerable and Safe samples were not equal β and the samplesβ¦ See the full description on the dataset page: https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced.
Code Vulnerability Balanced β CWE-Enriched Conversation Dataset
π Overview
This dataset is a balanced and shuffled version of ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune, which itself was derived from the original ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset (330k rows, sourced from DiverseVul + MITRE CWE enrichment).
The original fine-tuning dataset was imbalanced β the number of Vulnerable and Safe samples were not equal β and the samples were not shuffled. This dataset addresses both issues:
- β Balanced β Equal number of Vulnerable and Safe (patched) samples
- β Shuffled β Samples are randomly shuffled to prevent ordering bias during training
- β ShareGPT / FineTome format β Ready for fine-tuning with Unsloth, TRL, and similar frameworks
π― Use Cases
- Fine-tuning LLMs for security code review
- Training vulnerability detection models
- Building code-aware security assistants
- Research in automated static analysis and secure coding
π Dataset Statistics
π What Changed From the Source Dataset
ποΈ Data Format
Each row follows the ShareGPT conversation format with two turns:
{
"conversations": [
{
"from": "human",
"value": "Analyze the following code snippet and identify any security vulnerabilities...\n\n```c\n<source code>\n```"
},
{
"from": "gpt",
"value": "## Security Vulnerability Analysis\n\nβ οΈ This code sample is marked as **Vulnerable**.\n\n### π Vulnerability Classification\n- **CWE ID**: CWE-787\n- **Type**: Out-of-bounds Write\n- **Severity**: High\n..."
}
],
"source": "code_vulnerability_cwe",
"score": 4.8
}Fields
π Preprocessing Pipeline
Step 1 β Load
Load the source dataset from Hugging Face Hub (ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune).
Step 2 β Separate by Label
Split all rows into two groups:
- Vulnerable samples
- Safe (patched) samples
Step 3 β Balance
Undersample the larger group so that both groups have an equal number of rows, resulting in a 50/50 class distribution.
Step 4 β Shuffle
Combine the two balanced groups and shuffle the entire dataset with a fixed random seed (42) for reproducibility.
Step 5 β Split & Push
- 90% / 10% train-test split (random seed 42)
- Pushed to Hugging Face Hub in Parquet format
π CWE Categories Covered
The dataset covers a wide range of Common Weakness Enumeration types including:
π§ͺ Sample Example
Human Turn
Analyze the following code snippet and identify any security vulnerabilities.
Provide a detailed explanation of the vulnerability type, its severity,
potential impact, and the CWE classification.
static boolean ReadICCProfile(j_decompress_ptr jpeg_info) {
char magick[12];
ErrorManager *error_manager;
ExceptionInfo *exception;
Image *image;
...
}Assistant Turn
## Security Vulnerability Analysis
β οΈ This code sample is marked as **Vulnerable**.
### π Vulnerability Classification
- **CWE ID**: CWE-416
- **Type**: Use After Free
- **Severity**: High
- **Category**: Memory Corruption
- **Affected Languages**: C, C++
### π Description
The product reuses or references memory after it has been freed. At some point
afterward, the memory may be allocated again and saved in another pointer, while
the original pointer references a location somewhere within the new allocation.
### π₯ Potential Impact
- Execute Unauthorized Code or Commands
- Read Memory
- Modify Memory
- DoS: Crash, Exit, or Restart
### π‘οΈ Recommendation
Review the code for Use After Free patterns. Ensure proper bounds checking,
input validation, and memory management practices are applied as recommended
by the CWE guidelines for CWE-416.π Usage
Load with π€ Datasets
from datasets import load_dataset
dataset = load_dataset("ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced")
print(dataset)
# DatasetDict({
# train: Dataset({features: ['conversations', 'source', 'score'], num_rows: ...}),
# test: Dataset({features: ['conversations', 'source', 'score'], num_rows: ...})
# })Access a Sample
sample = dataset['train'][0]
# Print the human question (code to analyze)
print(sample['conversations'][0]['value'])
# Print the assistant answer (vulnerability analysis)
print(sample['conversations'][1]['value'])Fine-tuning with Unsloth / TRL
from trl import SFTTrainer
from unsloth import FastLanguageModel
# The dataset is already in ShareGPT format β compatible with
# most fine-tuning frameworks that support conversation datasets.
trainer = SFTTrainer(
model=model,
tokenizer=tokenizer,
train_dataset=dataset['train'],
dataset_text_field="conversations", # adjust per framework
...
)π Dataset Lineage
bstee615/diversevul
βββ> ChamaraVishwajithRajapaksha/Code_Vulnerability_Dataset
(330k rows, CWE-enriched via MITRE API)
βββ> ChamaraVishwajithRajapaksha/Code-Vulnerability-FineTune
(ShareGPT format, unbalanced, unshuffled)
βββ> ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced
(balanced + shuffled β this dataset)β οΈ Limitations
- Code samples are primarily in C and C++ β limited coverage of other languages
- Balancing is achieved by undersampling the majority class, so total row count is reduced compared to the source dataset
- The Safe samples represent patched/fixed versions, not inherently safe code β context matters
- CWE details describe the class of vulnerability, not a precise analysis of each individual function
- This dataset is intended for research and educational purposes
π License
This dataset is released under the MIT License, consistent with the source dataset license.
π Citation
If you use this dataset in your research, please cite the original source and this dataset:
@dataset{code_vulnerability_balanced,
title = {Code Vulnerability Balanced: CWE-Enriched Conversation Dataset},
author = {ChamaraVishwajithRajapaksha},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/datasets/ChamaraVishwajithRajapaksha/Code-Vulnerability-Balanced},
note = {Balanced and shuffled version of Code-Vulnerability-FineTune, in ShareGPT format}
}