datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
openswe-tasks-patched-v5static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub),
where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks.
OpenAI used the synth-vuln-fixes and fine-tuned
a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo.
More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13vulnerability-cwe-patch
Description
This dataset, CIRCL/vulnerability-cwe-patch, provides structured, real-world vulnerabilities enriched with CWE identifiers and corresponding patches from platforms like GitHub and GitLab. It is designed to support the development of tools for vulnerability classification, triage, and automated remediation. Each entry includes metadata such as CVE/GHSA ID, a description, CWE categorization, and links to verified patch commits with associated diff content and commit… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-cwe-patch.generate-readme-eval
Generate README Eval
The generate-readme-eval is a dataset (train split) and benchmark (test split) to evaluate the effectiveness of LLMs
when summarizing entire GitHub repos in form of a README.md file. The datset is curated from top 400 real Python repositories
from GitHub with at least 1000 stars and 100 forks. The script used to generate the dataset can be found here.
For the dataset we restrict ourselves to GH repositories that are less than 100k tokens in size to allow us to… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/generate-readme-eval.urban3d-patchespatchlet-embed-preprocessedtraining_03_05_patchPatchCamelyon
PatchCamelyon (PCam)
This is a reupload of the PatchCamelyon (PCam) dataset to make it more readily usable instead of manipulating H5 files. The original can be found in the author's Github repo.
If you use this dataset, please cite the original publications:
@inproceedings{veeling2018rotation,
title={Rotation Equivariant CNNs for Digital Pathology},
author={Veeling, Bastiaan S and Linmans, Jasper and Winkens, Jim and Cohen, Taco and Welling, Max},
booktitle={Medical Image… See the full description on the dataset page: https://huggingface.co/datasets/zacharielegault/PatchCamelyon.swe_gym_annotate_with_patchpatch_camelyonLLaVa-Instruct-150K-clip-vit-base-patch32github-patchesSynthMat_Patches_DSdiffllama_patch_tokenizedswe_rebench_patchedR2E-TestgenAgent-Patchesrl__24GPU_base__swe_rebench_patched_oracle__r2egym-nl2bash-stackPatchCamelyon
PatchCamelyon (PCam)
Description
The PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU
Why PCam
Fundamental… See the full description on the dataset page: https://huggingface.co/datasets/pavan316/PatchCamelyon.PatchCamelyon
PatchCamelyon (PCam)
Description
The PatchCamelyon benchmark is a new and challenging image classification dataset. It consists of 327.680 color images (96 x 96px) extracted from histopathologic scans of lymph node sections. Each image is annoted with a binary label indicating presence of metastatic tissue. PCam provides a new benchmark for machine learning models: bigger than CIFAR10, smaller than imagenet, trainable on a single GPU
Why PCam
Fundamental… See the full description on the dataset page: https://huggingface.co/datasets/KE9037/PatchCamelyon.R2EGym-VerifierTrajectories-PatchOnlygithub-patches-genesysimport re
import json
from datasets import load_dataset
PROMPT_TEMPLATE = """\
We are currently solving the following issue within our repository. Here is the issue text:
--- BEGIN ISSUE ---
{issue}
--- END ISSUE ---
Below are some code segments, each from a relevant file. One or more of these files may contain bugs.
--- BEGIN FILES ---
{file_context}
--- END FILES ---
Please first localize the bug based on the issue statement, and then generate a patch according to the `git diff` format… See the full description on the dataset page: https://huggingface.co/datasets/rasdani/github-patches-genesys.PatchBench
PatchBench
PatchBench is a benchmark for evaluating AI agents on realistic vulnerability patching tasks: 213 tasks drawn from 32 popular GitHub C/C++ projects. It selects vulnerabilities whose ground-truth fixes lie outside the crash stack, and uses vulnerability transplant plus code mutation to mitigate surface-level fixes and patch memorization.
This repository holds the task metadata, one row per task to identify the project, the exact repository state, the crash, and the… See the full description on the dataset page: https://huggingface.co/datasets/ai-sec-lab/PatchBench.mtg-scryfall-cropped-art-embeddings-siglip-so400m-patch14-384rl__24GPU_shaped__swe_rebench_patched_oracle__r2egym-nl2bash-stackmscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13wikitext-103-raw-v1_sents_min_len10_max_len30_openai_clip-vit-base-patch32cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15Patchnoisseur
Patchnoisseur
A connoisseur's cellar of CVEs: every NVD CVE joined to its fixing-commit
diff (when one could be found), its NVD description, and its
associated CWE(s) (id, name, short description) — served as a single
Parquet dataset.
351 884 CVEs · 25 015 with a real git diff attached · CVE-1999 → CVE-2026
· ~744 MB on disk (zstd-compressed Parquet, sharded ~300 MB each).
What's in it
One row per CVE in the NVD feed. CVEs without a retrievable patch are… See the full description on the dataset page: https://huggingface.co/datasets/michoo42/Patchnoisseur.a3-rl-DCAgent_r2egym-patched-full-oracle
