CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /openswe-tasks-patched-v5text10K<n<100K0 likes892 downloads5mo agoHugging Face02patched-codes /static-analysis-evalA dataset of 76 Python programs taken from real Python open source projects (top 100 on GitHub), where each program is a file that has exactly 1 vulnerability as detected by a particular static analyzer (Semgrep), used in the paper Patched MOA: optimizing inference for diverse software development tasks. OpenAI used the synth-vuln-fixes and fine-tuned a new version of gpt-4o is now the SOTA on this benchmark. More details and code is available from their repo. More details on the benchmark… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/static-analysis-eval.textn<1K20 likes730 downloads1y agoHugging Face03closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13image10M<n<100M0 likes533 downloads4y agoHugging Face04CIRCL /vulnerability-cwe-patch Description This dataset, CIRCL/vulnerability-cwe-patch, provides structured, real-world vulnerabilities enriched with CWE identifiers and corresponding patches from platforms like GitHub and GitLab. It is designed to support the development of tools for vulnerability classification, triage, and automated remediation. Each entry includes metadata such as CVE/GHSA ID, a description, CWE categorization, and links to verified patch commits with associated diff content and commit… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-cwe-patch.text1K<n<10K4 likes507 downloads2mo agoHugging Face05patched-codes /generate-readme-eval Generate README Eval The generate-readme-eval is a dataset (train split) and benchmark (test split) to evaluate the effectiveness of LLMs when summarizing entire GitHub repos in form of a README.md file. The datset is curated from top 400 real Python repositories from GitHub with at least 1000 stars and 100 forks. The script used to generate the dataset can be found here. For the dataset we restrict ourselves to GH repositories that are less than 100k tokens in size to allow us to… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/generate-readme-eval.textsummarizationn<1K3 likes452 downloads2y agoHugging Face06anaumghori /patchlet-embed-preprocessedimage100K<n<1M0 likes362 downloads7mo agoHugging Face07mlfoundations-dev /swe_gym_annotate_with_patchtext10K<n<100K0 likes262 downloads2y agoHugging Face08Martingkc /LLaVa-Instruct-150K-clip-vit-base-patch32text100K<n<1M0 likes260 downloads5mo agoHugging Face09prakanda /SynthMat_Patches_DStext100K<n<1M0 likes257 downloads2y agoHugging Face10rasdani /github-patchestext10K<n<100K0 likes254 downloads1y agoHugging Face11DCAgent /swe_rebench_patchedtext1K<n<10K0 likes230 downloads7mo agoHugging Face12DCAgent /rl__24GPU_base__swe_rebench_patched_oracle__r2egym-nl2bash-stacktext10K<n<100K0 likes203 downloads6mo agoHugging Face13R2E-Gym /R2E-TestgenAgent-Patchestextn<1K1 likes201 downloads1y agoHugging Face14R2E-Gym /R2EGym-VerifierTrajectories-PatchOnlytext1K<n<10K0 likes154 downloads2y agoHugging Face15rasdani /github-patches-genesysimport re import json from datasets import load_dataset PROMPT_TEMPLATE = """\ We are currently solving the following issue within our repository. Here is the issue text: --- BEGIN ISSUE --- {issue} --- END ISSUE --- Below are some code segments, each from a relevant file. One or more of these files may contain bugs. --- BEGIN FILES --- {file_context} --- END FILES --- Please first localize the bug based on the issue statement, and then generate a patch according to the `git diff` format… See the full description on the dataset page: https://huggingface.co/datasets/rasdani/github-patches-genesys.text10K<n<100K0 likes148 downloads1y agoHugging Face16ai-sec-lab /PatchBench PatchBench PatchBench is a benchmark for evaluating AI agents on realistic vulnerability patching tasks: 213 tasks drawn from 32 popular GitHub C/C++ projects. It selects vulnerabilities whose ground-truth fixes lie outside the crash stack, and uses vulnerability transplant plus code mutation to mitigate surface-level fixes and patch memorization. This repository holds the task metadata, one row per task to identify the project, the exact repository state, the crash, and the… See the full description on the dataset page: https://huggingface.co/datasets/ai-sec-lab/PatchBench.texttext-generationn<1K0 likes144 downloads3d agoHugging Face17TrevorJS /mtg-scryfall-cropped-art-embeddings-siglip-so400m-patch14-384image10K<n<100K0 likes132 downloads2y agoHugging Face18closji /cc12m_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-15image10M<n<100M0 likes131 downloads4y agoHugging Face19open-athena /rl__24GPU_shaped__swe_rebench_patched_oracle__r2egym-nl2bash-stacktext10K<n<100K0 likes131 downloads6mo agoHugging Face20closji /mscoco_train_2014_openai_clip-vit-base-patch32_image_image_retrieval_pairs_2022-09-13tabular10M<n<100M0 likes130 downloads4y agoHugging Face21closji /wikitext-103-raw-v1_sents_min_len10_max_len30_openai_clip-vit-base-patch32text1M<n<10M0 likes128 downloads4y agoHugging Face22michoo42 /Patchnoisseur Patchnoisseur A connoisseur's cellar of CVEs: every NVD CVE joined to its fixing-commit diff (when one could be found), its NVD description, and its associated CWE(s) (id, name, short description) — served as a single Parquet dataset. 351 884 CVEs · 25 015 with a real git diff attached · CVE-1999 → CVE-2026 · ~744 MB on disk (zstd-compressed Parquet, sharded ~300 MB each). What's in it One row per CVE in the NVD feed. CVEs without a retrievable patch are… See the full description on the dataset page: https://huggingface.co/datasets/michoo42/Patchnoisseur.tabulartext-classification100K<n<1M0 likes127 downloads4mo agoHugging Face23open-athena /a3-rl-DCAgent_r2egym-patched-full-oracletext10K<n<100K0 likes126 downloads4mo agoHugging Face24DCAgent /swe_rebench_patched_oracletext1K<n<10K0 likes125 downloads7mo agoHugging Face25rasdani /github-patches-decontaminated# removed all repos of SWE-bench and RepoBench repos = [ "astropy", "django", "flask", "matplotlib", "seaborn", "requests", "xarray", "pylint", "pytest", "scikit-learn", "sphinx", "sympy", ] text10K<n<100K0 likes124 downloads1y agoHugging Face26Transluce /act_patch_llama_3.1_8b_counterfact Training Language Models to Explain Their Own Computations Paper | Code This dataset contains activation patching results used for training explainer models to predict how internal interventions affect target model outputs. It was introduced in the paper "Training Language Models to Explain Their Own Computations". Dataset Summary The dataset covers the Activation Patching task for the Llama-3.1-8B target model, where explainer models learn to predict the effects of… See the full description on the dataset page: https://huggingface.co/datasets/Transluce/act_patch_llama_3.1_8b_counterfact.texttext-generation100K<n<1M0 likes115 downloads9mo agoHugging Face27open-athena /r2egym-patched-full-oracle-qwen3.5-122b-131k-opencode-literal-rescue-traces Agent trace dataset Decoding the literal token IDs The prompt_token_ids / completion_token_ids / logprobs columns are the verbatim tokens the serving engine emitted, stored PER AGENT STEP as a list-of-lists (one inner list per turn). To turn them back into text you MUST use the exact tokenizer the model was served with — a generic same-family tokenizer will decode word tokens to garbage. Served model / tokenizer source: Qwen/Qwen3.5-122B-A10B-FP8 from transformers… See the full description on the dataset page: https://huggingface.co/datasets/open-athena/r2egym-patched-full-oracle-qwen3.5-122b-131k-opencode-literal-rescue-traces.text1K<n<10K0 likes114 downloads1mo agoHugging Face28DCAgent /Kimi-2.5-swe_rebench_patched-maxeps-32ktext10K<n<100K0 likes109 downloads5mo agoHugging Face29pykale /bln600-img-patch BLN600 Image Patches This dataset provides BLN600's image patches for fine-tuning vision-language models on post-OCR correction, introduced in "Image-Informed Post-OCR Correction with Vision-Language Models" (EMNLP 2026 Findings). Each patch corresponds to a text sequence in BLN600 and is cropped from the Gale British Library Newspapers collection, using word-level bounding boxes from the collection's ALTO XML OCR layout data. It is intended to be used alongside the code and… See the full description on the dataset page: https://huggingface.co/datasets/pykale/bln600-img-patch.imageimage-to-text10K<n<100K0 likes103 downloads28d agoHugging Face30yurkes /patch_tasks_vllm Dataset Card for Patch-Based Visual Question Answering Dataset Dataset Details Dataset Description This dataset contains approximately 305,000 triplets of question, answer, and image designed for patch-based visual reasoning tasks. A standard question in this dataset is formatted as follows: Image Grid: The image is divided into a 4x4 grid of 16 equal-sized patches. Patches are numbered sequentially from the top-left corner and moving right, then down to the… See the full description on the dataset page: https://huggingface.co/datasets/yurkes/patch_tasks_vllm.imageimage-text-to-text100K<n<1M3 likes94 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.