datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cve-proof-corpus
CVE Proof Corpus
Six real vulnerability classes, each with a machine-checkable proof that the shipped fix
eliminates it — and a checker that shares no code with whatever produced the proof.
Every record carries the safety relation, the guard the upstream project shipped, the declared
attacker domain, and the nonnegative multipliers that prove the guard implies safety. All six verify.
pip install "certkit@git+https://github.com/nickharris808/certkit@main"
python verify.py… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/cve-proof-corpus.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.cv22_azeros
FLEURS (Lhotse cuts)
Each language is a separate config. Load a single language's cuts as a HF Dataset of raw manifest records with, e.g.:
from datasets import load_dataset
ds = load_dataset("your-org/REPO_NAME", "bg_bg", split="train")
If audio shards (recording.NNNNN.tar) are present alongside the cuts, the LANG/SPLIT/ folder is a valid Lhotse Shar directory. Download it (e.g. via snapshot_download) and load with Lhotse directly:
from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/sonalsannigrahi/cv22_azeros.kepler-arc-agi-3-traces
Kepler 1.0 ARC-AGI-3 trace corpus
Run artifacts from Kepler 1.0, an open-source agent harness for the 25 public
ARC-AGI-3 games. A stock CLI coding agent
encodes its theory of each game as an executable world_model.py, certifies it
against the full recorded interaction history, plans inside the certified
model, and acts through a guarded channel that voids the plan on the first
misprediction.
Project page ·
Code ·
Paper ·
Integrity record
The canonical release contains two… See the full description on the dataset page: https://huggingface.co/datasets/cveinnt/kepler-arc-agi-3-traces.cv-corpus-25.0-ja
Mozilla Common Voice 25.0 - Japanese Test Set (Complete)
Dataset Description
Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace.
Key Features
Size: 9,019 validated test utterances
Coverage: 100% of official Common Voice 25.0 Japanese test split
Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.freebsd-cvs-archive
📦 FreeBSD CVS Archive (C/C++)
Dataset Summary
FreeBSD CVS Archive (C/C++) is a large-scale dataset of source code extracted from the historical FreeBSD CVS repository. The dataset focuses on C and C++ source files, providing structured samples suitable for code modeling, analysis, and benchmarking.
Each sample includes:
the dataset source
commit year
extracted code content
token count (computed using GPT tokenizer)
This dataset is designed for:
code language modeling… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/freebsd-cvs-archive.tb4-clm-cv-embeddings-8k
Terminal-Bench 4.0 CLM embeddings
Native PyTorch embeddings for the 66-task, five-candidate Fable 5.1 MAX
Terminal-Bench 4.0 evaluation job. These files support task-disjoint 3-fold CLM
training and evaluation with the unified release-branch scripts.
Contents
evaluation/: 14,144 state/action pairs from all 330 trajectories and 66 tasks.
train/: 9,198 pairs from the 191 successful trajectories (52 tasks).
index.json: the 330-trial Harbor index used for… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/tb4-clm-cv-embeddings-8k.cv_chunked_tokenized
cv_chunked_tokenized
This is a gated Uzbek tokenized speech dataset from instinct-org.
This repository contains tokenized or prepared speech data for text-to-speech training workflows.
Language
Primary language: uz (Uzbek)
Intended Use
text-to-speech training
Internal dataset curation, quality checks, and model evaluation
Research or commercial use only after access approval and license review
Data Notes
Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.cvefixes_for_ml@inproceedings{bhandari2021:cvefixes,
title = {{CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software}},
booktitle = {{Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering (PROMISE '21)}},
author = {Bhandari, Guru and Naseer, Amara and Moonen, Leon},
year = {2021},
pages = {10},
publisher = {{ACM}},
doi = {10.1145/3475960.3475985},
copyright = {Open Access}… See the full description on the dataset page: https://huggingface.co/datasets/ijakenorton/cvefixes_for_ml.cvs-act
CVS-Act: Action Recommendation for Critical View of Safety Assessment
Dataset Description
CVS-Act is a surgical action recommendation dataset for laparoscopic cholecystectomy grounded in Critical View of Safety (CVS) assessment. Each example corresponds to a CVS transition example and contains structured action recommendations over the current task label space for the left instrument, right instrument, and camera, with the original other actor field preserved when… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/cvs-act.CV1-QStb21-clm-cv-embeddings-8k
Terminal-Bench 2.1 CLM embeddings
Native PyTorch Qwen3-8B embeddings of every step of the Terminal-Bench 2.1 Harbor job
tbench-2-1-fable-5__xhigh-claude-code-simple (Claude Fable 5, reasoning effort xhigh,
Claude Code 2.1.167; 89 tasks x 5 rollouts). They support task-disjoint 3-fold CLM training
and best-of-5 evaluation with the release-branch scripts; the matching heads are
Contrastive-LM/tb21-clm-cv-heads-8k.
Contents
evaluation/: 7,160 state/action pairs from… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/tb21-clm-cv-embeddings-8k.
