CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nickh007 /cve-proof-corpus CVE Proof Corpus Six real vulnerability classes, each with a machine-checkable proof that the shipped fix eliminates it — and a checker that shares no code with whatever produced the proof. Every record carries the safety relation, the guard the upstream project shipped, the declared attacker domain, and the nonnegative multipliers that prove the guard implies safety. All six verify. pip install "certkit@git+https://github.com/nickharris808/certkit@main" python verify.py… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/cve-proof-corpus.tabulartext-classificationn<1K0 likes2.1k downloads19d agoHugging Face02AILab-CVC /SEED-Data-Edit-Part1-Openimages SEED-Data-Edit SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data: Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs). Part-2: Real-world scenario data collected from the internet (52K editing pairs). Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.tabulartext-to-image1M<n<10M10 likes1.2k downloads2y agoHugging Face03sonalsannigrahi /cv22_azeros FLEURS (Lhotse cuts) Each language is a separate config. Load a single language's cuts as a HF Dataset of raw manifest records with, e.g.: from datasets import load_dataset ds = load_dataset("your-org/REPO_NAME", "bg_bg", split="train") If audio shards (recording.NNNNN.tar) are present alongside the cuts, the LANG/SPLIT/ folder is a valid Lhotse Shar directory. Download it (e.g. via snapshot_download) and load with Lhotse directly: from huggingface_hub import snapshot_download… See the full description on the dataset page: https://huggingface.co/datasets/sonalsannigrahi/cv22_azeros.tabular1M<n<10M0 likes348 downloads3mo agoHugging Face04cveinnt /kepler-arc-agi-3-traces Kepler 1.0 ARC-AGI-3 trace corpus Run artifacts from Kepler 1.0, an open-source agent harness for the 25 public ARC-AGI-3 games. A stock CLI coding agent encodes its theory of each game as an executable world_model.py, certifies it against the full recorded interaction history, plans inside the certified model, and acts through a guarded channel that voids the plan on the first misprediction. Project page · Code · Paper · Integrity record The canonical release contains two… See the full description on the dataset page: https://huggingface.co/datasets/cveinnt/kepler-arc-agi-3-traces.tabularn<1K0 likes216 downloads19d agoHugging Face05FluidInference /cv-corpus-25.0-ja Mozilla Common Voice 25.0 - Japanese Test Set (Complete) Dataset Description Complete Japanese test set from Mozilla Common Voice Corpus 25.0. This dataset contains all 9,019 validated test samples, compared to the partial 2,334-sample version previously available on HuggingFace. Key Features Size: 9,019 validated test utterances Coverage: 100% of official Common Voice 25.0 Japanese test split Multi-speaker: Diverse set of speakers with demographic metadata… See the full description on the dataset page: https://huggingface.co/datasets/FluidInference/cv-corpus-25.0-ja.audioautomatic-speech-recognition1K<n<10K0 likes182 downloads6mo agoHugging Face06TheFinAI /freebsd-cvs-archive 📦 FreeBSD CVS Archive (C/C++) Dataset Summary FreeBSD CVS Archive (C/C++) is a large-scale dataset of source code extracted from the historical FreeBSD CVS repository. The dataset focuses on C and C++ source files, providing structured samples suitable for code modeling, analysis, and benchmarking. Each sample includes: the dataset source commit year extracted code content token count (computed using GPT tokenizer) This dataset is designed for: code language modeling… See the full description on the dataset page: https://huggingface.co/datasets/TheFinAI/freebsd-cvs-archive.tabular100K<n<1M0 likes113 downloads6mo agoHugging Face07Contrastive-LM /tb4-clm-cv-embeddings-8k Terminal-Bench 4.0 CLM embeddings Native PyTorch embeddings for the 66-task, five-candidate Fable 5.1 MAX Terminal-Bench 4.0 evaluation job. These files support task-disjoint 3-fold CLM training and evaluation with the unified release-branch scripts. Contents evaluation/: 14,144 state/action pairs from all 330 trajectories and 66 tasks. train/: 9,198 pairs from the 191 successful trajectories (52 tasks). index.json: the 330-trial Harbor index used for… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/tb4-clm-cv-embeddings-8k.tabularn<1K0 likes25 downloads17h agoHugging Face08instinct-org /cv_chunked_tokenizedgated cv_chunked_tokenized This is a gated Uzbek tokenized speech dataset from instinct-org. This repository contains tokenized or prepared speech data for text-to-speech training workflows. Language Primary language: uz (Uzbek) Intended Use text-to-speech training Internal dataset curation, quality checks, and model evaluation Research or commercial use only after access approval and license review Data Notes Contains tokenized speech… See the full description on the dataset page: https://huggingface.co/datasets/instinct-org/cv_chunked_tokenized.tabulartext-to-speech10K<n<100K0 likes17 downloads1mo agoHugging Face09ijakenorton /cvefixes_for_ml@inproceedings{bhandari2021:cvefixes, title = {{CVEfixes: Automated Collection of Vulnerabilities and Their Fixes from Open-Source Software}}, booktitle = {{Proceedings of the 17th International Conference on Predictive Models and Data Analytics in Software Engineering (PROMISE '21)}}, author = {Bhandari, Guru and Naseer, Amara and Moonen, Leon}, year = {2021}, pages = {10}, publisher = {{ACM}}, doi = {10.1145/3475960.3475985}, copyright = {Open Access}… See the full description on the dataset page: https://huggingface.co/datasets/ijakenorton/cvefixes_for_ml.tabular10K<n<100K0 likes9 downloads1y agoHugging Face10BrachioLab /cvs-act CVS-Act: Action Recommendation for Critical View of Safety Assessment Dataset Description CVS-Act is a surgical action recommendation dataset for laparoscopic cholecystectomy grounded in Critical View of Safety (CVS) assessment. Each example corresponds to a CVS transition example and contains structured action recommendations over the current task label space for the left instrument, right instrument, and camera, with the original other actor field preserved when… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/cvs-act.tabulartext-classification1K<n<10K0 likes9 downloads4mo agoHugging Face11Byne /CV1-QSgatedimage10K<n<100K0 likes2 downloads10mo agoHugging Face12Contrastive-LM /tb21-clm-cv-embeddings-8k Terminal-Bench 2.1 CLM embeddings Native PyTorch Qwen3-8B embeddings of every step of the Terminal-Bench 2.1 Harbor job tbench-2-1-fable-5__xhigh-claude-code-simple (Claude Fable 5, reasoning effort xhigh, Claude Code 2.1.167; 89 tasks x 5 rollouts). They support task-disjoint 3-fold CLM training and best-of-5 evaluation with the release-branch scripts; the matching heads are Contrastive-LM/tb21-clm-cv-heads-8k. Contents evaluation/: 7,160 state/action pairs from… See the full description on the dataset page: https://huggingface.co/datasets/Contrastive-LM/tb21-clm-cv-embeddings-8k.tabularn<1K0 likes4h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.