datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
assembly101
Assembly101
Assembly101 is a procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles. Participants work without fixed instructions, and the sequences feature rich and natural variations in action ordering, mistakes, and corrections. Assembly101 is the first multi-view action dataset, with simultaneous static (8) and egocentric (4) recordings. Sequences are annotated with more than 100K coarse and 1M fine-grained… See the full description on the dataset page: https://huggingface.co/datasets/cvml-nus/assembly101.WavCaps
WavCaps
WavCaps is a ChatGPT-assisted weakly-labelled audio captioning dataset for audio-language multimodal research, where the audio clips are sourced from three websites (FreeSound, BBC Sound Effects, and SoundBible) and a sound event detection dataset (AudioSet Strongly-labelled Subset).
Paper: https://arxiv.org/abs/2303.17395
Github: https://github.com/XinhaoMei/WavCaps
Statistics
Data Source
# audio
avg. audio duration (s)avg. text length
FreeSound… See the full description on the dataset page: https://huggingface.co/datasets/cvssp/WavCaps.CVPR-BiomedSegFMThis repository contains the BiomedSegFM dataset, a crucial resource for the CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation.
Foundation Models for Interactive 3D Biomedical Image Segmentation (Homepage)
Foundation Models for Text-guided 3D Biomedical Image Segmentation (Homepage)
CVPR 2025 Competition: Foundation Models for 3D Biomedical Image Segmentation
Highly recommend watching the webinar recording to learn about the task settings and… See the full description on the dataset page: https://huggingface.co/datasets/junma/CVPR-BiomedSegFM.new-york-smells
New York Smells: A Large Multimodal Dataset for Olfaction
While olfaction is central to how animals perceive the world, this rich chemical
sensory modality remains largely inaccessible to machines. One key bottleneck is the
lack of diverse, multimodal olfactory data collected in natural settings. We present
New York Smells, a large-scale dataset of paired image and olfactory signals
captured in-the-wild. Our dataset contains 7,000 smell-image pairs from 3,500 distinct
objects… See the full description on the dataset page: https://huggingface.co/datasets/cvlab/new-york-smells.PocketQubeobelics_seed2_tokensPart of the OBELISC data set, including 32 Million samples, please refer to dataset.py to use this data
CV-Bench
Cambrian Vision-Centric Benchmark (CV-Bench)
This repository contains the Cambrian Vision-Centric Benchmark (CV-Bench), introduced in Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.
Files
The test*.parquet files contain the dataset annotations and images pre-loaded for processing with HF Datasets.
These can be loaded in 3 different configurations using… See the full description on the dataset page: https://huggingface.co/datasets/nyu-visionx/CV-Bench.SwissCubeSportsSlomo-CVS
🎥 SportsSloMo-CVS Dataset
This repository contains the dataset presented in the paper Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor.
The Complementary Vision Sensor (CVS), known as Tianmouc, captures synchronized RGB frames together with high-frame-rate, multi-bit spatial difference (SD, encoding structural edges) and temporal difference (TD, encoding motion cues) data within a single RGB exposure. This dataset facilitates research in RGB… See the full description on the dataset page: https://huggingface.co/datasets/mypThu/SportsSlomo-CVS.cvdp-benchmark-datasetImportant please see "Files and versions" above for full list of files in the CVDP dataset.
Please see LICENSE and NOTICE for licensing information. See CHANGELOG for changes.
This is the Comprehensive Verilog Design Problems (CVDP) benchmark dataset to use with the CVDP infrastructure on GitHub.
ComicsPAP
Comics: Pick-A-Panel
Updated val and test on 25/02/2025
This is the dataset for the ICDAR 2025 Competition on Comics Understanding in the Era of Foundational Models.
Please, check out our 🚀 arxiv paper 🚀 for more information 😊
The competition is hosted in the Robust Reading Competition website and the leaderboard is available here.
The dataset contains five subtask or skills:
Sequence Filling
Given a sequence of comic panels, a missing panel, and a set of option panels, the… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/ComicsPAP.SEED-Bench
SEED-Bench Card
Benchmark details
Benchmark type:
SEED-Bench is a large-scale benchmark to evaluate Multimodal Large Language Models (MLLMs).
It consists of 19K multiple choice questions with accurate human annotations, which
covers 12 evaluation dimensions including the comprehension of both the image and video modality.
Benchmark date:
SEED-Bench was collected in July 2023.
Paper or resources for more information:
https://github.com/AILab-CVC/SEED-Bench
License:… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Bench.cvm-corpus
CVM Filings Corpus (PT-BR)
Brazilian CVM regulatory filings in Portuguese, cleaned and chunked for language-model pretraining. Built for DAPT on financial Portuguese.
Contents
Path
What
output/corpus.jsonl
Full corpus (7.2 GB). Chunk schema: text, company, cnpj, category, subject, date, year, document_id, chunk_id, extraction_quality.
output/corpus-250M.jsonl
DSIR-selected 250M-token subset (token count by chars÷4 proxy ≈ 150M whitespace tokens): 187… See the full description on the dataset page: https://huggingface.co/datasets/heitorrosa/cvm-corpus.cv_corpus_v22
Dataset Card for Common Voice Corpus 22.0
This dataset is an unofficial version of the Mozilla Common Voice Corpus 22. It was downloaded and converted from the project's website https://commonvoice.mozilla.org/.
NOTE: currently converting to parquet for convenience.. WIP
Languages
Abkhaz, Albanian, Amharic, Arabic, Armenian, Assamese, Asturian, Azerbaijani, Basaa, Bashkir, Basque, Belarusian, Bengali, Breton, Bulgarian, Cantonese, Catalan, Central Kurdish, Chinese… See the full description on the dataset page: https://huggingface.co/datasets/mort666/cv_corpus_v22.Long-Horizon-GUI-DatasetBoilingBench-CV
BoilingBench-CV Dataset
Version: v0.1.0
Maintainer: NED3 Laboratory, University of Arkansas
License: CC BY 4.0
DOI: 10.5281/zenodo.22264378
Mirror of the Zenodo deposit of 3 September 2026, published here because most
users of these data work in the Hugging Face ecosystem. The file set was
verified identical to the deposit at upload time: 7,147 files, 4.20 GB
uncompressed.
Authors
Hari Pandey (University of Arkansas), Manohar Bongarala (Purdue University),
Christy… See the full description on the dataset page: https://huggingface.co/datasets/UARK-NED3/BoilingBench-CV.cve-proof-corpus
CVE Proof Corpus
Six real vulnerability classes, each with a machine-checkable proof that the shipped fix
eliminates it — and a checker that shares no code with whatever produced the proof.
Every record carries the safety relation, the guard the upstream project shipped, the declared
attacker domain, and the nonnegative multipliers that prove the guard implies safety. All six verify.
pip install "certkit@git+https://github.com/nickharris808/certkit@main"
python verify.py… See the full description on the dataset page: https://huggingface.co/datasets/nickh007/cve-proof-corpus.cv-v1.0-segment
CommonVoice v1 Phone-Segment Alignments
Phone-level time alignments for 10 languages of Mozilla Common Voice,
packaged in a canonical segmentation schema with embedded 16 kHz audio. The
phone boundaries come from the charsiu/cv_ali
release of MFA alignments; the audio and transcripts come from
Common Voice Corpus 13.0 (2023-03-09).
Dataset summary
lang
train rows
train hrs
val rows
val hrs
test rows
test hrs
en
1,008,669
1,354.0
3,537
4.9
1,285
1.7
rw… See the full description on the dataset page: https://huggingface.co/datasets/changelinglab/cv-v1.0-segment.shinka-cvdp-benchmark-fulldoc3D-dataset
doc3D
Doc3D is the first 3D dataset focused on document unwarping with realistic paper warping and renderings.
It contains 100k images with the following ground-truths:
3D Coordinates
Depth
UV
Backward Mapping
Albedo
Normals
Checkerboard
Useful links:
More details of the data usage instructions are available in the GitHub repo:
https://github.com/cvlab-stonybrook/doc3D-dataset
Link to the training code: https://github.com/cvlab-stonybrook/DewarpNet
Link to the… See the full description on the dataset page: https://huggingface.co/datasets/StonyBrook-CVLab/doc3D-dataset.cvqa
About CVQA
CVQA is a culturally diverse multilingual VQA benchmark consisting of over 10,000 questions from 39 country-language pairs. The questions in CVQA are written in both the native languages and English, and are categorized into 10 diverse categories.
This data is designed for use as a test set. Please submit your submission here to evaluate your model performance. CVQA is constructed through a collaborative effort led by a team of researchers from MBZUAI. Read more about… See the full description on the dataset page: https://huggingface.co/datasets/afaji/cvqa.SEED-Data-Edit-Part1-Openimages
SEED-Data-Edit
SEED-Data-Edit is a hybrid dataset for instruction-guided image editing with a total of 3.7 image editing pairs, which comprises three distinct types of data:
Part-1: Large-scale high-quality editing data produced by automated pipelines (3.5M editing pairs).
Part-2: Real-world scenario data collected from the internet (52K editing pairs).
Part-3: High-precision multi-turn editing data annotated by humans (95K editing pairs, 21K multi-turn rounds with a maximum of 5… See the full description on the dataset page: https://huggingface.co/datasets/AILab-CVC/SEED-Data-Edit-Part1-Openimages.cvc-clinicdb
CVC-ClinicDB (file mirror)
612 frames extracted from 29 colonoscopy sequences, each with a pixel-level
polyp mask (Bernal et al., 2015). Native resolution 384x288.
This is a plain file mirror, not a datasets-format repo: images and masks
are stored as files so any path-based dataloader can consume them directly
after snapshot_download. The dataset viewer is disabled for that reason.
Layout
CVC-ClinicDB/
Original/*.png # 612 RGB frames, 384x288
Ground… See the full description on the dataset page: https://huggingface.co/datasets/berkaytrhn/cvc-clinicdb.cvsearch_hr8kVisualizations of CVSearch
Citation
@misc{li2026cvsearchempoweringmultimodalllms,
title={CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception},
author={Liupeng Li and Haoqian Kang and Zhenyu Lu and Jinpeng Wang and Bin Chen and Ke Chen and Yaowei Wang},
year={2026},
eprint={2605.23655},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.23655},
}… See the full description on the dataset page: https://huggingface.co/datasets/tothanhdat/cvsearch_hr8k.DocVQA-2026
DocVQA 2026 | ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains
Building upon previous DocVQA benchmarks, this evaluation dataset introduces challenging reasoning questions over a diverse collection of documents spanning eight domains, including business reports, scientific papers, slides, posters, maps, comics, infographics, and engineering drawings.
By expanding coverage to new document domains and… See the full description on the dataset page: https://huggingface.co/datasets/VLR-CVC/DocVQA-2026.SPADES
SPADES dataset
SPADES - SPAcecraft Pose Estimation Dataset using Event Sensing, a unique and new space dataset designed to advance spacecraft pose estimation research. SPADES dataset contains two categories of data: Synthetic and Real.
Synthetic dataset focuses on simulating RGB images of a satellite target—in this case, Proba-2—by moving a spacecraft model along predefined trajectories within the simulator’s camera field of view. To generate realistic imagery, the Unreal… See the full description on the dataset page: https://huggingface.co/datasets/CVI2-UniLU/SPADES.CVPR_2024_Papers
Dataset Card for cvpr2024_papers
This is a FiftyOne dataset with 2379 samples.
The dataset consists of images of the first page for accepted papers to CVPR 2024, plus their abstract and other metadata.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'split', 'max_samples', etc
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/CVPR_2024_Papers.cvefixes
CVEfixes Security Vulnerabilities Dataset
Security vulnerability data from CVEfixes v1.0.8 with 12,987 vulnerability fix records across 11,726 unique CVEs and 4,205 repositories.
Contains CVE metadata (descriptions, CVSS scores, CWE classifications), git commit data, and code diffs showing vulnerable vs fixed code.
Usage
from datasets import load_dataset
dataset = load_dataset("hitoshura25/cvefixes")
Citation
If you use this dataset, please cite the original… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/cvefixes.gta-data-files-universalcv-project
