datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DeepTumorVQA_2.0
DeepTumorVQA v2
3D abdominal-CT diagnostic Visual Question Answering benchmark with 42
clinical subtypes and 438K total QA pairs (10K curated benchmark + 428K
training pool). Includes pre-extracted 2D and video modalities, 20K agent
training trajectories with tool-use traces, and a paper-locked leaderboard.
Resources
📄 Paper (arXiv)
https://arxiv.org/abs/2605.09679
💻 Code (GitHub)
https://github.com/Schuture/DeepTumorVQA
🤗 Dataset (this… See the full description on the dataset page: https://huggingface.co/datasets/tumor-vqa/DeepTumorVQA_2.0.AV-Deepfake1M
AV-Deepfake1M
This is the official repository for the paper
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset.
Abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most
advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting
high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.DeepSea-MOT
DeepSea MOT
DeepSea MOT is a benchmark dataset for multi-object tracking on deep-sea video.
Dataset Description
DeepSea MOT consists of 4 video sequences (2 midwater, 2 benthic) with a total of 2,400 frames and 57,376 annotated objects comprising 188 tracks. The videos were captured by the Monterey Bay Aquarium Research Institute (MBARI) using remotely operated vehicles (ROVs) Doc Ricketts and Ventana in deep-sea environments, showcasing a variety of marine species and… See the full description on the dataset page: https://huggingface.co/datasets/MBARI-org/DeepSea-MOT.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/GT-Neuronext/human-motion-tracking-deeplabcut.human-motion-tracking-deeplabcutThis dataset is used to adapt DeepLabCut for Human motion tracking.
Structure of the dataset
videos contains 100+ videos of 4 candidates recorded during a game of darts.
labeled-data contains labels on the corresponding frames of the videos. These labels are used to adapt DeepLabCut for human motion tracking. Under labeled-data there are 2 folders for every video.
video_name has all the relevant frames extracted from the video, xy coordinates of the labels in the csv file and the… See the full description on the dataset page: https://huggingface.co/datasets/pratikshapai/human-motion-tracking-deeplabcut.deepfake-videos-dataset
DeepFake Videos for detection tasks
Dataset consists of 10,000+ files featuring 7,000+ people, providing a comprehensive resource for research in deepfake detection and deepfake technology. It includes real videos of individuals with AI-generated faces overlaid, specifically designed to enhance liveness detection systems.
By utilizing this dataset, researchers can advance their understanding of deepfake generation and improve the performance of detection methods. - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/deepfake-videos-dataset.CVQAD
MSU Compression Dataset Description
We developed LEHA-CVQAD dataset to evaluate full-reference and no-reference video quality metrics. Here we share the open part of the whole compression artifacts dataset (1,962 out of 6,240 videos). The hidden part is only available to benchmark-support personnel for testing metric performance. All videos are of mostly FullHD resolution, YUV420, and 10-15 seconds duration. Fps values are 24, 25, 30, 39, 50, and 60.
Subjective quality scores are… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/CVQAD.deepfake_video_datain-the-wild-deepfake
mohammedph197/in-the-wild-deepfake
Media collected by a deepfake-dataset pipeline, published for annotation.
One row per item, its media referenced by URL:
column
meaning
media_url
public URL of the file in this repo; the media is not distributed in the table
media_type
video, audio, image, or unknown
Files are content-addressed: a file's name is the SHA-256 of its bytes, so identical media appears once however many source records pointed at it.
Deepfake-Eval-2024
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
Deepfake-Eval-2024 is an in-the-wild deepfake dataset. Deepfake-Eval-2024 contains 44 hours of videos, 56.5 hours of audio, and 1,975 images, encompassing contemporary manipulation technologies, diverse media content, 88 different website sources, and 52 different languages. Deepfake-Eval-2024 contains manually labeled real and fake media. Deepfake-Eval-2024 is designed to facilitate deepfake… See the full description on the dataset page: https://huggingface.co/datasets/nuriachandra/Deepfake-Eval-2024.deepfake-videoso100_test2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 15469,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/deepinsand/so100_test2.DeepSound-V1quest-bimanual-dataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/deepakshankar94/quest-bimanual-data.whitecolor-cuboidbluenotepad-7dof-dataset-v2SO101_LAR_gripper_with_adapter_deepThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 20,
"total_frames": 10557,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:20"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Hugo-Castaing/SO101_LAR_gripper_with_adapter_deep.AV-Deepfake1M-PlusPlus
AV-Deepfake1M++
The dataset used for the 2025 1M-Deepfakes Detection Challenge.
Task 1 Video-Level Deepfake Detection:
Given an audio-visual sample containing a single speaker, the task is to identify if the video is a deepfake or real.
Task 2 Deepfake Temporal Localization:
Given an audio-visual sample containing a single speaker, the task is to find out the timestamps [start, end] in which the manipulation is done.
The assumption here is that from the perspective of spreading… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M-PlusPlus.high-camera-updated-v2deep-fake-detection-dfd-entire-original-datasetdeep-fake-detection-dfd-entire-original-datasettDeepFake_Extracted_Face_Imagesbi_so101_fold_10inch_4camThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so_follower",
"total_episodes": 603,
"total_frames": 541612,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 30,
"splits": {
"train": "0:603"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/deepreach/bi_so101_fold_10inch_4cam.multicolor-cuboidnotepad-7dof-dataset-v2RideSafe-400
RideSafe-400
Dataset Description
Dataset Summary
RideSafe-400 is a dataset of annotated dashcam videos designed specifically for detecting traffic violations involving motorized two-wheelers, such as helmet non-compliance and triple riding. The dataset was created to address the lack of publicly available resources tailored to these safety violations. It supports tasks like violation detection, traffic safety analysis, and automated E-ticket generation.… See the full description on the dataset page: https://huggingface.co/datasets/DeepBug/RideSafe-400.social_media_deepfakesSpatial-Scene-Synthetic-Datasetbluetowel-7dof-dataset-30fpsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"left_shoulder_pan.pos",
"left_shoulder_lift.pos",
"left_elbow_flex.pos",
"left_wrist_flex.pos",
"left_wrist_yaw.pos",
"left_wrist_roll.pos"… See the full description on the dataset page: https://huggingface.co/datasets/deepakshankar94/bluetowel-7dof-dataset-30fps.rppg-deepfake-detectionbulk-v1-progress-sample
bulk_v1 progress sample
A random look at the clips still on disk (withdrawn rejects are excluded). Not a
release: the whole dataset is replaced on every run. selection is kept/rejected
once a slice is gated, and not gated yet only for slices that never ran the gate.
drawn: 2026-09-24 07:18, seed 20260924, 3 pair-set(s) per family x owner x slice, slices 2 and up
selection: kept / rejected after GATE_DONE or the drop ledger (or refill kept lists); not gated yet only if the slice… See the full description on the dataset page: https://huggingface.co/datasets/Deep-Fake/bulk-v1-progress-sample.agilex_arrange_word_DEEPThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 10,
"total_frames": 13731,
"total_tasks": 1,
"total_videos": 30,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_arrange_word_DEEP.
