datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
in-the-wild-deepfake
mohammedph197/in-the-wild-deepfake
Media collected by a deepfake-dataset pipeline, published for annotation.
One row per item, its media referenced by URL:
column
meaning
media_url
public URL of the file in this repo; the media is not distributed in the table
media_type
video, audio, image, or unknown
Files are content-addressed: a file's name is the SHA-256 of its bytes, so identical media appears once however many source records pointed at it.
deepfake-ecg-small
ECG Dataset
This repository contains an small version of the ECG dataset: https://huggingface.co/datasets/deepsynthbody/deepfake_ecg, split into training, validation, and test sets. The dataset is provided as CSV files and corresponding ECG data files in .asc format. The ECG data files are organized into separate folders for the train, validation, and test sets.
Folder Structure
.
├── train.csv
├── validate.csv
├── test.csv
├── train
│ ├── file_1.asc
│ ├── file_2.asc… See the full description on the dataset page: https://huggingface.co/datasets/deepsynthbody/deepfake-ecg-small.human-perception-audio-deepfake-2026
Human Audio Deepfake Perception 2026
A large-scale listening study evaluating how well humans detect modern audio
deepfakes. The dataset contains 35,532 deepfake-detection judgments from
1,768 anonymous participants across 138 TTS and voice-conversion systems,
collected via a publicly accessible online listening game in 2025–2026.
This is the successor to the 2021 ASVspoof-2019 perception study
(Müller, Pizzi & Williams, 2022)
and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.deepfakeDeepfake-Eval-2024-Protocals
WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
Download the Protocols
Install the datasets package:
pip install datasets
Log in with your Hugging Face account:
huggingface-cli login
Load the dataset in Python:
from datasets import load_dataset
# Download from HF and cache
ds = load_dataset("xxuan-speech/Deepfake-Eval-2024-Protocals")
Statistics of Deepfake-Eval-2024 Benchmark
Dataset
Total
Real
Fake… See the full description on the dataset page: https://huggingface.co/datasets/xxuan-speech/Deepfake-Eval-2024-Protocals.Sap_Kush_Med_Deepfake
Sap_Kush_Med_Deepfake Dataset
Paired medical-image forgery lineages across six modalities. Every lineage is
one source image, one mask, one seed: the arms differ only in what was done
inside the mask, so a comparison between arms isolates the manipulation rather
than an encoding artefact.
3956 lineages, 30385 files, 8.49 GiB.
v2 adds a removal arm grounded in human annotation for three more modalities
(endoscopy, ultrasound, MRI). v1 had removal for CT only.
What the… See the full description on the dataset page: https://huggingface.co/datasets/Kanhaiyya/Sap_Kush_Med_Deepfake.deepfake-detection-dataset-v2
Deepfake Detection Dataset V2
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images
CAM visualization images
CAM overlay images
Comparison images
Labels (real/fake)
Confidence scores
Image captions
Technical and non-technical… See the full description on the dataset page: https://huggingface.co/datasets/saakshigupta/deepfake-detection-dataset-v2.Deep-Fake-videodeep_fake_videos_valdeepfake-detection-realDeep-Fake-Imagedeep_fake_videosDeep-Fake-DetectionDeep-Fake-videoDeep-Fake-video
