datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BRSpeech-DF
🗣️ BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese
🧩 Description
BRSpeech-DF is the first publicly available dataset for deepfake speech detection in Portuguese, covering both Brazilian and European variants.
It contains 459,000 audio samples, including both real and synthetic speech generated using multiple zero-shot text-to-speech (TTS) models.
This dataset aims to foster the development of more robust, inclusive, and multilingual audio deepfake… See the full description on the dataset page: https://huggingface.co/datasets/AKCIT-Deepfake/BRSpeech-DF.IndicTTS-Deepfake-Challenge-Data
IndicTTS Deepfake Detection Challenge
Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip.
🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data
This is the official dataset for the challenge and must be used for training and evaluation.
📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.deepfake_video_datain-the-wild-deepfake
mohammedph197/in-the-wild-deepfake
Media collected by a deepfake-dataset pipeline, published for annotation.
One row per item, its media referenced by URL:
column
meaning
media_url
public URL of the file in this repo; the media is not distributed in the table
media_type
video, audio, image, or unknown
Files are content-addressed: a file's name is the SHA-256 of its bytes, so identical media appears once however many source records pointed at it.
deepfake-ecg-small
ECG Dataset
This repository contains an small version of the ECG dataset: https://huggingface.co/datasets/deepsynthbody/deepfake_ecg, split into training, validation, and test sets. The dataset is provided as CSV files and corresponding ECG data files in .asc format. The ECG data files are organized into separate folders for the train, validation, and test sets.
Folder Structure
.
├── train.csv
├── validate.csv
├── test.csv
├── train
│ ├── file_1.asc
│ ├── file_2.asc… See the full description on the dataset page: https://huggingface.co/datasets/deepsynthbody/deepfake-ecg-small.image-splicing-deepfake-mix-newro-dia-deepfake-audiodeepfake_iitm_rawdeepfake_iitm_rawvoxcpm_deepfake_dataset
voxcpm_deepfake_dataset
547 hours of synthetic speech deepfakes — 377,932 clips across 9 corpora and 6 languages (Mandarin, English, Spanish, French, Italian, Japanese), generated via voice cloning and hundreds of designed voice profiles.
Generated with VoxCPM-0.5B, built to train/evaluate speech deepfake detectors (part of the Forensics model family).
Built to improve detection of voice cloning, deepfakes, highly realistic synthetic voices, voice conversion/changing, and… See the full description on the dataset page: https://huggingface.co/datasets/eliya/voxcpm_deepfake_dataset.Deepfake_dataset_cybersentinal
Dataset Card for OpenFake
OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on politically and socially salient content where misinformation risk is highest. It pairs real photographs with synthetic counterparts produced by a wide range of frontier proprietary generators, open-source diffusion models, and community fine-tunes. A separate in-the-wild test set is sourced from Reddit to evaluate detector performance on naturally circulated… See the full description on the dataset page: https://huggingface.co/datasets/vjjoshi23/Deepfake_dataset_cybersentinal.bulk-v1-progress-sample
bulk_v1 progress sample
A random look at the clips still on disk (withdrawn rejects are excluded). Not a
release: the whole dataset is replaced on every run. selection is kept/rejected
once a slice is gated, and not gated yet only for slices that never ran the gate.
drawn: 2026-09-25 09:30, seed 20260925, 3 pair-set(s) per family x owner x slice, slices 2 and up
selection: kept / rejected after GATE_DONE or the drop ledger (or refill kept lists); not gated yet only if the slice… See the full description on the dataset page: https://huggingface.co/datasets/Deep-Fake/bulk-v1-progress-sample.deepfake-identity-isolated-fakeDeepfake-leonardo-stablecoghuman-perception-audio-deepfake-2026
Human Audio Deepfake Perception 2026
A large-scale listening study evaluating how well humans detect modern audio
deepfakes. The dataset contains 35,532 deepfake-detection judgments from
1,768 anonymous participants across 138 TTS and voice-conversion systems,
collected via a publicly accessible online listening game in 2025–2026.
This is the successor to the 2021 ASVspoof-2019 perception study
(Müller, Pizzi & Williams, 2022)
and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.deepfakeface-inpaintingseq-deepfake-facial-attributes-train
seq-deepfake-facial-attributes-train
Mirror of the exact bmcore v24 holdout subset. Source: rshaojimmy/Seq-DeepFake.
Mirror the exact local subset; do not expand to the full upstream archive. StyleGAN originals are synthetic per the v24 correction.
Contains 401 synthetic image files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 189f969af3bf6d06613535e19923c9a1445670a6.
Original source card
license:… See the full description on the dataset page: https://huggingface.co/datasets/34data/seq-deepfake-facial-attributes-train.deepfake-identity-isolated-realdeepfakeqwen3_deepfake_dataset
qwen3_deepfake_dataset
515.5 hours of synthetic speech deepfakes — 364,640 clips across 9 corpora and 6 languages (Mandarin, English, Spanish, French, Italian, Japanese), generated via voice cloning and hundreds of designed voice profiles.
Generated with Qwen3-TTS, built to train/evaluate speech deepfake detectors (part of the Forensics model family).
Built to improve detection of voice cloning, deepfakes, highly realistic synthetic voices, voice conversion/changing, and similar… See the full description on the dataset page: https://huggingface.co/datasets/eliya/qwen3_deepfake_dataset.seq-deepfake-facial-attributes-test
seq-deepfake-facial-attributes-test
Mirror of the exact bmcore v24 holdout subset. Source: rshaojimmy/Seq-DeepFake.
Mirror the exact local subset; do not expand to the full upstream archive. StyleGAN originals are synthetic per the v24 correction.
Contains 160 synthetic image files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: 189f969af3bf6d06613535e19923c9a1445670a6.
Original source card
license:… See the full description on the dataset page: https://huggingface.co/datasets/34data/seq-deepfake-facial-attributes-test.sowaiba01-deepfake-fake
sowaiba01-deepfake-fake
Mirror of the exact bmcore v24 holdout subset. Source: Sowaiba01/Deepfake.
All 5,426 local fake basenames match the public fake/ subset. Face swaps remain semisynthetic.
Contains 5426 semisynthetic image files. This mirror repackages the media; it does not grant additional rights.
Source revision reviewed: d8b189507fc026b3e0c794a06c318cb13d0c8743.
Original source card
license: mit
task_categories:
- image-classification… See the full description on the dataset page: https://huggingface.co/datasets/34data/sowaiba01-deepfake-fake.deepfake-detection-dataset-v3
Deepfake Detection Dataset V3
This dataset contains images and detailed explanations for training and evaluating deepfake detection models. It includes original images, manipulated images, confidence scores, and comprehensive technical and non-technical explanations.
Dataset Structure
The dataset consists of:
Original images (image)
CAM visualization images (cam_image)
CAM overlay images (cam_overlay)
Comparison images (comparison_image)
Labels (label): Binary… See the full description on the dataset page: https://huggingface.co/datasets/saakshigupta/deepfake-detection-dataset-v3.deepfakeface-text2imgimage-splicing-deepfake-mixDeepfake-Eval-2024-Protocals
WaveSP-Net: Learnable Wavelet-Domain Sparse Prompt Tuning for Speech Deepfake Detection
Download the Protocols
Install the datasets package:
pip install datasets
Log in with your Hugging Face account:
huggingface-cli login
Load the dataset in Python:
from datasets import load_dataset
# Download from HF and cache
ds = load_dataset("xxuan-speech/Deepfake-Eval-2024-Protocals")
Statistics of Deepfake-Eval-2024 Benchmark
Dataset
Total
Real
Fake… See the full description on the dataset page: https://huggingface.co/datasets/xxuan-speech/Deepfake-Eval-2024-Protocals.omnivoice_deepfake_dataset
omnivoice_deepfake_dataset
210.4 hours of synthetic speech deepfakes — 179,868 clips across 6 corpora and 5 languages (Mandarin, Spanish, French, Italian, Japanese), generated via voice cloning and hundreds of designed voice profiles.
Generated with OmniVoice, built to train/evaluate speech deepfake detectors (part of the Forensics model family).
Built to improve detection of voice cloning, deepfakes, highly realistic synthetic voices, voice conversion/changing, and similar… See the full description on the dataset page: https://huggingface.co/datasets/eliya/omnivoice_deepfake_dataset.deepfakecss-deepfake-datasetspeech-deepfake-detection-40k
