datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AV-Deepfake1M
AV-Deepfake1M
This is the official repository for the paper
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset.
Abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most
advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting
high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic voice… See the full description on the dataset page: https://huggingface.co/datasets/garystafford/deepfake-audio-detection.Deepfake
DeepGuard Deepfake Dataset
A paired deepfake dataset for training and benchmarking deepfake detection models.
Generated as part of the DeepGuard AI Project (2026).
📥 2,301+ all-time downloads
Dataset Statistics
Fake images: 5426 (face-swapped using InsightFace inswapper_128)
Real images: 5426 (original paired faces)
Total: 10852 images
Format: JPEG, high quality (95%)
Generation method: InsightFace inswapper_128 (ONNX runtime)
Why This Dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Sowaiba01/Deepfake.DeepFakeFace---
license: apache-2.0
---
The dataset accompanying the paper
"Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models".
[Website] [paper] [GitHub].
Introduction
Welcome to the DeepFakeFace (DFF) dataset! Here we present a meticulously curated collection of artificial celebrity faces, crafted using cutting-edge diffusion models.
Our aim is to tackle the rising challenge posed by deepfakes in today's digital landscape.
Here are some example images in… See the full description on the dataset page: https://huggingface.co/datasets/OpenRL/DeepFakeFace.BRSpeech-DF
🗣️ BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese
🧩 Description
BRSpeech-DF is the first publicly available dataset for deepfake speech detection in Portuguese, covering both Brazilian and European variants.
It contains 459,000 audio samples, including both real and synthetic speech generated using multiple zero-shot text-to-speech (TTS) models.
This dataset aims to foster the development of more robust, inclusive, and multilingual audio deepfake… See the full description on the dataset page: https://huggingface.co/datasets/AKCIT-Deepfake/BRSpeech-DF.IndicTTS-Deepfake-Challenge-Data
IndicTTS Deepfake Detection Challenge
Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip.
🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data
This is the official dataset for the challenge and must be used for training and evaluation.
📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.NTIRE-RobustAIGenDetection-train
Training set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly indistinguishable from real photos in many cases, which creates serious challenges for trust, authenticity, forensics, and content… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-train.deepfake-detection-challenge
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/191fa07121/deepfake-detection-challenge.deepfake_detection_dataset_urdu
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset
This repository contains the Urdu Deepfake Audio Dataset introduced in the ACL 2024 paper "Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset".
The dataset focuses on two spoofing attacks – Tacotron and VITS TTS – and includes bonafide audio samples for comparison. The dataset construction ensures phonemic cover and balance, making it suitable for training… See the full description on the dataset page: https://huggingface.co/datasets/CSALT/deepfake_detection_dataset_urdu.deepfake-videos-dataset
DeepFake Videos for detection tasks
Dataset consists of 10,000+ files featuring 7,000+ people, providing a comprehensive resource for research in deepfake detection and deepfake technology. It includes real videos of individuals with AI-generated faces overlaid, specifically designed to enhance liveness detection systems.
By utilizing this dataset, researchers can advance their understanding of deepfake generation and improve the performance of detection methods. - Get the data… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/deepfake-videos-dataset.DeepfakeBenchCVQAD
MSU Compression Dataset Description
We developed LEHA-CVQAD dataset to evaluate full-reference and no-reference video quality metrics. Here we share the open part of the whole compression artifacts dataset (1,962 out of 6,240 videos). The hidden part is only available to benchmark-support personnel for testing metric performance. All videos are of mostly FullHD resolution, YUV420, and 10-15 seconds duration. Fps values are 24, 25, 30, 39, 50, and 60.
Subjective quality scores are… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/CVQAD.deepfake_video_datain-the-wild-deepfake
mohammedph197/in-the-wild-deepfake
Media collected by a deepfake-dataset pipeline, published for annotation.
One row per item, its media referenced by URL:
column
meaning
media_url
public URL of the file in this repo; the media is not distributed in the table
media_type
video, audio, image, or unknown
Files are content-addressed: a file's name is the SHA-256 of its bytes, so identical media appears once however many source records pointed at it.
popular-deepfakesdeepfake-celeba
MetaFLOS Deepfake Dataset (CelebA Face Deepfake Generated Images)
Flux2-Klein generated deepfake images from CelebA face descriptions, for deepfake detection / comparison research.
Content
19,867 generated images (full CelebA validation split)
Resolution: 256×256
Each corresponds to a CelebA real face (see real_orig field in manifest)
Generation style: snapshot/crop candid-photo look (not studio portrait) — off-center framing, subject possibly touching or cut by… See the full description on the dataset page: https://huggingface.co/datasets/tjw/deepfake-celeba.20K_real_and_deepfake_images_PCAThis dataset contains the test images used to evaluate our deepfake detection framework. It originally contained 20,000 real and deepfake images, but as some 2600 files are protected by the UK Crown and we do not have a permission to reproduced them, so these files were removed.
Our framework contains 4 machine learning models, which feed in the original images, error-level analysis (ELA) images, noise analysis (NA) images and Principal Component Analysis (PCA) images.
The models were created… See the full description on the dataset page: https://huggingface.co/datasets/ts0pwo/20K_real_and_deepfake_images_PCA.NTIRE-RobustAIGenDetection-val
Validation set for NTIRE 2026 Robust AI-Generated Image Detection in the Wild (updated)
Note: This is an updated version of the dataset. For challenge submissions, please make sure you use this version.
Robust AI-Generated Image Detection in the Wild Challenge is organized as a part of the New Trends in Image Restoration and Enhancement Workshop in conjunction with CVPR 2026.
Challenge overview
Text-to-image (T2I) models have made synthetic images nearly… See the full description on the dataset page: https://huggingface.co/datasets/deepfakesMSU/NTIRE-RobustAIGenDetection-val.nepali-audio-deepfake-datasetdeepfake-videoDeepfake-Eval-2024
Deepfake-Eval-2024: A Multi-Modal In-the-Wild Benchmark of Deepfakes Circulated in 2024
Deepfake-Eval-2024 is an in-the-wild deepfake dataset. Deepfake-Eval-2024 contains 44 hours of videos, 56.5 hours of audio, and 1,975 images, encompassing contemporary manipulation technologies, diverse media content, 88 different website sources, and 52 different languages. Deepfake-Eval-2024 contains manually labeled real and fake media. Deepfake-Eval-2024 is designed to facilitate deepfake… See the full description on the dataset page: https://huggingface.co/datasets/nuriachandra/Deepfake-Eval-2024.deepfake-and-real-imagesreal-vs-fake-human-voice-deepfake-audio
Deepfake Audio Dataset
Dataset contains 5,000 audio files, comprising both authentic human recordings and synthetic** AI-generated voice** samples. It designed for advanced research in deepfake detection, focusing on detecting fake voices and generated speech analysis. Specifically engineered to challenge voice authentication systems, it supports the development of robust models for real vs fake human voice recognition.
By utilizing this dataset, researchers and developers can… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/real-vs-fake-human-voice-deepfake-audio.Deepfake-Identity-Isolated-Dataset-PreP
Identity-Isolated Deepfake Face Images Dataset.
A rigorously preprocessed, identity-aware deepfake detection dataset of 142,837 face images, constructed for training generalized deepfake detectors across Face Swap and Entire Face Synthesis manipulation categories.
Built as part of the DFDS project - an open-source deepfake detection API for FinTech KYC identity verification. The full system is available on GitHub at DFDS-XAI
Dataset Summary.
Property
Value… See the full description on the dataset page: https://huggingface.co/datasets/ThinothW/Deepfake-Identity-Isolated-Dataset-PreP.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic… See the full description on the dataset page: https://huggingface.co/datasets/koyyalamudiraghavendra/deepfake-audio-detection.Deepfake-Eval-2024-audiodeepfake_face_classification
Deepfake Face Image Classification Dataset
This dataset is curated for the classification of deepfake face images, 16060 each real and fake images. It is derived from the DF40 dataset(only the test data), which includes 40 distinct deepfake techniques, facilitating the detection of state-of-the-art deepfakes and AI-generated content. (paperswithcode.com)
Dataset Overview
The dataset contains 32134 total images.
The dataset is divided into two categories:
Fake Images: 16… See the full description on the dataset page: https://huggingface.co/datasets/pujanpaudel/deepfake_face_classification.Deepfakeaudio
Dataset Card for "Deepfakeaudio"
More Information needed
AV-Deepfake1M-PlusPlus
AV-Deepfake1M++
The dataset used for the 2025 1M-Deepfakes Detection Challenge.
Task 1 Video-Level Deepfake Detection:
Given an audio-visual sample containing a single speaker, the task is to identify if the video is a deepfake or real.
Task 2 Deepfake Temporal Localization:
Given an audio-visual sample containing a single speaker, the task is to find out the timestamps [start, end] in which the manipulation is done.
The assumption here is that from the perspective of spreading… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M-PlusPlus.deepfake-ecg-small
ECG Dataset
This repository contains an small version of the ECG dataset: https://huggingface.co/datasets/deepsynthbody/deepfake_ecg, split into training, validation, and test sets. The dataset is provided as CSV files and corresponding ECG data files in .asc format. The ECG data files are organized into separate folders for the train, validation, and test sets.
Folder Structure
.
├── train.csv
├── validate.csv
├── test.csv
├── train
│ ├── file_1.asc
│ ├── file_2.asc… See the full description on the dataset page: https://huggingface.co/datasets/deepsynthbody/deepfake-ecg-small.
