deep-fake
AV-Deepfake1M
AV-Deepfake1M
This is the official repository for the paper
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset.
Abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most
advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting
high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.deepfake-audio-detection
Deepfake Audio Detection Dataset (v4)
Dataset Description
This dataset contains 1,866 audio samples (933 real, 933 synthetic) for training deepfake audio detection models. It is specifically designed for binary classification tasks to distinguish between authentic human speech and AI-generated synthetic audio.
What's New in v4
52% larger: Increased from 1,224 to 1,866 samples (642 new samples)
Expanded TTS coverage: Added Hume AI as 6th synthetic voice… See the full description on the dataset page: https://huggingface.co/datasets/garystafford/deepfake-audio-detection.Deepfake
DeepGuard Deepfake Dataset
A paired deepfake dataset for training and benchmarking deepfake detection models.
Generated as part of the DeepGuard AI Project (2026).
📥 2,301+ all-time downloads
Dataset Statistics
Fake images: 5426 (face-swapped using InsightFace inswapper_128)
Real images: 5426 (original paired faces)
Total: 10852 images
Format: JPEG, high quality (95%)
Generation method: InsightFace inswapper_128 (ONNX runtime)
Why This Dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Sowaiba01/Deepfake.DeepFakeFace---
license: apache-2.0
---
The dataset accompanying the paper
"Robustness and Generalizability of Deepfake Detection: A Study with Diffusion Models".
[Website] [paper] [GitHub].
Introduction
Welcome to the DeepFakeFace (DFF) dataset! Here we present a meticulously curated collection of artificial celebrity faces, crafted using cutting-edge diffusion models.
Our aim is to tackle the rising challenge posed by deepfakes in today's digital landscape.
Here are some example images in… See the full description on the dataset page: https://huggingface.co/datasets/OpenRL/DeepFakeFace.BRSpeech-DF
🗣️ BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese
🧩 Description
BRSpeech-DF is the first publicly available dataset for deepfake speech detection in Portuguese, covering both Brazilian and European variants.
It contains 459,000 audio samples, including both real and synthetic speech generated using multiple zero-shot text-to-speech (TTS) models.
This dataset aims to foster the development of more robust, inclusive, and multilingual audio deepfake… See the full description on the dataset page: https://huggingface.co/datasets/AKCIT-Deepfake/BRSpeech-DF.IndicTTS-Deepfake-Challenge-Data
IndicTTS Deepfake Detection Challenge
Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip.
🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data
This is the official dataset for the challenge and must be used for training and evaluation.
📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.
