datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AV-Deepfake1M
AV-Deepfake1M
This is the official repository for the paper
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset.
Abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most
advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting
high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.AV-Deepfake1M-PlusPlus
AV-Deepfake1M++
The dataset used for the 2025 1M-Deepfakes Detection Challenge.
Task 1 Video-Level Deepfake Detection:
Given an audio-visual sample containing a single speaker, the task is to identify if the video is a deepfake or real.
Task 2 Deepfake Temporal Localization:
Given an audio-visual sample containing a single speaker, the task is to find out the timestamps [start, end] in which the manipulation is done.
The assumption here is that from the perspective of spreading… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M-PlusPlus.LAV-DF
Localized Audio Visual DeepFake Dataset (LAV-DF)
This repo is the dataset for the DICTA paper Do You Really Mean That? Content Driven Audio-Visual
Deepfake Dataset and Multimodal Method for Temporal Forgery Localization
(Best Award), and the journal paper "Glitch in the Matrix!": A Large Scale Benchmark for Content Driven Audio-Visual
Forgery Detection and Localization submitted to CVIU.
LAV-DF Dataset
Download
To use this LAV-DF dataset, you should… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/LAV-DF.
