datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amateur_drawings-controlnet-dataset
Dataset Card for "amateur_drawings-controlnet-dataset"
WIP... Come back later....
ControlNetsAV-Deepfake1M
AV-Deepfake1M
This is the official repository for the paper
AV-Deepfake1M: A Large-Scale LLM-Driven Audio-Visual Deepfake Dataset.
Abstract
The detection and localization of highly realistic deepfake audio-visual content are challenging even for the most
advanced state-of-the-art methods. While most of the research efforts in this domain are focused on detecting
high-quality deepfake images and videos, only a few works address the problem of the localization of small… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M.WT-Data-Project
WT-DATA-PROJECT.DATA
Data collected in wt-data-project.
Repository
Info
wt-data-project.data
wt-data-project.web
wt-data-project.visualization
Visualization
The visualization part is in this webpage.
Other links… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/WT-Data-Project.VIDIT-Depth-ControlNet
VIDIT Dataset
This is a version of the VIDIT dataset equipped for training ControlNet using depth maps conditioning.
VIDIT includes 390 different Unreal Engine scenes, each captured with 40 illumination settings, resulting in 15,600 images. The illumination settings are all the combinations of 5 color temperatures (2500K, 3500K, 4500K, 5500K and 6500K) and 8 light directions (N, NE, E, SE, S, SW, W, NW). Original image resolution is 1024x1024.
We include in this version only the… See the full description on the dataset page: https://huggingface.co/datasets/Nahrawy/VIDIT-Depth-ControlNet.controlnet-testingcontrolnetopen_pose_controlnet
Dataset for training controlnet models with conditioning images as Human Pose
the entries have been taken from this dataset
ptx0/photo-concept-bucket
the open pose images have been generated with
controlnet_aux
for the scripts to download the files, generate the open pose and the dataset please refer to:
raulc0399/dataset_scripts
AV-Deepfake1M-PlusPlus
AV-Deepfake1M++
The dataset used for the 2025 1M-Deepfakes Detection Challenge.
Task 1 Video-Level Deepfake Detection:
Given an audio-visual sample containing a single speaker, the task is to identify if the video is a deepfake or real.
Task 2 Deepfake Temporal Localization:
Given an audio-visual sample containing a single speaker, the task is to find out the timestamps [start, end] in which the manipulation is done.
The assumption here is that from the perspective of spreading… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/AV-Deepfake1M-PlusPlus.mscoco-controlnet-cannyVIDIT-FAID-Depth-ControlNet
Dataset Card for "VIDIT-FAID-Depth-ControlNet"
More Information needed
fluxdev_controlnet_16kcontrolnet-color-palette-20K
ControlNet Color Palette Dataset
This dataset contains resized images (512×512) and pixelated conditioning maps
generated using a 8×8 grid.
VIDIT-Depth-ControlNet-E
Dataset Card for "VIDIT-Depth-ControlNet-E"
More Information needed
controlnet-color-palette-20K-e1mscoco-controlnet-canny-less-colorsdw_pose_controlnet
Important Notice
This is a copy of raulc0399/open_pose_controlnet replacing openpose conditioning images
with DW pose.
Dataset for training controlnet models with conditioning images as Human Pose
the entries have been taken from this dataset
ptx0/photo-concept-bucket
the open pose images have been generated with
controlnet_aux
for the scripts to download the files, generate the openpose and the dataset please refer to:
raulc0399/dataset_scripts
Openpose to… See the full description on the dataset page: https://huggingface.co/datasets/dimitribarbot/dw_pose_controlnet.FAID-Depth-ControlNet
A Dataset of Flash and Ambient Illumination Pairs from the Crowd
This is a version of the A Dataset of Flash and Ambient Illumination Pairs from the Crowd dataset equipped for training ControlNet using depth maps conditioning.
The dataset includes 2775 pairs of flash light and ambient light images. It includes images of people, shelves, plants, toys, rooms and objects.
Captions were generated using the BLIP-2, Flan T5-xxl model.
Depth maps were generated using the GLPN fine-tuned on… See the full description on the dataset page: https://huggingface.co/datasets/Nahrawy/FAID-Depth-ControlNet.Fashion_controlnet_dataset_V3
Dataset Card for "Fashion_controlnet_dataset_V3"
More Information needed
controlnet-color-palette-7K
ControlNet Color Palette Dataset
This dataset contains resized images (768×768) and pixelated conditioning maps
generated using a 8×8 grid.
APT-36K-poses-controlnet-dataset
Dataset Card for "APT-36K-poses-controlnet-dataset"
More Information needed
sample_controlnet_dataset
ControlNet training
this dataset is subset of fill_50k dataset just to test the finetuning logic.
TODO:
add text data
Controlnet_dart_v2_sft_img_resizecontrolnet-color-palette-5K
ControlNet Color Palette Dataset
This dataset contains resized images (768×768) and pixelated conditioning maps
generated using a 8×8 grid.
LAV-DF
Localized Audio Visual DeepFake Dataset (LAV-DF)
This repo is the dataset for the DICTA paper Do You Really Mean That? Content Driven Audio-Visual
Deepfake Dataset and Multimodal Method for Temporal Forgery Localization
(Best Award), and the journal paper "Glitch in the Matrix!": A Large Scale Benchmark for Content Driven Audio-Visual
Forgery Detection and Localization submitted to CVIU.
LAV-DF Dataset
Download
To use this LAV-DF dataset, you should… See the full description on the dataset page: https://huggingface.co/datasets/ControlNet/LAV-DF.MIIW-Depth-ControlNet
Dataset Card for "MIIW-Depth-ControlNet"
More Information needed
controlnet_datainstructpix2pix-controlnet
Dataset Summary
This dataset is designed for training models to follow edit instructions on images.
Data Structure
Each sample contains:
original_prompt: the initial text prompt.
original_image: image generated with the Stable Diffusion model using the original_prompt.
edit_prompt: textual instruction describing the modification.
edited_prompt: reformulated prompt used for the generation of edited_image.
edited_image: image generated with the Stable Diffusion… See the full description on the dataset page: https://huggingface.co/datasets/iamlucaconti/instructpix2pix-controlnet.seait_ControlNet1-1-modules-safetensorsThis is the model files for ControlNet 1.1.
This model card will be filled in a more detailed way after 1.1 is officially merged into ControlNet.
dog-poses-controlnet-dataset
Dataset Card for "dog-poses-controlnet-dataset"
More Information needed
