datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sd3_5_fine_sixcard
DreamBooth training example
DreamBooth is a method to personalize text2image models like stable diffusion given just a few(3~5) images of a subject.
The train_dreambooth.py script shows how to implement the training procedure and adapt it for stable diffusion.
Running locally with PyTorch
Installing the dependencies
Before running the scripts, make sure to install the library's training dependencies:
Important
To make sure you can successfully run the latest… See the full description on the dataset page: https://huggingface.co/datasets/yyyzzzzyyy/sd3_5_fine_sixcard.sd3-tssd3_5_dataset700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Preference_Dataset
Rapidata Image Generation Preference Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3.Flux_SD3_MJ_Dalle_Human_Alignment_Dataset
NOTE: A newer version of this dataset is available Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Alignment_Dataset
Rapidata Image Generation Alignment Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Coherence dataset: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
Link to the Preference dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Alignment_Dataset.sd3-images
Stable Diffusion 3 Images
A dataset of 1:1 images generated by Stable Diffusion 3, through glif.app.
Find the prompts in prompts.json, they correspond to the image based on number, for example the first element in the JSON array is x, then the image you're looking for is 0.jpg, and so on.
Prompts sourced from MohamedRashad/midjourney-detailed-prompts.
Data
You can find enhanced images by Gigapixel AI in the enhanced folder; these are the same 1024x1024 quality, but… See the full description on the dataset page: https://huggingface.co/datasets/leafspark/sd3-images.Flux_SD3_MJ_Dalle_Human_Coherence_Dataset
NOTE: A newer version of this dataset is available: Imagen3_Flux1.1_Flux1_SD3_MJ_Dalle_Human_Coherence_Dataset
Rapidata Image Generation Coherence Dataset
This Dataset is a 1/3 of a 2M+ human annotation dataset that was split into three modalities: Preference, Coherence, Text-to-Image Alignment.
Link to the Preference dataset: https://huggingface.co/datasets/Rapidata/700k_Human_Preference_Dataset_FLUX_SD3_MJ_DALLE3
Link to the Text-2-Image Alignment dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/Flux_SD3_MJ_Dalle_Human_Coherence_Dataset.DiNa-LRM-SD35m-HPSv3-Preprocess-Dataridgelora-cross-sensor-sd302d-f-to-m-20260825
RidgeLoRA-FP: SD302A-F to SD302D-M cross-sensor experiment
This public archive contains the leakage-controlled direct cross-sensor
experiment used to evaluate whether Stage-2 synthetic target-sensor images
help recognition on a physically different real sensor.
Locked protocol
Source/condition sensor: NIST SD302A device F.
Target sensor: NIST SD302D device M.
Identity: subject:finger-position; the same fingers exist across both
collections.
Subject split: 160… See the full description on the dataset page: https://huggingface.co/datasets/LamTNguyen/ridgelora-cross-sensor-sd302d-f-to-m-20260825.DMin_sd3_medium_lora_r4_caching_8846Implementation for "DMin: Scalable Training Data Influence Estimation for Diffusion Models".
Influence Function, Influence Estimation and Training Data Attribution for Diffusion Models.
Github, Paper
sd3_5_dposd3.5-pickapic-subsetautotree_automl_eye_movements_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_eye_movements_gosdt_l512_d3_sd3"
More Information needed
autotree_automl_default-of-credit-card-clients_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_default-of-credit-card-clients_gosdt_l512_d3_sd3"
More Information needed
autotree_automl_pol_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_pol_gosdt_l512_d3_sd3"
More Information needed
imagenet-latents-sd3.5-vae-e2e-lr2e-5-400ksd3_resultsautotree_automl_heloc_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_heloc_gosdt_l512_d3_sd3"
More Information needed
autotree_automl_MagicTelescope_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_MagicTelescope_gosdt_l512_d3_sd3"
More Information needed
sd3-controlnet-resultssd3.5_pickapicautotree_automl_electricity_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_electricity_gosdt_l512_d3_sd3"
More Information needed
autotree_automl_Diabetes130US_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_Diabetes130US_gosdt_l512_d3_sd3"
More Information needed
sd3-distilled-dataDSPart_SD3autotree_automl_covertype_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_covertype_gosdt_l512_d3_sd3"
More Information needed
flickr-sd3-realautotree_automl_credit_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_credit_gosdt_l512_d3_sd3"
More Information needed
autotree_automl_california_gosdt_l512_d3_sd3
Dataset Card for "autotree_automl_california_gosdt_l512_d3_sd3"
More Information needed
flickr-sd3-fake
