stable-diffusion-3
Stable_Diffusion_3_RecaptionThis dataset is the one specified in the stable diffusion 3 paper which is composed of the ImageNet dataset and the CC12M dataset.
I used the ImageNet 2012 train/val data and captioned it as specified in the paper: "a photo of a 〈class name〉" (note all ids are 999,999,999)
CC12M is a dataset with 12 million images created in 2021. Unfortunately the downloader provided by Google has many broken links and the download takes forever.
However, some people in the community publicized the dataset.… See the full description on the dataset page: https://huggingface.co/datasets/gmongaras/Stable_Diffusion_3_Recaption.white_paper_holdout_3___stable-diffusion-xl-base-1.0FeatureCoding-StableDiffusion3.5ControlNet-AfterCodec@misc{pang2026largemodelfeaturecoding,
title={Towards Large Model Feature Coding},
author={Youwei Pang and Changsheng Gao and Dong Liu and Huchuan Lu and Weisi Lin},
year={2026},
eprint={2605.24025},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.24025},
}
FeatureCoding-StableDiffusion3.5Large@misc{pang2026largemodelfeaturecoding,
title={Towards Large Model Feature Coding},
author={Youwei Pang and Changsheng Gao and Dong Liu and Huchuan Lu and Weisi Lin},
year={2026},
eprint={2605.24025},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2605.24025},
}
LSDIR_stable_diffusion_3_fp16stabilityai-stable-diffusion-3-medium-diffusers_fp16_no_cpu
