datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
civitaiCivitAI-Flux-PromptsThis is a dataset consisting of prompts from the top 50k most reacted and most commented images from CivitAI.
The only used images are images generated by models utilizing the T5 Text Encoder, thus mostly using natural prose or something close to that.
Those prompts were sanitized and the shorter ones removed. 150 of the most often reoccuring (quality tags, unnecessary names, etc...) tags have been removed. This resulted in 7.5k images and prompts in total.
For those 7.5k images, short natural… See the full description on the dataset page: https://huggingface.co/datasets/Aconexx/CivitAI-Flux-Prompts.civitai_top_10000_imagesCivitAI-SD-Prompts
CivitAI SD Prompts
Dataset is a collection of both synthetic and organic descriptions of images with prompts for Stable Diffusion series of models associated to them.
It works best with SDXL, but should probably work with others too.
Synthetic part
"Synthetic" means that the end result of the sample (SD prompt) was produced by a large language model (Claude Opus).
Select characters from CharGen v2 datasets were used as prompts for Claude Opus to generate their appearance… See the full description on the dataset page: https://huggingface.co/datasets/NewEden/CivitAI-SD-Prompts.civitai_top10kCivitAI-Prompts-Sharegpt
