datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
DALL-E-Prompts-OpenAI-ChatGPT
Dataset Card for Dataset Name
Dataset Summary
This dataset has been generated using Prompt Generator for OpenAI's DALL-E.
Languages
English
Dataset Structure
1.000.000 Prompts
midjourney-prompts-highquality
Thank you to the Akash Network for sponsoring this project and providing A100s/H100s for compute!
About
A filtered version of the vivym/midjourney-prompts dataset
Filtering criteria
top 10% in length (assuming that longer prompts = more effort and higher quality)
used on an image to be upscaled (assuming that users are more likely to upscale an image that is aesthetically pleasing)
used on midjourney version 5.0+
deduplicated
Run yourself
filter.py script… See the full description on the dataset page: https://huggingface.co/datasets/gaodrew/midjourney-prompts-highquality.PromptedArtistIdentificationDataset
Prompted Artist Identification Dataset
Website | Paper | GitHub
Identifying Prompted Artist Names from Generated Images
Grace Su, Sheng-Yu Wang, Aaron Hertzmann, Eli Shechtman, Jun-Yan Zhu, Richard Zhang
arXiv, 2025
Prompted Artist Identification Benchmark. We introduce the first large-scale benchmark for identifying prompted artist names from generated images. The benchmark covers four axes of generalization that match realistic use cases: (1) Artists: we collect artists… See the full description on the dataset page: https://huggingface.co/datasets/cmu-gil/PromptedArtistIdentificationDataset.
