datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lexica-stable-diffusion-v1-5
Stable Diffusion Dataset
This is a set of about 80,000 Image-Prompt pairs generated by stable-diffusion-v1-5.
The Prompts come from dataset Stable-Diffusion-Prompts which filtered and extracted from the image finder for Stable Diffusion: "Lexica.art".
lexica_dataset
LexicaDataset
LexicaDataset is a large-scale text-to-image prompt dataset shared in [USENIX'24] Prompt Stealing Attacks Against Text-to-Image Generation Models.
It contains 61,467 prompt-image pairs collected from Lexica.
All prompts are curated by real users and images are generated by Stable Diffusion.
Data collection details can be found in the paper.
Data Splits
We randomly sample 80% of a dataset as the training dataset and the rest 20% as the testing dataset.… See the full description on the dataset page: https://huggingface.co/datasets/vera365/lexica_dataset.lexica_6kMore Information needed
lexical_diff_bangla_assamese_v2
Overview
The dataset is composed of 1000 images containing Bangla and Assamese text.
Bangla and Assamese are closely related languages with sharing the
Bengali–Assamese script and have similar lexical
constructs (https://en.wikipedia.org/wiki/Bengali%E2%80%93Assamese_script#cite_note-MajR-22).
The text is this dataset has been generated from Open English Bible
(https://openenglishbible.org/oeb/2022.1/OEB-2022.1-US.txt) by:
Selecting the first 500 lines containg between 50 and… See the full description on the dataset page: https://huggingface.co/datasets/anrikus/lexical_diff_bangla_assamese_v2.lexical_diff_bangla_assamese
Overview
The dataset is composed of 1000 images containing Bangla and Assamese text.
Bangla and Assamese are closely related languages with sharing the
Bengali–Assamese script and have similar lexical
constructs (https://en.wikipedia.org/wiki/Bengali%E2%80%93Assamese_script#cite_note-MajR-22).
The text is this dataset has been generated from Open English Bible
(https://openenglishbible.org/oeb/2022.1/OEB-2022.1-US.txt) by:
Selecting the first 500 lines containg between 50 and… See the full description on the dataset page: https://huggingface.co/datasets/anrikus/lexical_diff_bangla_assamese.
