datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.drawvla-prompt-validation-clean
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-clean.Prompt_Tuning_Datasets_with_Foreground
⭐ Dataset Introduction
The standard datasets (except ImageNet) used for CLIP-based Prompt Tuning research (e.g., CoOp).
Based on the original datasets, this repository adds foreground segmentation masks (generated by SEEM) of all raw images.
For the foreground masks, the RGB value of the foreground region is [255, 255, 255], and the background region is [0, 0, 0].
The shorter side is always fixed to 512 px, and the scaling ratio is the same as that of the… See the full description on the dataset page: https://huggingface.co/datasets/JREion/Prompt_Tuning_Datasets_with_Foreground.drawvla-prompt-validation
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation.drawvla-prompt-validation-v3
DrawVLA — Sketch-Prompt Validation
Circle (which) + arrow (where) + caption (what) visual instructions overlaid on
LIBERO observations, each labelled with a binary
verdict for training a prompt validator or a self-checking VLA:
right — every channel is correct and exactly one reading survives; execute.
wrong — a channel is incorrect or the deictic prompt remains under-determined;
reject. Formerly ambiguous prompts are retained in this class.
All captions are name-free L2/L3… See the full description on the dataset page: https://huggingface.co/datasets/shibuina/drawvla-prompt-validation-v3.image-prompt-injection
Image-Based Prompt Injection Dataset
Synthetic dataset for prompt injection attacks on Large Vision-Language Models (LVLMs).
Team
Name
GitHub
Ritik Sinha
@Ritik1207-ind
Siddhant Kumar
@siddhantkumar101
Udit Dadhich
@UditDadhich
GitHub Repository: prompt-injection-attacks-on-LVLMS
Dataset Details
4,859 labeled samples
4 attack types: typographic, structural, adversarial, metadata
4 injection goals: jailbreak, exfiltration… See the full description on the dataset page: https://huggingface.co/datasets/Reet1207/image-prompt-injection.Civitai-2m-prompts[EDIT]
This dataset is now obsolete, you should use AdamCodd/Civitai-8m-prompts instead, as it includes all the contents of this dataset along with improved data preprocessing.
A subset of hanruijiang/civitai-stable-diffusion-2.5m database from Civitai API containing only hash, url, nsfwLevel, nsfw, stats, prompt, negativePrompt fields. All items of the dataset have a prompt/negativePrompt populated. It contains exactly 2129933 items.
If you want to support me, you can here.
AI_Assisted_Self_Images_With_Prompts_And_Personality_Tests
Digital Mirror of the Soul - AI-Assisted Self-Images with Prompts and Psychological Questionnaires
This dataset originates from a study that examines the intersection of artificial intelligence, psychology, and art. It provides a comprehensive collection of AI-generated images and textual prompts from participants engaging in a task designed to express their self-image. This work is ideal for researchers in the fields of clinical and art psychology, and data science, offering a… See the full description on the dataset page: https://huggingface.co/datasets/Hipnotalamusz/AI_Assisted_Self_Images_With_Prompts_And_Personality_Tests.instagram-fashion-prompts-v1
Instagram Fashion Prompts v1
Instagram Fashion Prompts v1 is a small, high–quality image dataset built for training and testing fashion and lifestyle models for social media content.
This first version contains:
15 high–resolution images (portrait & full–body)
1 digital muse / model
Multiple locations: city streets at night, Paris rooftops, beaches, desert, rooftops, luxury interiors
Outfits: elegant evening gowns, latex dresses, floral dresses, bikinis, sporty outfits and… See the full description on the dataset page: https://huggingface.co/datasets/Octavian-labs/instagram-fashion-prompts-v1.prompt2model-examples
Prompt2Model Toy Examples
Product: Prompt2Model:
a language-guided vision model factory. A typed pipeline (prompt, dataset config, training,
calibration/conformal abstain, ONNX export, an optional distill/quantize step with an
accuracy-floor gate, and a hard-case flywheel).
What this is (and isn't)
This is not a benchmark dataset. Prompt2Model has no natural "own" benchmark corpus the way a
task-specific product does. What's uploaded here is the repository's own… See the full description on the dataset page: https://huggingface.co/datasets/Dhi-Technologies/prompt2model-examples.text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/bofbofai/text-to-image-prompts.PromptedArtistIdentificationDataset
Prompted Artist Identification Dataset
Website | Paper | GitHub
Identifying Prompted Artist Names from Generated Images
Grace Su, Sheng-Yu Wang, Aaron Hertzmann, Eli Shechtman, Jun-Yan Zhu, Richard Zhang
arXiv, 2025
Prompted Artist Identification Benchmark. We introduce the first large-scale benchmark for identifying prompted artist names from generated images. The benchmark covers four axes of generalization that match realistic use cases: (1) Artists: we collect artists… See the full description on the dataset page: https://huggingface.co/datasets/cmu-gil/PromptedArtistIdentificationDataset.PromptedArtistIdentificationDataset-ViewSamples
Sample Dataset Viewer for Prompted Artist Identification Dataset
Website | Paper | GitHub
Identifying Prompted Artist Names from Generated Images
Grace Su, Sheng-Yu Wang, Aaron Hertzmann, Eli Shechtman, Jun-Yan Zhu, Richard Zhang
arXiv, 2025
Description
This page serves as a viewer for sample images from the Prompted Artist Identification Dataset.
Please visit the main dataset page for a description of the full dataset.
The entire benchmark dataset consists of 1.95… See the full description on the dataset page: https://huggingface.co/datasets/cmu-gil/PromptedArtistIdentificationDataset-ViewSamples.
