datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.gpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.messy-prompt-datasets
🎨 Messy Prompt Dataset
🎨 A mixed collection of AI image prompts (500+). A bit of everything — raw and uncurated. Truly open source: No login, no ads, no redirection. Just pure data for AI creators.
This project is a growing collection of diverse image generation prompts gathered from social platforms like Twitter/X. The entire dataset contains 500+ images, all structured into a comprehensive dataset.
Due to GitHub's limitations with large file storage, the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/messy-prompt-datasets.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.cyberseceval3-visual-prompt-injection
Dataset Card for CyberSecEval 3 - Visual Prompt Injection Benchmark
Dataset Details
Dataset Description
This dataset provides a multimodal benchmark for visual prompt injection, with text/image inputs. It is part of CyberSecEval 3, the third edition of Meta's flagship suite of security benchmarks for LLMs to measure cybersecurity risks and capabilities across multiple domains.
Language(s): English
License: MIT
Dataset Sources
Repository: Link… See the full description on the dataset page: https://huggingface.co/datasets/facebook/cyberseceval3-visual-prompt-injection.product-photography-v1-tiny-prompts-tasks-collage-filteredbad_prompt
Negative Embedding / Textual Inversion
Idea
The idea behind this embedding was to somehow train the negative prompt as an embedding, thus unifying the basis of the negative prompt into one word or embedding.
Side note: Embedding has proven to be very helpful for the generation of hands! :)
Usage
To use this embedding you have to download the file aswell as drop it into the "\stable-diffusion-webui\embeddings" folder.
Please put the embedding in… See the full description on the dataset page: https://huggingface.co/datasets/Nerfgun3/bad_prompt.NijiJourney-Prompt-Pairs
NijiJourney Prompt Pairs
A dataset containing txt2img prompt pairs for training on diffusion models
The final goal of this dataset is to create an OpenJourney like model but with NijiJourney images
All-Prompt-Jailbreakhle-no-img-prompt-completion-formatstream-prompt-upsample-runssdxl_images_easy_prompts-artists-seed1GimbalDiffusion-Prompt-Entanglement
GimbalDiffusion Prompt Entanglement Benchmark
The 380-sample benchmark used to measure pitch/content entanglement in GimbalDiffusion: Gravity-Aware Camera Control for Video Generation.
Project page
Paper
Poly Haven
Contents
prompt_entanglement.tar: the complete benchmark in one archive, unpacking directly into the layout below.
test_index.json and test_samples/: 380 prompts, seeds, camera matrices, intrinsics, requested pitch angles, and Poly Haven panorama… See the full description on the dataset page: https://huggingface.co/datasets/lefreud/GimbalDiffusion-Prompt-Entanglement.GEMRec-PromptBook
GEMRec-18k -- Prompt Book
This is the official image dataset for the paper Towards Personalized Prompt-Model Retrieval for Generative Recommendation.
Dataset Intro
GEMRec-18K is a prompt-model interaction dataset with 18K images generated by 200 publicly-available generative models paired with a diverse set of 90 textual prompts. We randomly sampled a subset of 197 models from the full set of models (all finetuned from Stable Diffusion) on Civitai according to the… See the full description on the dataset page: https://huggingface.co/datasets/MAPS-research/GEMRec-PromptBook.force-prompting-dataset-creationdhivehi-image-bbox-prompt
Dhivehi Image Bounding Box Prompt Dataset
This dataset, alakxender/dhivehi-image-bbox-prompt, contains 58,738 images annotated with COCO-style bounding boxes and Dhivehi (Thaana script) text, along with layout categories such as Text, Title, Picture, Caption, and Columns. It is designed for OCR, document layout analysis, and multimodal vision–language research focused on Dhivehi.
Dataset
Each row includes:
image — the RGB image (preserved original dimensions)
width… See the full description on the dataset page: https://huggingface.co/datasets/alakxender/dhivehi-image-bbox-prompt.stable-diffusion-prompts-stats-full-uncensoredPromptfixData
Model Scources
Dataset: https://huggingface.co/datasets/yeates/PromptfixData
Github: https://github.com/yeates/PromptFix
Paper: https://arxiv.org/pdf/2405.16785
Project Page: https://www.yongshengyu.com/PromptFix-Page/
Model Usage
The PromptFix dataset is intended solely for research purposes.
Please note that the PromptFix dataset is curated from open-source research projects and publicly available photo libraries. By using our dataset, you automatically agree to… See the full description on the dataset page: https://huggingface.co/datasets/yeates/PromptfixData.All-Prompt-Jailbreakeval2_all_promptsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval2_all_prompts.spro-optimized-prompts-fullvg150_train_sgg_promptThis repository contains the VG150 dataset transformed into datasets format with keys: "image_id", "image", "prompt_open", "prompt_close", "objects", and "relationships".
It was presented in the paper Compile Scene Graphs with Reinforcement Learning.
Code: https://github.com/gpt4vision/R1-SGG
All-Prompt-JailbreakPrompt_Protections
Protection Protections
This dataset contains a number of snippets and short extentions to add to the system prompt of bots and GPTs to persuade the model not to reveal it's instructions to the user. It's not a perfect solution, but sometimes a little clever prompting is all you need :)
All-Prompt-JailbreakAll-Prompt-Jailbreaksd-prompt-image-in-the-wild-counterfeitfr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "fr3",
"total_episodes": 301,
"total_frames": 276017,
"total_tasks": 2,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 60,
"splits": {
"train": "0:301"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/yio-ye2004/fr3_pickplace_extended_new_cmd_SYNC_part_corrected_with_prompt_test.
