datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/GokuScraper/seedance-2-prompts-datasets.gpt-image-2-prompts-datasets
🖼️ GPT Image 2 Prompt Dataset
🖼️ The ultimate GPT Image 2 prompt dataset (5GB+). 15,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for OpenAI's GPT Image 2 model and the resulting generated images. The entire dataset exceeds 5GB and contains 15,000+ images, all structured into a comprehensive dataset.
Due to… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/gpt-image-2-prompts-datasets.nano-banana-pro-prompts-datasets
🖼️ Nano Banana Pro Prompt Dataset
🖼️ The ultimate Nano Banana Pro prompt dataset (6GB+). 26,000+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for Nano Banana Pro AI image model and the resulting generated images. The entire dataset exceeds 6GB and contains 26,000+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/nano-banana-pro-prompts-datasets.veo3-video-prompts
Veo 3 Video Generation Dataset
English | Português do Brasil
English
Summary
A collection of AI-generated videos created with Google's Veo 3 family of models. Each record contains the original text prompt, the model variant used, the generated video, and (when applicable) the input reference image. Videos are organized into one configuration per model variant.
Videos: 5,811
Input images: 1,354
Configurations: 6
Language of prompts: multilingual… See the full description on the dataset page: https://huggingface.co/datasets/artificialguybr/veo3-video-prompts.product-photography-v1-tiny-prompts-tasks-collage-filteredsdxl_images_easy_prompts-artists-seed1stable-diffusion-prompts-stats-full-uncensoredepisodes
My Weird Prompts - Episode Dataset
The production record of every episode of the
My Weird Prompts podcast: the transcript, links to
the published episode, a description of the prompt that started it, and the
generation telemetry for how it was made - model, pipeline version, GPU, timings
and compute cost.
5,393 episodes. Synced daily from the production database.
from datasets import load_dataset
ds = load_dataset("My-Weird-Prompts/episodes", split="train")
Which… See the full description on the dataset page: https://huggingface.co/datasets/My-Weird-Prompts/episodes.spro-optimized-prompts-fulleval2_all_promptsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval2_all_prompts.images-promptsSpeech-To-Text-System-Prompts-2
Speech To Text System Prompt Library
This repository provides a collection of system prompts designed to transform and refine text captured using speech-to-text technologies.
By passing STT outputs through large language models with these specialized prompts, you can achieve cleaner, more structured, and purpose-specific text formats.
📋 The Idea
Here is the basic implementation. I don't pretend that this is the stuff of high AI engineering. But it does create quite… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Speech-To-Text-System-Prompts-2.midjourney-prompts
midjourney-prompts
Description
This dataset contains the cleaned midjourney prompts from Midjourney.
Total prompts: 9,085,397
Version
Count
5.2
2,272,465
5.1
2,060,106
5.0
3,530,770
4.0
1,204,384
3.0
14,991
2.0
791
1.0
1,239
Style
Count
default
8,874,181
raw
177,953
expressive
27,919
scenic
2,146
cute
2,036
original
511
1980s-photo-prompts
1980s Photo Prompts
Ten image-to-image prompts for the viral "1980s photo" trend, each paired with the image it actually produced on the first attempt. Generated on September 16, 2026 with Claude Imagine using GPT Image 2.5 and Nano Banana 2. No retouching, no cherry-picking.
Every prompt follows the same rule: keep my face, facial features and skin tone exactly the same, change everything else to 1985. The identity line comes first, then a specific year, then hair, clothes… See the full description on the dataset page: https://huggingface.co/datasets/claudeimagine/1980s-photo-prompts.image-prompts-rawawesome-text2video-prompts
Rapidata Video Generation Preference Dataset
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset contains prompts for video generation for 14 different categories. They were collected with a combination of manual prompting and ChatGPT 4o. We provide one example sora video generation for each video.
Overview
Categories and Comments
Object Interactions Scenes: Basic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/awesome-text2video-prompts.cc_news_promptsourceeval2_all_prompts_fixed_266This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval2_all_prompts_fixed_266.midjourney-detailed-prompts
Midjourney: Detailed Prompts
This dataset is my attempt in providing a high quality text-to-image dataset with detailed and several levels of prompting for images.
Hope it helps anyone in his research ^^
Thanks goes to ...
midjourney-images dataset
Qwen-VL-Max for descriping images in huge detail.
Command R for long & short prompt generation
sdxl_images_sb_prompts-multi_artist-seed1Shakespearean-Text-Transformation-Prompts
Shakespeare GPT (Shakespearean Text Generation Prompts)
Welcome to what might be the internet's largest collection of prompts for rewriting text in Shakespearean English! This repository contains a variety of prompts designed to transform modern text into the style of Shakespeare, organized by format and purpose.
These prompts can be used with any AI tool that accepts custom instructions. A user interface may be forthcoming for those who feel the need to do this regularly.… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Shakespearean-Text-Transformation-Prompts.awesome_hunyuanImage_prompts
Awesome HunyuanImage Prompts
Built with Hugging Face AI Sheets
Do you want to master HunyuanImage 3, one of the best open models for generating images? This resource is for you.
HunyuanImage 3's Prompt Handbook is a great resource for learning how to prompt the model. Unfortunately, it is in Chinese, so I have created this resource for the open community.
It contains:
All prompts in HunyuanImage 3's Prompt Handbook are organized by category.
Their translation into English… See the full description on the dataset page: https://huggingface.co/datasets/dvilasuero/awesome_hunyuanImage_prompts.parti-prompts-sdxl-1.0
Dataset Card for "parti-promtps-sdxl-1.0"
More Information needed
seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/zzyyun/seedance-2-prompts-datasets.editing_promptsUR5e_metaquest_rubberduck_varying-prompts_full-mixture_phase0This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.image": {
"dtype": "image",
"shape": [
224,
224,
3
],
"names": [
"height",
"width",
"channels"
]
},
"observation.images.wrist_image": {… See the full description on the dataset page: https://huggingface.co/datasets/nsimonato25/UR5e_metaquest_rubberduck_varying-prompts_full-mixture_phase0.DALL-E-Prompts-OpenAI-ChatGPT
Dataset Card for Dataset Name
Dataset Summary
This dataset has been generated using Prompt Generator for OpenAI's DALL-E.
Languages
English
Dataset Structure
1.000.000 Prompts
UR5e_metaquest_rubberduck_varying-prompts_full-mixture_phase0-1-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.images.image": {
"dtype": "image",
"shape": [
224,
224,
3
],
"names": [
"height",
"width",
"channels"
]
},
"observation.images.wrist_image": {… See the full description on the dataset page: https://huggingface.co/datasets/nsimonato25/UR5e_metaquest_rubberduck_varying-prompts_full-mixture_phase0-1-2.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/yemsrach3723/seedance-2-prompts-datasets.seedance-2-prompts-datasets
🎞️ Seedance-2-prompts-datasets
🎞️ The ultimate Seedance-2 video prompt dataset (50GB+). 8100+ video generation prompts with full metadata and preview frames. Truly open source: No login, no ads, no redirection. Just pure data for AI video creators.
This project is a massive collection of prompts used for Bytedance's Seedance 2.0 and the resulting generated videos. The entire dataset exceeds 50GB and contains 8100+ videos, all structured into a comprehensive dataset.
Due… See the full description on the dataset page: https://huggingface.co/datasets/Jason990/seedance-2-prompts-datasets.
