KavinduHansaka/prompt-gen-10k-flux-sdxl
Prompt Generation Dataset (10K Narrative for Flux / SDXL) This dataset (prompt_gen_final_10k.jsonl and prompt_gen_final_10k.csv) was used to train and fine-tune image-prompt models such as KavinduHansaka/Llama-3.2-1B-ImageGen. It contains 10,000 curated narrative prompt samples designed for image generation models like Stable Diffusion XL and Flux.Unlike raw tag-based datasets, the target field provides natural paragraphs (≈80–100 words) that describe cinematic scenes with… See the full description on the dataset page: https://huggingface.co/datasets/KavinduHansaka/prompt-gen-10k-flux-sdxl.
Prompt Generation Dataset (10K Narrative for Flux / SDXL)
This dataset (prompt_gen_final_10k.jsonl and prompt_gen_final_10k.csv) was used to train and fine-tune image-prompt models such as KavinduHansaka/Llama-3.2-1B-ImageGen.
It contains 10,000 curated narrative prompt samples designed for image generation models like Stable Diffusion XL and Flux. Unlike raw tag-based datasets, the target field provides natural paragraphs (≈80–100 words) that describe cinematic scenes with attributes (lighting, textures, mood, etc.) woven into the prose.
Dataset Format
The dataset includes the following fields:
input– short seed description or raw captionstyle– original style stringstyle_normalized– normalized style category (cinematic, realism, fantasy, anime, etc.)negative– raw negative tags (original)negative_target– cleaned negative prompt text (separate field for conditioning)attributes– additional properties (lighting, mood, quality, etc.)aspect_ratio– target aspect ratio (original metadata, not intarget)width– image width in pixels (original metadata)height– image height in pixels (original metadata)target– final narrative prompt used for training (~80–100 words, cinematic style)_aug– marker for augmented samples (used to reach 10k)
Example
Usage
You can load the dataset directly with 🤗 Datasets:
from datasets import load_dataset
# Load JSONL
ds = load_dataset("KavinduHansaka/prompt-gen-10k-flux-sdxl", split="train")
print(ds[0])