datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
text-to-image-promptsIf you have questions about this dataset , feel free to ask them on the fusion-discord : https://discord.gg/8TVHPf6Edn
This collection contains sets from the fusion-t2i-ai-generator on perchance.
This datset is used in this notebook: https://huggingface.co/datasets/codeShare/text-to-image-prompts/tree/main/Google%20Colab%20Notebooks
To see the full sets, please use the url "https://perchance.org/" + url
, where the urls are listed below:
_generator
gen_e621
fusion-t2i-e621-tags-1… See the full description on the dataset page: https://huggingface.co/datasets/codeShare/text-to-image-prompts.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Citation
@article{zou2026advancing,
title = {Advancing Aesthetic Image Generation via Composition Transfer},
author = {Zou, Kai and Zhao, Zhiwei and Liu, Bin and Yu, Nenghai},
journal = {International Journal of Computer Vision},
volume = {134},
pages = {252},
year = {2026},
doi = {10.1007/s11263-026-02862-8},
url = {https://doi.org/10.1007/s11263-026-02862-8}… See the full description on the dataset page: https://huggingface.co/datasets/jackyhate/text-to-image-2M.surya-ocr-500-image-to-textsurya-ocr-1K-image-to-textText_to_Image
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
midjourney-texttoimage
Dataset Card for Midjourney User Prompts & Generated Images (250k)
Dataset Summary
General Context
Midjourney is an independent research lab whose broad mission is to "explore new mediums of thought". In 2022, they launched a text-to-image service that, given a natural language prompt, produces visual depictions that are faithful to the description. Their service is accessible via a public Discord server, where users interact with a Midjourney bot. When issued… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/midjourney-texttoimage.text-to-image-prompts
The dataset of the most popular text-to-image prompts.
Dataset Details
Dataset Description
Curated by: kazimir.ai
Funded by [optional]: [More Information Needed]
Shared by [optional]: https://kazimir.ai
License: apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Free to use.
Dataset Structure
CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.midjourney-texttoimage-new
Dataset Card for Midjourney User Prompts & Generated Images (250k)
Dataset Summary
General Context
Midjourney is an independent research lab whose broad mission is to "explore new mediums of thought". In 2022, they launched a text-to-image service that, given a natural language prompt, produces visual depictions that are faithful to the description. Their service is accessible via a public Discord server, where users interact with a Midjourney bot. When issued… See the full description on the dataset page: https://huggingface.co/datasets/nateraw/midjourney-texttoimage-new.text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image
model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers
part_id 1-20 (20,000 source images) before filtering. The same content is
also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")
Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.code-image-to-text
Code Snippet Image → Text
A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of
transcribing an image of a code snippet back into its source text — syntax-aware OCR.
Each example pairs a syntax-highlighted PNG of code with the exact code text that
produced it. It spans 8 programming languages and deliberately mixes two capture types:
block — a complete function / unit (6–45 lines).
fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.Detonate_Text_To_Imageghibli-TextToImageX2I-text-to-image
X2I Dataset
Project Page: https://vectorspacelab.github.io/OmniGen/
Github: https://github.com/VectorSpaceLab/OmniGen
Paper: https://arxiv.org/abs/2409.11340
Model: https://huggingface.co/Shitao/OmniGen-v1
To achieve robust multi-task processing capabilities, it is essential to train the OmniGen on large-scale and diverse datasets. However, in the field of unified image generation, a readily available dataset has yet to emerge. For this reason, we have curated a large-scale… See the full description on the dataset page: https://huggingface.co/datasets/yzwang/X2I-text-to-image.text_to_imagefashion_text_to_image
annotations_creators:
- machine-generated
language:
- en
language_creators:
- other
multilinguality:
- monolingual
pretty_name: "Fashion captions"
size_categories:
- n<100K
tags: []
task_categories:
- text-to-image
task_ids: []
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/duyngtr16061999/fashion_text_to_image.trending-text-to-image
CivitAI Improved Prompts Dataset
This dataset contains trending AI-generated images from CivitAI with Flux-improved prompts for better generation results.
Dataset Format (JSONL)
Each line contains a JSON object with:
id: Original image ID from CivitAI
improved_prompt: Flux-enhanced version of the prompt
category: Automatically determined theme category
All original CivitAI metadata including:
Original prompt and negative prompt
Model information
Image URL and… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/trending-text-to-image.diagram_image_to_text
Dataset Card for "diagram_image_to_text"
More Information needed
dior_text_to_imageUltra-Prompts-Text-To-Imagetext-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/NovaBrown/text-to-image-2M.synthetic-text-to-image-prompts-100k
Synthetic Text-to-Image Prompt Pairs
Platform-neutral prompts. No claim of conversion or output-quality performance.
Dataset summary
Rows: 100,000
Columns: 14
Train/validation/test: 80,000 / 10,000 / 10,000
Synthetic: yes, every row
Generation seed: 550031
Intended uses
Model prototyping, pipeline testing, schema experiments, and educational demonstrations.
Limitations
This dataset is synthetic and must not be represented as… See the full description on the dataset page: https://huggingface.co/datasets/synthdataq9x260918/synthetic-text-to-image-prompts-100k.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.Chemistry_text_to_image
Dataset Card for "Chemistry_text_to_image"
More Information needed
TREC-2023-Text-to-Image
Dataset Card for "TREC-2023-Text-to-Image"
More Information needed
visdrone-text-to-imageText-To-Image-Test-Prompts
Text To Image Test Prompt Library
A comprehensive collection of evaluation prompts for testing text-to-image AI models across diverse parameters and use cases.
Overview
This repository contains a structured set of test prompts designed to evaluate the capabilities of text-to-image generation models. Rather than focusing on formal evaluation metrics, these prompts are intended for end users who want to test how well a model might perform for their specific use cases.… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Text-To-Image-Test-Prompts.llama-3.2-random-images-to-textThis dataset contains 39020 unique image anotated pairs.
The images in the dataset are copywrited. and should only be used in acordance with the copywrite law.
For example I belive training LLM falls under fair use policies. Check with your lawyers.
Each image is anotated by the Meta llama 3.2 vison model. This first upload is part of the larger 5,000,000 image dataset.
The parquet files have the actual image inside them savd in raw bytes. they have the responce from the llm and also a unique… See the full description on the dataset page: https://huggingface.co/datasets/mylesgoose/llama-3.2-random-images-to-text.aid-text-to-imagediffusion.4.text_to_image
Dataset Card for "diffusion.4.text_to_image"
More Information needed
Chemistry_text_to_image_BASE64
