datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Text_to_Image
Dataset Card
Dataset in ImagenHub.
Citation
Please kindly cite our paper if you use our code, data, models or results:
@article{ku2023imagenhub,
title={ImagenHub: Standardizing the evaluation of conditional image generation models},
author={Max Ku and Tianle Li and Kai Zhang and Yujie Lu and Xingyu Fu and Wenwen Zhuang and Wenhu Chen},
journal={arXiv preprint arXiv:2310.01596},
year={2023}
}
text-to-image-prompts
The dataset of the most popular text-to-image prompts.
Dataset Details
Dataset Description
Curated by: kazimir.ai
Funded by [optional]: [More Information Needed]
Shared by [optional]: https://kazimir.ai
License: apache-2.0
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]
Uses
Free to use.
Dataset Structure
CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Kazimir-ai/text-to-image-prompts.text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image
model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers
part_id 1-20 (20,000 source images) before filtering. The same content is
also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")
Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.code-image-to-text
Code Snippet Image → Text
A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of
transcribing an image of a code snippet back into its source text — syntax-aware OCR.
Each example pairs a syntax-highlighted PNG of code with the exact code text that
produced it. It spans 8 programming languages and deliberately mixes two capture types:
block — a complete function / unit (6–45 lines).
fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.Detonate_Text_To_Imageghibli-TextToImagetext_to_imagefashion_text_to_image
annotations_creators:
- machine-generated
language:
- en
language_creators:
- other
multilinguality:
- monolingual
pretty_name: "Fashion captions"
size_categories:
- n<100K
tags: []
task_categories:
- text-to-image
task_ids: []
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/duyngtr16061999/fashion_text_to_image.trending-text-to-image
CivitAI Improved Prompts Dataset
This dataset contains trending AI-generated images from CivitAI with Flux-improved prompts for better generation results.
Dataset Format (JSONL)
Each line contains a JSON object with:
id: Original image ID from CivitAI
improved_prompt: Flux-enhanced version of the prompt
category: Automatically determined theme category
All original CivitAI metadata including:
Original prompt and negative prompt
Model information
Image URL and… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/trending-text-to-image.diagram_image_to_text
Dataset Card for "diagram_image_to_text"
More Information needed
dior_text_to_imagetext-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.Chemistry_text_to_image
Dataset Card for "Chemistry_text_to_image"
More Information needed
TREC-2023-Text-to-Image
Dataset Card for "TREC-2023-Text-to-Image"
More Information needed
llama-3.2-random-images-to-textThis dataset contains 39020 unique image anotated pairs.
The images in the dataset are copywrited. and should only be used in acordance with the copywrite law.
For example I belive training LLM falls under fair use policies. Check with your lawyers.
Each image is anotated by the Meta llama 3.2 vison model. This first upload is part of the larger 5,000,000 image dataset.
The parquet files have the actual image inside them savd in raw bytes. they have the responce from the llm and also a unique… See the full description on the dataset page: https://huggingface.co/datasets/mylesgoose/llama-3.2-random-images-to-text.aid-text-to-imageChemistry_text_to_image_BASE64visdrone-text-to-imageimage-to-text-ocr-serp-snapshot
Image-to-Text OCR SERP Snapshot
An English-language, US Google organic-results snapshot for image to text converter, collected on 2026-09-22.
Files
image-to-text-ocr-serp-snapshot.csv is the machine-readable row-by-row dataset.
comparison.md is the same comparison as a readable Markdown table with methodology and limitations.
How to read the dataset
Each row is one result. The classification column distinguishes the Android app-store listing from… See the full description on the dataset page: https://huggingface.co/datasets/phoenix11000/image-to-text-ocr-serp-snapshot.winogroud_text_to_image
Dataset Card for "winogroud_text_to_image"
More Information needed
mm_diagram_image_to_text16xModdedMinecraft-TextToImage
Minecraft 16x Text-to-Image Dataset (Captioned)
Description
This dataset contains over 1 million Minecraft textures in 16x16 resolution. It has been specifically processed and captioned for training generative AI models (Text-to-Image).
Each image is paired with a descriptive natural language caption derived from the original file labels, enabling AI models to learn the relationship between Minecraft concepts (blocks, items, tools) and their pixel-art representation.… See the full description on the dataset page: https://huggingface.co/datasets/NathMen12/16xModdedMinecraft-TextToImage.image-to-textTREC-2023-Image-to-Text
Dataset Card for "TREC-2023-Image-to-Text"
More Information needed
speech_to_image_text_transformationcarla_image_to_text_datasetocr-image-to-textimage-description_text_to_image_BASE64diffusion.4.text_to_image.book
Dataset Card for "diffusion.4.text_to_image.book"
More Information needed
Text_to_ImageText to Image
