datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
surya-ocr-500-image-to-textSpeech-To-Text-System-Prompts-2
Speech To Text System Prompt Library
This repository provides a collection of system prompts designed to transform and refine text captured using speech-to-text technologies.
By passing STT outputs through large language models with these specialized prompts, you can achieve cleaner, more structured, and purpose-specific text formats.
📋 The Idea
Here is the basic implementation. I don't pretend that this is the stuff of high AI engineering. But it does create quite… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Speech-To-Text-System-Prompts-2.surya-ocr-1K-image-to-textOCR-Tibetan_line_to_text_benchmark
Tibetan OCR-line-to-text Benchmark Dataset
This repository hosts a line-to-text benchmark dataset to evaluate and compare Tibetan OCR models. The dataset includes diverse scripts, writing styles, and print methods, enabling comprehensive testing across multiple domains.
💽 Datasets Overview
Features:
filename: Name of the file.
label: Ground truth text.
image_url: URL of the image.
BDRC_work_id: BDRC scan id for specific works.
char_len: Character count of the text.
script:… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR-Tibetan_line_to_text_benchmark.text-to-image-diffusiondb-2M
DiffusionDB text-to-image subset
A cleaned, safety-filtered image-prompt dataset for training a text-to-image
model, built from DiffusionDB.
Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers
part_id 1-20 (20,000 source images) before filtering. The same content is
also kept on the 20k-subset branch.
Load it with:
load_dataset("whosouravsharma/text-to-image-diffusiondb-2M")
Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.code-image-to-text
Code Snippet Image → Text
A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of
transcribing an image of a code snippet back into its source text — syntax-aware OCR.
Each example pairs a syntax-highlighted PNG of code with the exact code text that
produced it. It spans 8 programming languages and deliberately mixes two capture types:
block — a complete function / unit (6–45 lines).
fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.Detonate_Text_To_Imagetext_to_imagecaptcha-to-text
eKYB Captcha Labeled Dataset
Generated at: 2026-05-10T17:09:52Z
Summary
Source metadata CSV: /Users/huynhthanhdat/Workspace/iNexus/eKYB/ocr-captcha-finetuned/datasets/hf_dataset/source/metadata.csv
Rows total in source: 10954
Rows with label: 9000
Rows unlabeled: 1954
Rows skipped (missing image): 0
Exported labeled samples: 9000
Splits
train: 8100
validation: 900
test: 0
Files
HF imagefolder standard layout:
train/*.png… See the full description on the dataset page: https://huggingface.co/datasets/dathuynh1108/captcha-to-text.trending-text-to-image
CivitAI Improved Prompts Dataset
This dataset contains trending AI-generated images from CivitAI with Flux-improved prompts for better generation results.
Dataset Format (JSONL)
Each line contains a JSON object with:
id: Original image ID from CivitAI
improved_prompt: Flux-enhanced version of the prompt
category: Automatically determined theme category
All original CivitAI metadata including:
Original prompt and negative prompt
Model information
Image URL and… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/trending-text-to-image.diagram_image_to_text
Dataset Card for "diagram_image_to_text"
More Information needed
dior_text_to_imageOCR-Tibetan_line_to_text_benchmark
Tibetan OCR-line-to-text Benchmark Dataset
This repository hosts a line-to-text benchmark dataset to evaluate and compare Tibetan OCR models. The dataset includes diverse scripts, writing styles, and print methods, enabling comprehensive testing across multiple domains.
💽 Datasets Overview
Features:
filename: Name of the file.
label: Ground truth text.
image_url: URL of the image.
BDRC_work_id: BDRC scan id for specific works.
char_len: Character count of the text.
script:… See the full description on the dataset page: https://huggingface.co/datasets/adhia/OCR-Tibetan_line_to_text_benchmark.text-to-image-2M
text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset
Overview
text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.Chemistry_text_to_image
Dataset Card for "Chemistry_text_to_image"
More Information needed
text_to_img_street_sceneText-to-3D-Vehiclestext-to-sketchdiffusion.4.text_to_image
Dataset Card for "diffusion.4.text_to_image"
More Information needed
aid-text-to-imageText-to-Markdownchart_text_to_Base64visdrone-text-to-imageChemistry_text_to_image_BASE64image-description_text_to_image_BASE64maths_handwriting_to_textwinogroud_text_to_image
Dataset Card for "winogroud_text_to_image"
More Information needed
TREC-2023-Image-to-Text
Dataset Card for "TREC-2023-Image-to-Text"
More Information needed
mm_diagram_image_to_textcarla_image_to_text_dataset
