CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ketanpatil03 /surya-ocr-500-image-to-textimagen<1K1 likes1k downloads2mo agoHugging Face02danielrosehill /Speech-To-Text-System-Prompts-2 Speech To Text System Prompt Library This repository provides a collection of system prompts designed to transform and refine text captured using speech-to-text technologies. By passing STT outputs through large language models with these specialized prompts, you can achieve cleaner, more structured, and purpose-specific text formats. 📋 The Idea Here is the basic implementation. I don't pretend that this is the stuff of high AI engineering. But it does create quite… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Speech-To-Text-System-Prompts-2.imagen<1K1 likes306 downloads1y agoHugging Face03ketanpatil03 /surya-ocr-1K-image-to-textimage1K<n<10K1 likes256 downloads2mo agoHugging Face04openpecha /OCR-Tibetan_line_to_text_benchmark Tibetan OCR-line-to-text Benchmark Dataset This repository hosts a line-to-text benchmark dataset to evaluate and compare Tibetan OCR models. The dataset includes diverse scripts, writing styles, and print methods, enabling comprehensive testing across multiple domains. 💽 Datasets Overview Features: filename: Name of the file. label: Ground truth text. image_url: URL of the image. BDRC_work_id: BDRC scan id for specific works. char_len: Character count of the text. script:… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR-Tibetan_line_to_text_benchmark.image100K<n<1M4 likes234 downloads11mo agoHugging Face05whosouravsharma /text-to-image-diffusiondb-2M DiffusionDB text-to-image subset A cleaned, safety-filtered image-prompt dataset for training a text-to-image model, built from DiffusionDB. Built on Hugging Face Jobs directly from poloclub/diffusiondb. It covers part_id 1-20 (20,000 source images) before filtering. The same content is also kept on the 20k-subset branch. Load it with: load_dataset("whosouravsharma/text-to-image-diffusiondb-2M") Note on the repo name: despite "2M" in the name, this is a small slice of… See the full description on the dataset page: https://huggingface.co/datasets/whosouravsharma/text-to-image-diffusiondb-2M.imagetext-to-image10K<n<100K0 likes202 downloads1mo agoHugging Face06anisiraj /code-image-to-text Code Snippet Image → Text A multimodal dataset for fine-tuning vision-language models (VLMs) on the task of transcribing an image of a code snippet back into its source text — syntax-aware OCR. Each example pairs a syntax-highlighted PNG of code with the exact code text that produced it. It spans 8 programming languages and deliberately mixes two capture types: block — a complete function / unit (6–45 lines). fragment — a contiguous partial view (3–14 lines) that may start or… See the full description on the dataset page: https://huggingface.co/datasets/anisiraj/code-image-to-text.imageimage-to-text10K<n<100K0 likes166 downloads3mo agoHugging Face07DetonateT2I /Detonate_Text_To_Imageimage10K<n<100K0 likes159 downloads2y agoHugging Face08Gufranzhcet21 /text_to_imageimage10K<n<100K1 likes92 downloads2y agoHugging Face09dathuynh1108 /captcha-to-text eKYB Captcha Labeled Dataset Generated at: 2026-05-10T17:09:52Z Summary Source metadata CSV: /Users/huynhthanhdat/Workspace/iNexus/eKYB/ocr-captcha-finetuned/datasets/hf_dataset/source/metadata.csv Rows total in source: 10954 Rows with label: 9000 Rows unlabeled: 1954 Rows skipped (missing image): 0 Exported labeled samples: 9000 Splits train: 8100 validation: 900 test: 0 Files HF imagefolder standard layout: train/*.png… See the full description on the dataset page: https://huggingface.co/datasets/dathuynh1108/captcha-to-text.imageimage-to-text1K<n<10K0 likes84 downloads5mo agoHugging Face10k-mktr /trending-text-to-image CivitAI Improved Prompts Dataset This dataset contains trending AI-generated images from CivitAI with Flux-improved prompts for better generation results. Dataset Format (JSONL) Each line contains a JSON object with: id: Original image ID from CivitAI improved_prompt: Flux-enhanced version of the prompt category: Automatically determined theme category All original CivitAI metadata including: Original prompt and negative prompt Model information Image URL and… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/trending-text-to-image.imagen<1K3 likes72 downloads10mo agoHugging Face11Kamizuru00 /diagram_image_to_text Dataset Card for "diagram_image_to_text" More Information needed imagen<1K15 likes68 downloads3y agoHugging Face12Westcott /dior_text_to_imageimage10K<n<100K0 likes47 downloads2y agoHugging Face13adhia /OCR-Tibetan_line_to_text_benchmark Tibetan OCR-line-to-text Benchmark Dataset This repository hosts a line-to-text benchmark dataset to evaluate and compare Tibetan OCR models. The dataset includes diverse scripts, writing styles, and print methods, enabling comprehensive testing across multiple domains. 💽 Datasets Overview Features: filename: Name of the file. label: Ground truth text. image_url: URL of the image. BDRC_work_id: BDRC scan id for specific works. char_len: Character count of the text. script:… See the full description on the dataset page: https://huggingface.co/datasets/adhia/OCR-Tibetan_line_to_text_benchmark.image100K<n<1M0 likes45 downloads3mo agoHugging Face14rivisia /text-to-image-2M text-to-image-2M: A High-Quality, Diverse Text-to-Image Training Dataset Overview text-to-image-2M is a curated text-image pair dataset designed for fine-tuning text-to-image models. The dataset consists of approximately 2 million samples, carefully selected and enhanced to meet the high demands of text-to-image model training. The motivation behind creating this dataset stems from the observation that datasets with over 1 million samples tend to produce better… See the full description on the dataset page: https://huggingface.co/datasets/rivisia/text-to-image-2M.imagetext-to-image100K<n<1M0 likes34 downloads9mo agoHugging Face15VuongQuoc /Chemistry_text_to_image Dataset Card for "Chemistry_text_to_image" More Information needed image100K<n<1M17 likes32 downloads3y agoHugging Face16Trkkk /text_to_img_street_sceneimagen<1K0 likes29 downloads2y agoHugging Face17Grandmasterhaile /Text-to-3D-Vehiclesimage1K<n<10K0 likes29 downloads1mo agoHugging Face18vyrohith /text-to-sketchimagen<1K0 likes28 downloads2y agoHugging Face19lansinuote /diffusion.4.text_to_image Dataset Card for "diffusion.4.text_to_image" More Information needed imagen<1K0 likes20 downloads3y agoHugging Face20Westcott /aid-text-to-imageimage10K<n<100K0 likes20 downloads2y agoHugging Face21vidavox /Text-to-Markdownimagen<1K1 likes20 downloads1y agoHugging Face22LeroyDyer /chart_text_to_Base64image1K<n<10K2 likes19 downloads2y agoHugging Face23Westcott /visdrone-text-to-imageimagetext-to-image1K<n<10K0 likes18 downloads2y agoHugging Face24LeroyDyer /Chemistry_text_to_image_BASE64image1K<n<10K2 likes16 downloads2y agoHugging Face25LeroyDyer /image-description_text_to_image_BASE64image1K<n<10K3 likes15 downloads2y agoHugging Face26vikas117 /maths_handwriting_to_textimage1 likes14 downloads2y agoHugging Face27FarhatMay /winogroud_text_to_image Dataset Card for "winogroud_text_to_image" More Information needed imagen<1K0 likes13 downloads3y agoHugging Face28TREC-AToMiC /TREC-2023-Image-to-Text Dataset Card for "TREC-2023-Image-to-Text" More Information needed imagen<1K0 likes12 downloads3y agoHugging Face29nimapourjafar /mm_diagram_image_to_textimagen<1K2 likes12 downloads2y agoHugging Face30Sumukhdev /carla_image_to_text_datasetimage1K<n<10K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.