CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Reza2kn /persian-printed-ocr-3.5m Persian Printed OCR 3.5M A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733 rejected rows are excluded. The viewer exposes exactly image and label. Sources AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0) hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.imageimage-to-text1M<n<10M2 likes1.5k downloads2mo agoHugging Face02SoyVitou /62k-images-khmer-printed-dataset 62k Khmer-English Printed Dataset This repository contains a dataset of Khmer and English printed text images for training, validation, and testing. The dataset is stored in parquet format and managed using Git Large File Storage (LFS). Installation Prerequisites Before cloning this repository, make sure you have Git LFS installed: Install Git LFS Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/62k-images-khmer-printed-dataset.imagetext-generation10K<n<100K2 likes182 downloads2y agoHugging Face03medyas /arabic-ocr-printed-500k Arabic Printed OCR Lines — Synthetic, 500k A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and a documented label/decoding contract. Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.tabularimage-to-textn<1K0 likes40 downloads3mo agoHugging Face04arobin79 /bangla-ocr-validation_data_printed Bangla OCR Validation Dataset (Printed + Scanned) 📌 Description This dataset is a Bangla OCR validation dataset containing a mix of printed document images and their corresponding text annotations. It is designed to evaluate OCR and vision-language models on both clean digital text and scanned document images. 📊 Dataset Composition 1507 line-level images with text annotations 50 full-page document images with text Data includes: Printed/typed Bangla text… See the full description on the dataset page: https://huggingface.co/datasets/arobin79/bangla-ocr-validation_data_printed.imageimage-to-text1K<n<10K1 likes30 downloads5mo agoHugging Face05ud-synthetic /printed-usa-passports Introduction The Synthetic Printed USA Passports Dataset contains 9,600 AI-generated passport images designed for training OCR and computer vision models on identity documents. The dataset includes varied angles, lighting conditions, backgrounds, and distances, with structured metadata covering gender, age group, resolution, and more. All images are synthetically generated — no real personal data or biometric records are involved — making it a privacy-compliant solution for… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/printed-usa-passports.textimage-to-textn<1K1 likes23 downloads2mo agoHugging Face06thekamilya /kazakh-printed-dataset Kazakh Printed Dataset for OCR task Data Lineage This dataset was synthetically generated using issai/kazparc as the base. Since Kazakh OCR data is scarce, I developed a pipeline to transform digital Kazakh text into a printed-style dataset. Generation Process Source: Text samples were extracted from issai/kazparc. Augmentation & Stylization: Random Background Color: Simulates different lighting conditions by alternating between… See the full description on the dataset page: https://huggingface.co/datasets/thekamilya/kazakh-printed-dataset.image1K<n<10K0 likes23 downloads5mo agoHugging Face07ud-synthetic /printed-german-passports Introduction The Synthetic Printed German Passports Dataset contains 5,000 AI-generated passport images built for training OCR and computer vision models on printed identification documents. Each image is captured across 3 angles, 4 lighting conditions, 4 backgrounds, and 2 distances, with structured metadata covering passport ID, gender, age group, and more. Since all images are synthetically generated, the dataset contains no real personal data or biometric records — making it… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/printed-german-passports.textimage-to-textn<1K1 likes18 downloads2mo agoHugging Face08SoyVitou /khmer_font_printed_datasettext10K<n<100K0 likes13 downloads2y agoHugging Face09deepcopy /DonkeySmall-OCR-Cyrillic-Printed-8image1M<n<10M1 likes9 downloads1y agoHugging Face10AKKI-AFK /testset_printed_splitsimage10K<n<100K0 likes7 downloads1y agoHugging Face11electricsheepeurope /europe-owid-production-printed-books-half-century Production Printed Books Half Century | Europe (Our World in Data) 🇪🇺 77 observations · 11 Europe countries · 1475–1775 · Repackaged by Electric Sheep Europe TL;DR This dataset contains 77 observations of Production Printed Books Half Century data across 11 Europe countries, spanning 1475–1775. About the source Source: Our World in Data Publisher: Our World in Data License: cc-by-4.0 Topic: Production Printed Books Half Century… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-production-printed-books-half-century.tabulartabular-classificationn<1K0 likes7 downloads4mo agoHugging Face12mahesh006 /printed_text_ocrimage0 likes1 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.