CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LLMDH /post-ocr2text100K<n<1M6 likes21k downloads1y agoHugging Face02echo840 /OCRBenchGithub|Paper OCRBench has been accepted by Science China Information Sciences. image1K<n<10K25 likes19k downloads2y agoHugging Face03ElectronicHug /short_video_ocr_dataset Short Video OCR / ASR Dataset An actively curated research dataset for building OCR, ASR, subtitle-alignment, and video-transcript pipelines for short social videos. It combines source videos and extracted frames with human review artifacts and model-generated text candidates. The primary languages are Ukrainian and Russian; English or mixed-language content may also occur. Status: work in progress. Model outputs and pseudo-label candidates are not ground truth. Only… See the full description on the dataset page: https://huggingface.co/datasets/ElectronicHug/short_video_ocr_dataset.imageimage-to-text1K<n<10K0 likes11k downloads2h agoHugging Face04Reza2kn /persian-ocr-community-dataset-argilla Persian OCR community dataset - Argilla view Lightweight two-column view for Argilla. image is an HF-hosted asset URL and label is JSON containing spatial OCR objects. image1K<n<10K0 likes4.4k downloads2mo agoHugging Face05howard-hou /OCR-VQA Dataset Card for "OCR-VQA" More Information needed image100K<n<1M60 likes4k downloads3y agoHugging Face06bevaya /pubmed-ocr PubMed-OCR: PMC Open Access OCR Annotations PubMed-OCR is an OCR-centric corpus of scientific articles derived from PubMed Central Open Access PDFs. Each page is rendered to an image and annotated with Google Cloud Vision OCR, released in a compact JSON schema with word-, line-, and paragraph-level bounding boxes. Scale (release): 209.5K articles ~1.5M pages ~1.3B words (OCR tokens) This dataset is intended to support layout-aware modeling, coordinate-grounded QA, and… See the full description on the dataset page: https://huggingface.co/datasets/bevaya/pubmed-ocr.textimage-to-text1M<n<10M72 likes3.2k downloads8mo agoHugging Face07unsloth /LaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR image10K<n<100K92 likes2.7k downloads2y agoHugging Face08moondream /ia_ocrContains pages from documents sourced from the Internet Archive, transcribed by Pixtral. Not super accurate, but useful during pretraining. @misc{moondream_ia_ocr, author = {Vikhyat Korrapati}, title = {IA OCR Dataset}, year = {2025}, url = {https://huggingface.co/datasets/moondream/ia_ocr}, note = {Accessed: 2025-03-07} } image100K<n<1M28 likes2.4k downloads1y agoHugging Face09Reza2kn /persian-ocr-community-datasetimage10K<n<100K1 likes2.4k downloads2mo agoHugging Face10JustANormalTinkerer /animetext-ocrimageimage-to-text1M<n<10M1 likes2.1k downloads15d agoHugging Face11ling99 /OCRBench_v2image10K<n<100K20 likes2.1k downloads2y agoHugging Face12Salesforce /blip3-ocr-200m BLIP3-OCR-200M Dataset Overview The BLIP3-OCR-200M dataset is designed to address the limitations of current Vision-Language Models (VLMs) in processing and interpreting text-rich images, such as documents and charts. Traditional image-text datasets often struggle to capture nuanced textual information, which is crucial for tasks requiring complex text comprehension and reasoning. Key Features OCR Integration: The dataset incorporates Optical Character… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-ocr-200m.image10M<n<100M45 likes2k downloads2y agoHugging Face13freococo /ocr_arabic_books Arabic OCR Books Dataset (ocr_arabic_books) This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions. 📚 Master Book Inventory Total running pages in repository: 181,427 # Book Name (English) Book Name (Arabic) Subset / Config Name Page Count Image Index Range 1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.imageimage-to-text100K<n<1M4 likes1.9k downloads3mo agoHugging Face14SBB /sbb-dc-ocr Dataset Card for Berlin State Library OCR data Dataset Summary The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945. At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages. For each page with OCR text, the language has been determined by langid (Lui/Baldwin 2012). Supported Tasks and Leaderboards language-modeling: this dataset has the potential to be used… See the full description on the dataset page: https://huggingface.co/datasets/SBB/sbb-dc-ocr.textfill-mask1M<n<10M8 likes1.7k downloads4y agoHugging Face15Yesianrohn /OCR-Data OCR Text Detection and Recognition Dataset Dataset Description A large-scale, multi-source OCR dataset aggregating 14 public benchmarks for text detection and recognition in both scene images and handwritten documents. Each image is paired with: Transcribed text for each text region Bounding boxes (axis-aligned rectangles) for each text region Polygon coordinates (precise boundary points) for each text region The dataset is stored in HuggingFace Parquet format with… See the full description on the dataset page: https://huggingface.co/datasets/Yesianrohn/OCR-Data.imageobject-detection100K<n<1M8 likes1.6k downloads6mo agoHugging Face16Reza2kn /persian-printed-ocr-3.5m Persian Printed OCR 3.5M A unified corpus of 3,517,974 Persian printed OCR image/text pairs, selected from five public datasets using GlotLID v3. Only the accept bucket is included; 232,317 ambiguous and 190,733 rejected rows are excluded. The viewer exposes exactly image and label. Sources AliShafiee2003/persian-ocr-garshasp-70c — pinned revision 36bfdcdeac20c02231f4ee08472f80db2fc467bb (CC-BY-4.0) hezarai/parsynth-ocr-200k — pinned revision… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-printed-ocr-3.5m.imageimage-to-text1M<n<10M2 likes1.6k downloads2mo agoHugging Face17lbourdois /OCR-liboaccn-OPUS-MIT-5M-clean Description This dataset is a processed version of liboaccn/OPUS-MIT-5M to make it easier to use, particularly for a visual question answering task where answer is an OCR transcription.Specifically, the original dataset has been processed to provide the image directly as a PIL rather than a path in an image column.We've also created a question column containing around 40 prompts based on via tutoiement, vouvoiement and imperative forms. Note that this dataset contains only the… See the full description on the dataset page: https://huggingface.co/datasets/lbourdois/OCR-liboaccn-OPUS-MIT-5M-clean.imagevisual-question-answering100K<n<1M0 likes1.4k downloads1y agoHugging Face18rifathridoy /bengali-ocr-synthetic Bengali OCR Synthetic Dataset A high-quality synthetic Bengali OCR dataset for fine-tuning vision-language models like DeepSeek-OCR 2. Generated using 100+ professional Bengali Unicode fonts and 13K+ unique Bengali words with advanced text rendering via FreeType and HarfBuzz. Dataset Overview Language: Bengali (বাংলা) Task: Optical Character Recognition (OCR) Format: Conversation-based (vision-language) Total Samples: 30,000 Train: 27,007 samples Validation: 2,993… See the full description on the dataset page: https://huggingface.co/datasets/rifathridoy/bengali-ocr-synthetic.imageimage-to-text10K<n<100K0 likes1.1k downloads7mo agoHugging Face19lingamvamshikrishnareddy /ramanv-document-ocr-2gatedtext100K<n<1M2 likes1k downloads19d agoHugging Face20linxy /LaTeX_OCR LaTeX OCR 的数据仓库 本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。 如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++ 后续追加新的数据也会放在这个仓库 ~~ 原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR. 数据集 本仓库有 5 个数据集 small 是小数据集,样本数 110 条,用于测试 full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。 synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.imageimage-to-text100K<n<1M181 likes948 downloads2y agoHugging Face21Gflorent /deepseek-ocr-artifacts-test-XXimage1K<n<10K0 likes924 downloads10mo agoHugging Face22lmms-lab /OCRBench-v2image10K<n<100K12 likes864 downloads2y agoHugging Face23hezarai /parsynth-ocr-200kParsynthOCR is a synthetic dataset for Persian OCR. This version is a preview of the original 4 million samples dataset (ParsynthOCR-4M). Usage 🤗 Datasets from datasets import load_dataset dataset = load_dataset("hezarai/parsynth-ocr-200k") Hezar pip install hezar from hezar.data import Dataset dataset = Dataset.load("hezarai/parsynth-ocr-200k", split="train") imageimage-to-image100K<n<1M24 likes802 downloads2y agoHugging Face24himalaya-ai /ocr-document-processing-eval ocr_document_processing_eval Document-processing OCR proxy for digitization, KYC-like numeric fields, and RAG extraction checks. Repo: himalaya-ai/ocr-document-processing-eval Task: document_processing_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr:… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/ocr-document-processing-eval.imageimage-to-textn<1K1 likes723 downloads3mo agoHugging Face25albertobarnabo /synthetic-receipts-ocr synthetic-receipts-ocr 32,000 synthetic thermal receipts across 5 locales (US/UK/DE/IT/FR) — each a clean render plus a photo-degraded twin, with pixel-exact word boxes, full transcription, and structured KIE fields. Samples train-000357 (US), train-000073 (UK), eval-001179 (DE), train-000222 (IT), train-000711 (FR) — real dataset rows, not mockups. Each receipt is its sample's image_photo, cut out along its own homography quad; no retouching beyond composition. Built for… See the full description on the dataset page: https://huggingface.co/datasets/albertobarnabo/synthetic-receipts-ocr.imageimage-to-text10K<n<100K0 likes701 downloads2mo agoHugging Face26chuuhtetnaing /myanmar-ocr-dataset-for-vlm Myanmar OCR Dataset A synthetic OCR dataset for fine-tuning Vision Language Models (VLMs) on Myanmar (Burmese) text recognition. It contains page images paired with their ground-truth text, sourced from chuuhtetnaing/mm-lib-book-dataset and rendered into page images using various Myanmar fonts. Subsets Subset Description Details single_font Rendered with Pyidaungsu font only 437 books multi_font Rendered with 76 Myanmar fonts 3 books (ပဋ္ဌာန်းမြတ်ဒေသနာ၊… See the full description on the dataset page: https://huggingface.co/datasets/chuuhtetnaing/myanmar-ocr-dataset-for-vlm.image100K<n<1M2 likes680 downloads5mo agoHugging Face27himalaya-ai /indic-deva-ocr-eval indic_deva_eval Broad Indic Devanagari OCR benchmark across printed pages, digits, word crops, and handwriting. Repo: himalaya-ai/indic-deva-ocr-eval Task: indic_devanagari_ocr Main raw file: *.ocr.jsonl with image, ocr, source_repo, and language/provenance columns. Optional fine-tuning/eval file: *.sharegpt.json with messages and images. Core Columns id: unique sample identifier image: relative path to the image file ocr: ground-truth text label… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/indic-deva-ocr-eval.imageimage-to-text1K<n<10K1 likes652 downloads4mo agoHugging Face28Jiwon-Kang /Llama-Nemotron-VLM-Dataset-v1-OCR4image100K<n<1M1 likes575 downloads8mo agoHugging Face29amrosama /al-kawakib-magazine-ocr Al-Kawakib Magazine OCR Pages This dataset contains rendered Arabic magazine page images paired with page-level text and line-level bounding boxes. It is intended for OCR, document understanding, and VLM fine-tuning experiments. Fine-Tuning Notebook A standalone Google Colab notebook for DeepSeek-OCR 3B + TRL SFT is available here: Open the fine-tuning notebook in Colab The notebook can train on this dataset alone or on all three Arabic magazine OCR… See the full description on the dataset page: https://huggingface.co/datasets/amrosama/al-kawakib-magazine-ocr.imageimage-to-text10K<n<100K3 likes564 downloads2mo agoHugging Face30alphabot2 /07-08-2026_OCR_BimanualThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "aibot2", "total_episodes": 37, "total_frames": 17770, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 10, "splits": { "train": "0:37" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/alphabot2/07-08-2026_OCR_Bimanual.imagerobotics10K<n<100K0 likes522 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.