CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01freococo /ocr_arabic_books Arabic OCR Books Dataset (ocr_arabic_books) This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions. 📚 Master Book Inventory Total running pages in repository: 181,427 # Book Name (English) Book Name (Arabic) Subset / Config Name Page Count Image Index Range 1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.imageimage-to-text100K<n<1M4 likes2.2k downloads3mo agoHugging Face02freococo /synth_shamela_ocr_arabic_books Synthetic Arabic Books Dataset Structured book pages rendered dynamically with style, font, and degradation variations. imageimage-to-text1M<n<10M2 likes1.9k downloads2mo agoHugging Face03mohajesmaeili /Persian_Arabic_TextLine_Image_Ocr_Mediumimage100K<n<1M18 likes337 downloads1y agoHugging Face04loay /arabic-ocr-synthetic-scans-faker-300k Arabic OCR Synthetic Scans (Faker 300k) A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness. Dataset Summary Samples: ~300,000 synthetic Arabic document pages Image format: JPEG, ~800×1200 px (embedded in Parquet) Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.imageimage-to-text100K<n<1M7 likes290 downloads7mo agoHugging Face05sherif1313 /Historical-Arabic-Handwritten-OCRDescription A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image. No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.image1 likes223 downloads7mo agoHugging Face06JayanthMuthu /arabic-ocrimage10K<n<100K2 likes160 downloads2y agoHugging Face07ahmedheakl /arocrbench_arabicocrPlease see paper & code for more information: https://github.com/mbzuai-oryx/KITAB-Bench https://arxiv.org/abs/2502.14949 imagen<1K0 likes139 downloads2y agoHugging Face08aallail /arabic_ocr_synth_2image10K<n<100K0 likes128 downloads2y agoHugging Face09FatimahEmadEldin /Gutenberg-Arabic-OCR-HTML-Pages Gutenberg Arabic HTML-Page Dataset 📖 Dataset Description The Gutenberg Arabic HTML-Page Dataset is a large-scale, synthetically generated dataset designed for training and evaluating document understanding and Optical Character Recognition (OCR) models. The primary goal of this project is to provide a comprehensive resource of page images paired with their corresponding structured HTML ground truth, with a focus on the Arabic language. The dataset was created by… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Gutenberg-Arabic-OCR-HTML-Pages.image1K<n<10K3 likes123 downloads1y agoHugging Face10mohajesmaeili /Persian_Arabic_TextLine_Image_Ocr_Small Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional]… See the full description on the dataset page: https://huggingface.co/datasets/mohajesmaeili/Persian_Arabic_TextLine_Image_Ocr_Small.image100K<n<1M15 likes119 downloads1y agoHugging Face11saleh-c4 /arabic-ocr-imagesimagen<1K0 likes95 downloads1y agoHugging Face12TheRealOKAI /arabic-ocr-datasetv3image10K<n<100K5 likes80 downloads10mo agoHugging Face13craneset /arabic-ocr ocr-data Alshams, The largest Arabic OCR dataset at the word level. This dataset is specifically designed for fine-grained word-level OCR tasks, providing precise word-level bounding box annotations for each image. Each word is annotated with pixel-accurate localization, enabling tasks such as text detection, text recognition, and end-to-end OCR. This repository currently provides an Arabic OCR dataset. 📦 Available Datasets The technical specifications of each… See the full description on the dataset page: https://huggingface.co/datasets/craneset/arabic-ocr.imagen<1K1 likes62 downloads9mo agoHugging Face14TheRealOKAI /arabic_ocr_merged_datasetimagen<1K0 likes50 downloads10mo agoHugging Face15abzoo /arabic-names-synthetic-ocrimage1K<n<10K0 likes43 downloads23d agoHugging Face16sedra-hugface /arabic-handwritten-ocr-eval Arabic OCR Evaluation Dataset This dataset contains Arabic paragraph images paired with ground truth text. Dataset Structure images/ train.csv validation.csv test.csv Task OCR evaluation for Arabic printed text. Metrics Models are evaluated using: CER (Character Error Rate) WER (Word Error Rate) Size ~600 images Use case Evaluation of Arabic OCR models. imagen<1K1 likes38 downloads6mo agoHugging Face17melsiddieg /qari-arabic-ocr-10kimage10K<n<100K1 likes36 downloads10mo agoHugging Face18mustaphaelkady /arabic-ocr-books Scanned Book Page Images This dataset contains page images rendered from 9 PDF file(s). Dataset structure One image per PDF page. One subfolder per source PDF. data/metadata.jsonl contains technical provenance for each page. Image settings DPI: 300 Maximum width: 1600 Grayscale: True Contrast factor: 1.4 JPEG quality: 95 White-margin cropping: True Intended use OCR, document understanding, knowledge extraction, fine-tuning, and… See the full description on the dataset page: https://huggingface.co/datasets/mustaphaelkady/arabic-ocr-books.imageimage-to-text1K<n<10K0 likes31 downloads1mo agoHugging Face19Omar-youssef /arabic-ocr-markdown-dataset Arabic Document OCR Markdown Dataset Dataset Description This dataset contains 1,256 pairs of document images and their corresponding Markdown representations, specifically designed for Arabic document OCR tasks. The dataset is intended for training and evaluating models that convert document images into structured Markdown text (image-to-markdown OCR). Features The dataset consists of two main features: image: Document images in various formats markdown:… See the full description on the dataset page: https://huggingface.co/datasets/Omar-youssef/arabic-ocr-markdown-dataset.image1K<n<10K0 likes27 downloads9mo agoHugging Face20aallail /arabic_ocr_synth_3image1K<n<10K0 likes24 downloads2y agoHugging Face21Omar-youssef /arabic-ocr-datasetimagen<1K0 likes21 downloads2y agoHugging Face22Yousefmd /arabic_ocr_datasetimage10K<n<100K0 likes18 downloads11mo agoHugging Face23ahmedheakl /arabic_arabicocrimagen<1K0 likes17 downloads2y agoHugging Face24ahmedheakl /arabic_ocrisiimage1K<n<10K1 likes16 downloads2y agoHugging Face25Akshit77 /arabic-ocr-dataset-7imagen<1K2 likes15 downloads1y agoHugging Face26Akshit77 /arabic-ocr-dataset-2imagen<1K0 likes14 downloads1y agoHugging Face27oddadmix /Qari-OCR-0.2.2-Arabic-2B_Qari-0.2-eval-tashkilimagen<1K1 likes11 downloads5mo agoHugging Face28Akshit77 /arabic-ocr-dataset-5imagen<1K0 likes11 downloads1y agoHugging Face29hastyle /arabic-manuscript-ocrimage1K<n<10K0 likes10 downloads9mo agoHugging Face30TheRealOKAI /arabic-ocr-datasetv2imagen<1K0 likes8 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.