CoolFace
20 results

Arabic-OCR

freococo /ocr_arabic_books Arabic OCR Books Dataset (ocr_arabic_books) This repository is a structurally aligned Arabic Optical Character Recognition (OCR) dataset. It hosts a collection of classical Islamic books, organized by separate configurations (subsets) to support modular loading without filename collisions. 📚 Master Book Inventory Total running pages in repository: 181,427 # Book Name (English) Book Name (Arabic) Subset / Config Name Page Count Image Index Range 1… See the full description on the dataset page: https://huggingface.co/datasets/freococo/ocr_arabic_books.imageimage-to-text100K<n<1M4 likes1.9k downloads3mo agoHugging Facefreococo /synth_shamela_ocr_arabic_books Synthetic Arabic Books Dataset Structured book pages rendered dynamically with style, font, and degradation variations. imageimage-to-text1M<n<10M2 likes1.8k downloads2mo agoHugging FaceFatimahEmadEldin /deepseek-ocr-arabic-v11K<n<10K0 likes967 downloads8mo agoHugging Facesaleh-c4 /arabic-ocr-labelstextn<1K0 likes906 downloads2mo agoHugging Facemohajesmaeili /Persian_Arabic_TextLine_Image_Ocr_Mediumimage100K<n<1M18 likes328 downloads1y agoHugging Faceloay /arabic-ocr-synthetic-scans-faker-300k Arabic OCR Synthetic Scans (Faker 300k) A large-scale synthetic dataset of ~300,000 Arabic book pages generated to mimic real-world scanning imperfections. Designed for training Vision Language Models (VLMs) and OCR engines on structural layout analysis, font recognition, and document degradation robustness. Dataset Summary Samples: ~300,000 synthetic Arabic document pages Image format: JPEG, ~800×1200 px (embedded in Parquet) Text: Ground truth in UTF-8 with XML-style… See the full description on the dataset page: https://huggingface.co/datasets/loay/arabic-ocr-synthetic-scans-faker-300k.imageimage-to-text100K<n<1M7 likes307 downloads7mo agoHugging Face