CoolFace
Datasetpublic

medyas/arabic-ocr-printed-500k

Arabic Printed OCR Lines — Synthetic, 500k A general-purpose printed Arabic text-line recognition corpus: 500,000 train + 2,000 val line images with labels, built to fine-tune line-recognition models (PaddleOCR PP-OCR rec CTC/MultiHead, TrOCR, etc.). Real line-crop printed-Arabic data does not exist at this scale on the Hub, so this corpus is rendered synthetically with diverse fonts + real Arabic text and a documented label/decoding contract. Why this exists… See the full description on the dataset page: https://huggingface.co/datasets/medyas/arabic-ocr-printed-500k.

sourceHugging Facecc-by-sa-4.0updated 3mo agoView on Hugging Face
0likes44downloads
4 commits on main
2671aa13mo ago

Upload arabic_ocr_printed_v1.tar with huggingface_hub

medyas
8afa1453mo ago

Upload manifest.json with huggingface_hub

medyas
72da6e73mo ago

Upload README.md with huggingface_hub

medyas
ee9c95d3mo ago

initial commit

medyas