CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01saeid1999 /fa-en-ar-handwritten-ocr-v1 Multi-script Synthetic Handwritten OCR — fa / ar / en A large, clean, augmentation-rich synthetic handwriting dataset for training and benchmarking OCR / HTR models on Persian (fa), Arabic (ar) and English (en). Every line image ships with an exact Unicode transcription plus rich provenance metadata (writer style, font, ink, script direction, digit system). Page-level PAGE-XML and COCO ground truth support layout-aware training and evaluation out of the box. 1,000 rendered… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1.imageimage-to-text10K<n<100K0 likes412 downloads9d agoHugging Face02AnanOmri /handwritten-dates-numbers-ocrimagen<1K0 likes346 downloads2mo agoHugging Face03sherif1313 /Historical-Arabic-Handwritten-OCRDescription A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image. No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.image1 likes196 downloads7mo agoHugging Face04tell2jyoti /odia-handwritten-ocr Odia Handwritten OCR Dataset Dataset Description This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution. Dataset Summary Language: Odia (ଓଡ଼ିଆ) Task: Optical Character Recognition (OCR) Total Images: 182,152 Character Classes: 47 Image Format: Grayscale JPG (32x32 pixels) Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.textimage-classification100K<n<1M0 likes96 downloads9mo agoHugging Face05V4ldeLund /Faroese-Handwritten-OCR Faroese Handwritten OCR Public draft, version 0.1: an alignment pilot with 16 line-image/text pairs from one historical Faroese manuscript page. No rows are verified benchmark ground truth. The page is image 3 of D IV – Ániasar táttur, held by Landsbókasavnið (National Library of the Faroe Islands) and digitized on HandRit. The manuscript is associated with the scribe Jóhan Hendrik Schrøter (1842–1911). Proposed reference text is aligned from Eivind Weyhe's scholarly edition of… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/Faroese-Handwritten-OCR.imageimage-to-textn<1K0 likes95 downloads7d agoHugging Face06tahirkiller /multi-lingual-handwritten-ocr-datasetimage100K<n<1M0 likes66 downloads2mo agoHugging Face07sarmisarmitha /sinhala-handwritten-ocr-623imagen<1K0 likes65 downloads23d agoHugging Face08toghrultahirov /handwritten_text_ocrimage10K<n<100K7 likes51 downloads2y agoHugging Face09deepcopy /toghrultahirov-handwritten_text_ocrimage10K<n<100K2 likes47 downloads1y agoHugging Face10openpecha /OCR-Handwritten_Tibetan_Cursive Configuration: default Split: train Total Rows: 70,528 image_url Type: categorical Data Type: object Unique Values: 2 Value Distribution: Value Count Percentage Portrait 39,017 55.32% Landscape 31,511 44.68% Original README dataset_info: features: - name: image_name dtype: string - name: transcript dtype: string - name: image_url dtype: string - name: orientation dtype:… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR-Handwritten_Tibetan_Cursive.image10K<n<100K0 likes37 downloads1y agoHugging Face11sedra-hugface /arabic-handwritten-ocr-eval Arabic OCR Evaluation Dataset This dataset contains Arabic paragraph images paired with ground truth text. Dataset Structure images/ train.csv validation.csv test.csv Task OCR evaluation for Arabic printed text. Metrics Models are evaluated using: CER (Character Error Rate) WER (Word Error Rate) Size ~600 images Use case Evaluation of Arabic OCR models. imagen<1K1 likes36 downloads6mo agoHugging Face12hieudt0803 /ocr-math-formula-handwrittenimage10K<n<100K1 likes20 downloads1y agoHugging Face13lokeshe09 /LaTeX_OCR_Handwrittenimagen<1K0 likes9 downloads8mo agoHugging Face14BounharAbdelaziz /OCR_Handwritten_Khattgatedimage10K<n<100K0 likes8 downloads1y agoHugging Face15Dakh /Historical-Arabic-Handwritten-OCRgatedDescription A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image. No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/Dakh/Historical-Arabic-Handwritten-OCR.imagen<1K0 likes7 downloads6mo agoHugging Face16ademax /ocr_handwritten_vigated Dataset Card for "ocr_handwritten_vi" More Information needed image1K<n<10K1 likes1 downloads3y agoHugging Face17AadityaJain /OCR_handwrittengatedimage10K<n<100K1 likes1 downloads2y agoHugging Face18marianeft /handwritten_name_ocr_dataset0 likes1 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.