datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fa-en-ar-handwritten-ocr-v1
Multi-script Synthetic Handwritten OCR — fa / ar / en
A large, clean, augmentation-rich synthetic handwriting dataset for
training and benchmarking OCR / HTR models on Persian (fa), Arabic
(ar) and English (en). Every line image ships with an exact Unicode
transcription plus rich provenance metadata (writer style, font, ink, script
direction, digit system). Page-level PAGE-XML and COCO ground truth support
layout-aware training and evaluation out of the box.
1,000 rendered… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1.handwritten-dates-numbers-ocrHistorical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.Faroese-Handwritten-OCR
Faroese Handwritten OCR
Public draft, version 0.1: an alignment pilot with 16 line-image/text pairs from one
historical Faroese manuscript page. No rows are verified benchmark ground truth.
The page is image 3 of D IV – Ániasar táttur, held by Landsbókasavnið
(National Library of the Faroe Islands) and digitized on HandRit. The manuscript
is associated with the scribe Jóhan Hendrik Schrøter (1842–1911). Proposed
reference text is aligned from Eivind Weyhe's scholarly edition of… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/Faroese-Handwritten-OCR.multi-lingual-handwritten-ocr-datasetsinhala-handwritten-ocr-623handwritten_text_ocrtoghrultahirov-handwritten_text_ocrOCR-Handwritten_Tibetan_Cursive
Configuration: default
Split: train
Total Rows: 70,528
image_url
Type: categorical
Data Type: object
Unique Values: 2
Value Distribution:
Value
Count
Percentage
Portrait
39,017
55.32%
Landscape
31,511
44.68%
Original README
dataset_info:
features:
- name: image_name
dtype: string
- name: transcript
dtype: string
- name: image_url
dtype: string
- name: orientation
dtype:… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR-Handwritten_Tibetan_Cursive.arabic-handwritten-ocr-eval
Arabic OCR Evaluation Dataset
This dataset contains Arabic paragraph images paired with ground truth text.
Dataset Structure
images/
train.csv
validation.csv
test.csv
Task
OCR evaluation for Arabic printed text.
Metrics
Models are evaluated using:
CER (Character Error Rate)
WER (Word Error Rate)
Size
~600 images
Use case
Evaluation of Arabic OCR models.
ocr-math-formula-handwrittenLaTeX_OCR_HandwrittenOCR_Handwritten_KhattHistorical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/Dakh/Historical-Arabic-Handwritten-OCR.ocr_handwritten_vi
Dataset Card for "ocr_handwritten_vi"
More Information needed
OCR_handwritten
