handwritten-ocr
Arabic-English-handwritten-OCR-v3Arabic-English-handwritten-OCR-v3-i1-GGUFArabic-English-handwritten-OCR-v3-GGUFArabic-English-handwritten-OCR-Qwen3-VL-4B-i1-GGUFlumynax-ocr-trocr-large-handwrittenArabic-English-handwritten-OCR-Qwen3-VL-4B-GGUFArabic-handwritten-OCR-4bit-Qwen2.5-VL-3B-v3Arabic-handwritten-OCR-4bit-Qwen2.5-VL-3B-v2
fa-en-ar-handwritten-ocr-v1
Multi-script Synthetic Handwritten OCR — fa / ar / en
A large, clean, augmentation-rich synthetic handwriting dataset for
training and benchmarking OCR / HTR models on Persian (fa), Arabic
(ar) and English (en). Every line image ships with an exact Unicode
transcription plus rich provenance metadata (writer style, font, ink, script
direction, digit system). Page-level PAGE-XML and COCO ground truth support
layout-aware training and evaluation out of the box.
1,000 rendered… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1.handwritten-dates-numbers-ocrHistorical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.Faroese-Handwritten-OCR
Faroese Handwritten OCR
Public draft, version 0.1: an alignment pilot with 16 line-image/text pairs from one
historical Faroese manuscript page. No rows are verified benchmark ground truth.
The page is image 3 of D IV – Ániasar táttur, held by Landsbókasavnið
(National Library of the Faroe Islands) and digitized on HandRit. The manuscript
is associated with the scribe Jóhan Hendrik Schrøter (1842–1911). Proposed
reference text is aligned from Eivind Weyhe's scholarly edition of… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/Faroese-Handwritten-OCR.multi-lingual-handwritten-ocr-dataset
