datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
handwritten_cross-outs
HTR with Cross-out Words Dataset
This dataset is introduced in the paper:
"A Study of Handwritten Text Recognition with Cross-out Words"
DOI: https://doi.org/10.1007/s10032-026-00603-8
Overview
This dataset consists of handwritten word images produced by 12 different authors. It includes both clean (non-crossed-out) samples and crossed-out words, making it suitable for multiple handwriting-related research tasks.
The dataset introduces 7 distinct cross-out types… See the full description on the dataset page: https://huggingface.co/datasets/wahlinski/handwritten_cross-outs.handwritten-digit-dataset
Handwritten Digit Dataset
This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability.
Dataset Details
The images are pre-processed to match the standard machine learning format for digit recognition:
Dimensions: 28x28 pixels.
Format: Grayscale (single channel).
Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.Marathi_Handwritten
Dataset Card for Marathi Handwritten OCR Dataset
Dataset Summary
The Marathi Handwritten Text Dataset is a collection of handwritten text images in Marathi (देवनागरी लिपी),
aimed at supporting the development of Optical Character Recognition (OCR) systems, handwriting analysis tools,
and language research.The dataset was curated from native Marathi speakers to ensure a variety of handwriting styles and character variations.
The dataset contains 2520 images with two… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Marathi_Handwritten.pyu-handwritten-consonant-dataset
Myanmar’s Ancient Heritage: Pyu Handwritten Consonant Dataset
An open-access, systematically curated handwritten dataset of the 33 ancient Pyu consonants. This project serves as a foundational baseline benchmark to support digital humanities, paleographical preservation, and advanced computer vision tasks such as Optical Character Recognition (OCR). The dataset is modeled directly after canonical historical references documented by Thiripyanchi U Tha Myat.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pyu-handwritten-consonant-dataset.Marathi_Handwritten
Dataset Card for Marathi Handwritten OCR Dataset
Dataset Summary
The Marathi Handwritten Text Dataset is a collection of handwritten text images in Marathi (देवनागरी लिपी),
aimed at supporting the development of Optical Character Recognition (OCR) systems, handwriting analysis tools,
and language research.The dataset was curated from native Marathi speakers to ensure a variety of handwriting styles and character variations.
The dataset contains 2520 images with two… See the full description on the dataset page: https://huggingface.co/datasets/GodND/Marathi_Handwritten.
