datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
handwritten_cross-outs
HTR with Cross-out Words Dataset
This dataset is introduced in the paper:
"A Study of Handwritten Text Recognition with Cross-out Words"
DOI: https://doi.org/10.1007/s10032-026-00603-8
Overview
This dataset consists of handwritten word images produced by 12 different authors. It includes both clean (non-crossed-out) samples and crossed-out words, making it suitable for multiple handwriting-related research tasks.
The dataset introduces 7 distinct cross-out types… See the full description on the dataset page: https://huggingface.co/datasets/wahlinski/handwritten_cross-outs.zenodo-second-hand-fashion-v3
Second-Hand Fashion Dataset — wide (one row per garment)
Repack of Zenodo record 10.5281/zenodo.13788681 (Nauman et al., RISE + Wargön Innovation + Myrorna, CC-BY-4.0) into a one-row-per-garment wide layout so the HF dataset viewer shows every attribute — three images plus 25 metadata columns — on a single row.
Previous v3 releases stored one row per (garment, view) with satellite tables that had to be joined manually. That layout is preserved in git history if you need it; the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/zenodo-second-hand-fashion-v3.handwritten-digit-dataset
Handwritten Digit Dataset
This dataset contains a collection of handwritten digits (0-9) contributed by users through an interactive web-based drawing application. The dataset is continuously updated, reflecting real-world human handwriting variability.
Dataset Details
The images are pre-processed to match the standard machine learning format for digit recognition:
Dimensions: 28x28 pixels.
Format: Grayscale (single channel).
Processing: Each digit is cropped to… See the full description on the dataset page: https://huggingface.co/datasets/zentardev/handwritten-digit-dataset.handwriting-ocr
Handwriting OCR (Lance Format)
This Lance-formatted version of the Doctor's Handwritten Prescription BD dataset contains 4,680 cropped PNG images of handwritten medicine names from Bangladesh. Each row keeps the original image bytes with the medicine and generic-name labels, plus deterministic search metadata derived from those labels. The dataset contains three source-preserved splits: train, validation, and test.
[!NOTE]
Training note: The same samples appear repeatedly… See the full description on the dataset page: https://huggingface.co/datasets/lance-format/handwriting-ocr.d405-hand-surface-depth
D405 Hand-Surface Depth Dataset
A multi-user, multi-angle RGB-depth dataset for close-range hand-surface interaction research, captured with an Intel RealSense D405 stereo depth camera.
Paper: Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction (CVPR 2026)
Authors: Mukhiddin Toshpulatov, Wookey Lee, Suan Lee, Geehyuk Lee
Key Statistics
Metric
Value
Total RGB-depth pairs
53,300… See the full description on the dataset page: https://huggingface.co/datasets/muxiddin19/d405-hand-surface-depth.handwriting-signatures-datasetA dataset of handwritten signatures with text prompts, designed for LoRA fine-tuning of diffusion models to generate realistic personal signatures.
Handwritten Signatures Dataset (Processed)
This dataset contains preprocessed handwritten signatures designed for signature verification and LoRA fine-tuning of diffusion models (e.g., Stable Diffusion) for text-to-image tasks.
📌 Processing steps
Labeling – assigned using computer vision and AI models to group… See the full description on the dataset page: https://huggingface.co/datasets/LuisitoLuisito/handwriting-signatures-dataset.handwriting-signatures-datasetA dataset of handwritten signatures with text prompts, designed for LoRA fine-tuning of diffusion models to generate realistic personal signatures.
Handwritten Signatures Dataset (Processed)
This dataset contains preprocessed handwritten signatures designed for signature verification and LoRA fine-tuning of diffusion models (e.g., Stable Diffusion) for text-to-image tasks.
📌 Processing steps
Labeling – assigned using computer vision and AI models to group signatures… See the full description on the dataset page: https://huggingface.co/datasets/sl4shuur/handwriting-signatures-dataset.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.handwriting-signatures-datasetA dataset of handwritten signatures with text prompts, designed for LoRA fine-tuning of diffusion models to generate realistic personal signatures.
Handwritten Signatures Dataset (Processed)
This dataset contains preprocessed handwritten signatures designed for signature verification and LoRA fine-tuning of diffusion models (e.g., Stable Diffusion) for text-to-image tasks.
📌 Processing steps
Labeling – assigned using computer vision and AI models to group signatures… See the full description on the dataset page: https://huggingface.co/datasets/jade-kai/handwriting-signatures-dataset.handwriting-signatures-datasetA dataset of handwritten signatures with text prompts, designed for LoRA fine-tuning of diffusion models to generate realistic personal signatures.
Handwritten Signatures Dataset (Processed)
This dataset contains preprocessed handwritten signatures designed for signature verification and LoRA fine-tuning of diffusion models (e.g., Stable Diffusion) for text-to-image tasks.
📌 Processing steps
Labeling – assigned using computer vision and AI models to group signatures… See the full description on the dataset page: https://huggingface.co/datasets/ebixhaferaj/handwriting-signatures-dataset.Marathi_Handwritten
Dataset Card for Marathi Handwritten OCR Dataset
Dataset Summary
The Marathi Handwritten Text Dataset is a collection of handwritten text images in Marathi (देवनागरी लिपी),
aimed at supporting the development of Optical Character Recognition (OCR) systems, handwriting analysis tools,
and language research.The dataset was curated from native Marathi speakers to ensure a variety of handwriting styles and character variations.
The dataset contains 2520 images with two… See the full description on the dataset page: https://huggingface.co/datasets/Process-Venue/Marathi_Handwritten.pyu-handwritten-consonant-dataset
Myanmar’s Ancient Heritage: Pyu Handwritten Consonant Dataset
An open-access, systematically curated handwritten dataset of the 33 ancient Pyu consonants. This project serves as a foundational baseline benchmark to support digital humanities, paleographical preservation, and advanced computer vision tasks such as Optical Character Recognition (OCR). The dataset is modeled directly after canonical historical references documented by Thiripyanchi U Tha Myat.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/pyu-handwritten-consonant-dataset.amt-airframe-handbook-dataset
AMT Airframe Handbook Dataset
A comprehensive dataset extracted from the FAA Aviation Maintenance Technician (AMT) Airframe Handbook, containing text content and rendered page images suitable for training vision-language models.

Overview
This dataset was created using the doc-parser-engine - a production-grade document parsing engine with HuggingFace integration. The source document is the FAA… See the full description on the dataset page: https://huggingface.co/datasets/Remixonwin/amt-airframe-handbook-dataset.Marathi_Handwritten
Dataset Card for Marathi Handwritten OCR Dataset
Dataset Summary
The Marathi Handwritten Text Dataset is a collection of handwritten text images in Marathi (देवनागरी लिपी),
aimed at supporting the development of Optical Character Recognition (OCR) systems, handwriting analysis tools,
and language research.The dataset was curated from native Marathi speakers to ensure a variety of handwriting styles and character variations.
The dataset contains 2520 images with two… See the full description on the dataset page: https://huggingface.co/datasets/GodND/Marathi_Handwritten.
