datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.persian-handwritten-digits
Persian Handwritten Digits (Farsi)
80,000 grayscale images of handwritten Persian (Farsi) digits — ۰۱۲۳۴۵۶۷۸۹ —
organized as an ImageFolder dataset with 10 classes (0–9), 8,000 images per class.
Each image is a 28×28 grayscale PNG of a single digit.
Classes
Class
Count
0 (۰)
8,000
1 (۱)
8,000
2 (۲)
8,000
3 (۳)
8,000
4 (۴)
8,000
5 (۵)
8,000
6 (۶)
8,000
7 (۷)
8,000
8 (۸)
8,000
9 (۹)
8,000
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Mehdinmz/persian-handwritten-digits.Burmese-Handwritten-Sentence-Dataset
Burmese Handwritten Sentence Dataset (BHSD)
BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR), error analysis, robustness testing, and research on low-resource scripts.
The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers.
This dataset would not have been possible without its volunteers. Every handwritten image in BSHD exists… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-Handwritten-Sentence-Dataset.handwritten_cross-outs
HTR with Cross-out Words Dataset
This dataset is introduced in the paper:
"A Study of Handwritten Text Recognition with Cross-out Words"
DOI: https://doi.org/10.1007/s10032-026-00603-8
Overview
This dataset consists of handwritten word images produced by 12 different authors. It includes both clean (non-crossed-out) samples and crossed-out words, making it suitable for multiple handwriting-related research tasks.
The dataset introduces 7 distinct cross-out types… See the full description on the dataset page: https://huggingface.co/datasets/wahlinski/handwritten_cross-outs.English-Handwritten-Math-Notes-Dataset
English Handwritten Math Notes Dataset
This dataset contains high-resolution images of handwritten mathematical notes written in English. It includes problem statements, worked examples, formulas, and annotated derivations. The dataset supports AI research in handwriting recognition, mathematical OCR, and document understanding for STEM applications.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/English-Handwritten-Math-Notes-Dataset.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.100-handwritten-medical-recordsseraiki-handwritten-numerals
📊 Seraiki Handwritten Numbers (1–99)
Dataset Summary
This dataset contains handwritten Seraiki numbers from 1 to 99, written in words using the Perso-Arabic Seraiki script.It was created to support Optical Character Recognition (OCR) research and to promote AI development for the underrepresented Seraiki language.
Unlike digit-only datasets (e.g., MNIST), this dataset includes full word forms of numbers, making it suitable for sequence-based recognition tasks (TrOCR… See the full description on the dataset page: https://huggingface.co/datasets/tahaListens/seraiki-handwritten-numerals.Handwritten-Physics-Notes-Dataset
English Handwritten Physics Notes Dataset
This dataset contains high-resolution images of handwritten physics notes written in English. The collection includes theoretical explanations, formulas, diagrams, derivations, and problem-solving steps. It is designed to support AI research in handwriting recognition, scientific OCR, and document understanding for physics and STEM education.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Physics-Notes-Dataset.handwritten_text_detection
Handwritten text detection dataset
Data domain
The blanks were provided by youth organization "Armenian Club" (telegram, instagram ), Russia Moscow.
The text on blanks was written during dictation "Teladrutyun" in 2018
The blanks were labeled by Amir and Renal during research project in HSE MIEM
Dataset info
Contains labeled dictations blanks in YOLO format
91 image in total, 73 (80%) for train and 18 (20%) for test
No image alignment or any preprocess… See the full description on the dataset page: https://huggingface.co/datasets/armvectores/handwritten_text_detection.fa-en-ar-handwritten-ocr-v1
Multi-script Synthetic Handwritten OCR — fa / ar / en
A large, clean, augmentation-rich synthetic handwriting dataset for
training and benchmarking OCR / HTR models on Persian (fa), Arabic
(ar) and English (en). Every line image ships with an exact Unicode
transcription plus rich provenance metadata (writer style, font, ink, script
direction, digit system). Page-level PAGE-XML and COCO ground truth support
layout-aware training and evaluation out of the box.
1,000 rendered… See the full description on the dataset page: https://huggingface.co/datasets/saeid1999/fa-en-ar-handwritten-ocr-v1.handwritten-dates-numbers-ocrGujarati-Handwritten-Dataset
Gujarati Handwritten Word Dataset
This dataset is a subset of the IIIT-INDIC-HW-WORDS collection, specifically focused on the Gujarati language. It is designed for training and evaluating Handwritten Text Recognition (HTR) models.
Dataset Summary
The original IIIT-INDIC-HW-WORDS is a large-scale benchmark for Indic scripts. This Gujarati subset contains word-level images manually written by multiple annotators to capture natural variations in handwriting styles.… See the full description on the dataset page: https://huggingface.co/datasets/prthm29/Gujarati-Handwritten-Dataset.Handwritten-Chemistry-Notes-Dataset
English Handwritten Chemistry Notes Dataset
This dataset contains high-resolution images of handwritten chemistry notes written in English. The collection includes equations, reaction mechanisms, periodic table references, structural diagrams, and descriptive explanations. It supports AI research in handwriting recognition, chemical structure understanding, and document analysis for STEM and educational applications.
Contact
For queries or collaborations related to this… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Chemistry-Notes-Dataset.image-text_handwritten-bundesratsprotokolle_xix-xx!!!This data set does not contain Ground Truth!!!
--- Data has been automatically created, using ATR models ---
Dataset Card for transkribus-exports-74823-raw-xml
This dataset was created using pagexml-hf converter from Transkribus PageXML data.
Dataset Summary
This dataset contains 148494 samples across 1 split(s).
These are ''automatically'' transcribed pages.
Images provided by the Federal Archives (Schweizerisches Bundesarchiv, BAR)
For a description of the project… See the full description on the dataset page: https://huggingface.co/datasets/dh-unibe/image-text_handwritten-bundesratsprotokolle_xix-xx.Handwritten-Biology-Notes-Dataset
English Handwritten Biology Notes Dataset
This dataset contains high-resolution images of handwritten biology notes written in English. The collection includes labeled diagrams, definitions, explanations of biological processes, and annotated sketches. It supports AI research in handwriting recognition, diagram understanding, and document interpretation within the field of life sciences.
Contact
For queries or collaborations related to this dataset, contact:… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Biology-Notes-Dataset.Historical-Arabic-Handwritten-OCRDescription
A collection of rich historical Arabic text, spanning different geographies across centuries, is present in this dataset. Experts have meticulously transcribed forty historical pages, each five from a distinct book, providing the textual ground truth for each image.
No data as such has been made available publicly previously, up to our knowledge. This intends to contribute to deep learning OCR modeling and testing by practitioners and researchers interested in Arabic OCR and… See the full description on the dataset page: https://huggingface.co/datasets/sherif1313/Historical-Arabic-Handwritten-OCR.khmer-handwritten-dataset-4.2k4.2k images khmer handwritten Dataset
This repository contains a dataset of khmer handwritten dataset
Installation
Prerequisites
Before cloning this repository, make sure you have Git LFS installed:
Install Git LFS
Linux/macOS:curl -s https://packagecloud.io/install/repositories/github/git-lfs/script.deb.sh | sudo bash
sudo apt install git-lfs
Windows:
Download and install Git LFS from git-lfs.github.com
Clone the Repository
git clone… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/khmer-handwritten-dataset-4.2k.Gujarati-Handwritten-Dataset
Gujarati Handwritten Word Dataset
This dataset is a subset of the IIIT-INDIC-HW-WORDS collection, specifically focused on the Gujarati language. It is designed for training and evaluating Handwritten Text Recognition (HTR) models.
Dataset Summary
The original IIIT-INDIC-HW-WORDS is a large-scale benchmark for Indic scripts. This Gujarati subset contains word-level images manually written by multiple annotators to capture natural variations in handwriting styles.… See the full description on the dataset page: https://huggingface.co/datasets/zarvin114/Gujarati-Handwritten-Dataset.LAM-handwrittenOriginal: https://www.kaggle.com/datasets/vpippi/lam-dataset
About Dataset
The Ludovico Antonio Muratori (LAM) dataset is the largest line-level HTR dataset to date and contains 25,823 lines from Italian ancient manuscripts edited by a single author over 60 years. The dataset comes in two configurations: a basic splitting and a date-based splitting which takes into account the age of the author. The first setting is intended to study HTR on ancient documents in Italian, while… See the full description on the dataset page: https://huggingface.co/datasets/ianua/LAM-handwritten.100-handwritten-medical-recordsFrenchCensus-handwritten-texts
Source
This repository contains 3 datasets created within the POPP project (Project for the Oceration of the Paris Population Census) for the task of handwriting text recognition. These datasets have been published in Recognition and information extraction in historical handwritten tables: toward understanding early 20th century Paris census at DAS 2022.
The 3 datasets are called “Generic dataset”, “Belleville”, and “Chaussée d’Antin” and contains lines made from the extracted rows… See the full description on the dataset page: https://huggingface.co/datasets/agomberto/FrenchCensus-handwritten-texts.Korean_Handwritten_Notes_Dataset
Korean Handwritten Notes Dataset
This dataset contains high-resolution images of Korean handwritten notes, including personal notes, class notes, and informal writings. The dataset has been anonymized and curated to support AI research in handwriting recognition, OCR, and document understanding.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io
Supported Tasks
Task Categories:
Image… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Korean_Handwritten_Notes_Dataset.Handwritten-Mathematical-Expression-Convert-LaTeXHandwritten_Historical_SpanishICDAR_2025_Handwritten_Notes_Understanding_ChallengeRapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts
Dataset Attribution
The original dataset is available on Kaggle.
This dataset has been curated solely for ease of use within the Hugging Face ecosystem, with no intention of plagiarizing or copying the original work.
Please cite the original authors if you use this dataset.
Citation
@INPROCEEDINGS{8978005,
author={Chamchong, Rapeeporn and Gao, Wei and McDonnell, Mark D.},
booktitle={2019 International Conference on Document Analysis and Recognition (ICDAR)}… See the full description on the dataset page: https://huggingface.co/datasets/fwgpiyawudk/RapeepornChamchong_Thai_Handwritten_Scripts_from_Ancient_Manuscripts.Faroese-Handwritten-OCR
Faroese Handwritten OCR
Public draft, version 0.1: an alignment pilot with 16 line-image/text pairs from one
historical Faroese manuscript page. No rows are verified benchmark ground truth.
The page is image 3 of D IV – Ániasar táttur, held by Landsbókasavnið
(National Library of the Faroe Islands) and digitized on HandRit. The manuscript
is associated with the scribe Jóhan Hendrik Schrøter (1842–1911). Proposed
reference text is aligned from Eivind Weyhe's scholarly edition of… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/Faroese-Handwritten-OCR.LAM-handwrittenOriginal: https://www.kaggle.com/datasets/vpippi/lam-dataset
About Dataset
The Ludovico Antonio Muratori (LAM) dataset is the largest line-level HTR dataset to date and contains 25,823 lines from Italian ancient manuscripts edited by a single author over 60 years. The dataset comes in two configurations: a basic splitting and a date-based splitting which takes into account the age of the author. The first setting is intended to study HTR on ancient documents in Italian, while the… See the full description on the dataset page: https://huggingface.co/datasets/deepcopy/LAM-handwritten.aida-handwritten
Handwritten OCR training data from AIDA-project
Dataset Summary
This dataset contains handwritten textline images and their transcriptions from the AIDA-project. It is a subset of the full AIDA dataset, containing only the best-quality handwritten annotations — lines where the annotator was confident about every character. The majority of lines are in Finnish, with some Swedish, English, French, and German.
Supported Tasks
The dataset was created for… See the full description on the dataset page: https://huggingface.co/datasets/caveman273/aida-handwritten.
