handwritten
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.persian-handwritten-digits
Persian Handwritten Digits (Farsi)
80,000 grayscale images of handwritten Persian (Farsi) digits — ۰۱۲۳۴۵۶۷۸۹ —
organized as an ImageFolder dataset with 10 classes (0–9), 8,000 images per class.
Each image is a 28×28 grayscale PNG of a single digit.
Classes
Class
Count
0 (۰)
8,000
1 (۱)
8,000
2 (۲)
8,000
3 (۳)
8,000
4 (۴)
8,000
5 (۵)
8,000
6 (۶)
8,000
7 (۷)
8,000
8 (۸)
8,000
9 (۹)
8,000
Usage
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/Mehdinmz/persian-handwritten-digits.Burmese-Handwritten-Sentence-Dataset
Burmese Handwritten Sentence Dataset (BHSD)
BHSD is a sentence-level Burmese handwriting dataset developed for optical character recognition (OCR), handwritten text recognition (HTR), error analysis, robustness testing, and research on low-resource scripts.
The dataset was created by Ah Maung Oo and DatarrX through the voluntary contributions of 54 handwriting writers.
This dataset would not have been possible without its volunteers. Every handwritten image in BSHD exists… See the full description on the dataset page: https://huggingface.co/datasets/DatarrX/Burmese-Handwritten-Sentence-Dataset.handwritten_cross-outs
HTR with Cross-out Words Dataset
This dataset is introduced in the paper:
"A Study of Handwritten Text Recognition with Cross-out Words"
DOI: https://doi.org/10.1007/s10032-026-00603-8
Overview
This dataset consists of handwritten word images produced by 12 different authors. It includes both clean (non-crossed-out) samples and crossed-out words, making it suitable for multiple handwriting-related research tasks.
The dataset introduces 7 distinct cross-out types… See the full description on the dataset page: https://huggingface.co/datasets/wahlinski/handwritten_cross-outs.English-Handwritten-Math-Notes-Dataset
English Handwritten Math Notes Dataset
This dataset contains high-resolution images of handwritten mathematical notes written in English. It includes problem statements, worked examples, formulas, and annotated derivations. The dataset supports AI research in handwriting recognition, mathematical OCR, and document understanding for STEM applications.
Contact
For queries or collaborations related to this dataset, contact:
anoushka@kgen.io
abhishek.vadapalli@kgen.io… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/English-Handwritten-Math-Notes-Dataset.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.
