datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.odia-handwritten-ocr
Odia Handwritten OCR Dataset
Dataset Description
This dataset contains 182,152 handwritten Odia character images prepared for training OCR models. The dataset covers all 47 OHCS (Odia Handwritten Character Set) characters with balanced class distribution.
Dataset Summary
Language: Odia (ଓଡ଼ିଆ)
Task: Optical Character Recognition (OCR)
Total Images: 182,152
Character Classes: 47
Image Format: Grayscale JPG (32x32 pixels)
Splits: Train (145,717), Validation (18… See the full description on the dataset page: https://huggingface.co/datasets/tell2jyoti/odia-handwritten-ocr.clement-mk-handwritten
Clement — curated Macedonian conversations
700 instruction-following conversations in natively-styled Macedonian,
curated by a native speaker.
Provenance, stated plainly: the author wrote a seed set of conversations by
hand, derived exact prompt rules from them, and several frontier LLMs (Claude,
Gemini, DeepSeek) continued the generation under those rules. Every
conversation was then edited and reviewed by the author. Which rows are seeds
and which are generated was not… See the full description on the dataset page: https://huggingface.co/datasets/MartinV/clement-mk-handwritten.
