datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chagatai-HTR-Snippets-v1.0
Chagatai Handwritten Text Recognition Snippets Dataset (v1.0)
This dataset contains paired line-image snippets and ground-truth text transcriptions derived from historical Chagatai handwritten manuscripts from the Jarring Collection. It is specifically curated for training and evaluating Handwritten Text Recognition (HTR) and Optical Character Recognition (OCR) models on low-resource historical languages using Arabic-based scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/chagatai-htr/Chagatai-HTR-Snippets-v1.0.Luoyan-Chat-Snippets-CN
