CoolFace
Datasetpublic

himalaya-ai/nepali-proofreader

Nepali OCR Proofreading Dataset (Devanagari) Dataset Summary A Nepali-only (Devanagari script) text-correction dataset built for fine-tuning a small language model (target: HimalayaGPT 0.5B) as an OCR proofreader. Each example is a (corrupted, clean) pair: corrupted is Nepali text with OCR/handwriting-style errors (character confusions, missing matras, merged/split words, transposed or dropped characters), and clean is the correct text it should map to. The set is… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/nepali-proofreader.

sourceHugging Facemitupdated 2mo agoView on Hugging Face
0likes36downloads

himalaya-ai/nepali-proofreader · main · files are served by the source, never re-hosted here