datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.Devanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/rhythmjain30/Devanagari-Characters-Image.devanagari_pretraindevanagari_pretraindevanagari-ocr-datasetDevanagari_PreTrainCorpusDevanagari-Characters-Image
Devanagari Characters Image Dataset
Dataset Summary
The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for:
Vowels (स्वर)
Consonants (व्यंजन)
Matra combinations (e.g., का, कि, की, कु)
Hindi numerals (०-९)
The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Yash141414/Devanagari-Characters-Image.devanagari_charater_handwrittenEmotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.devanagari_and_roman_digits
Dataset Card for Dataset Name
The OCR Digits Dataset consists of 20,000 high-quality images of digit combinations captured under various conditions. This dataset is designed to support research in optical character recognition, particularly for multi-digit recognition tasks.
Dataset Details
Citation
BibTeX:
@dataset{SumitYadav2025OCRDigits,
author = {[Sumit Yadav]},
title = {OCR Digits Dataset: A Collection of 20,000 Multi-Digit(Roman and… See the full description on the dataset page: https://huggingface.co/datasets/rockerritesh/devanagari_and_roman_digits.devanagari-glue-ocrdevanagari_ocr_pretrain
devanagari_ocr_pretrain
Incrementally compiled OCR dataset for Devanagari/Nepali adaptation.
Repo: himalaya-ai/devanagari_ocr_pretrain
Preset: devanagari_general_ocr
Raw rows use image and ocr columns plus source and language provenance.
Image paths are relative to the dataset root.
Generated by scripts/compile_ocr_datasets.py --upload-to-hf.
sanskrit-multitask-devanagariLID201_Devanagari_Script_Languages_Identificationdevanagari_ocr_dataset
Dataset Card: Devanagari Compiled Dataset (ShareGPT)
Dataset Description
This is a curated compilation of 7 public Devanagari OCR datasets, filtered for script purity and repackaged into the ShareGPT conversation format for fine-tuning vision-language models like GLM-OCR.
The dataset is distributed as 42 compressed image batches (to enable manageable downloads) alongside a single consolidated JSON annotation file — devanagari_ocr.json. No fixed… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/devanagari_ocr_dataset.devanagari-ocr-synthetic-80k
Devanagari OCR Synthetic 80K
80,000 synthetically rendered Devanagari (Hindi) text-line images paired with
their ground-truth transcription, intended for training / fine-tuning OCR and
text-recognition models (e.g. TrOCR, Donut, CRNN-CTC).
Each image is a single line of Hindi text rendered with a randomly chosen
font and font size.
Dataset Structure
column
type
description
image
Image
rendered text-line image (RGB PNG, 900x64 px)
text
string… See the full description on the dataset page: https://huggingface.co/datasets/iamkushagratomar/devanagari-ocr-synthetic-80k.pali-english-devanagari
Dataset Card for "pali-english-devanagari"
More Information needed
Inhouse_DevanagariIndicQA-devanagariDevanagari-OCR-ICL-Benchmark
Devanagari Post-OCR Correction Benchmark
A benchmark for evaluating post-OCR correction systems on Hindi and Marathi
text rendered in Devanagari script, accompanying the paper
"Evaluating In-Context Learning and Retrieval Strategies for Devanagari
Post-OCR Correction" (Bhandari and Harit, 2026).
Summary
20,000 evaluation pairs (10k Hindi, 10k Marathi) of (OCR-corrupted, ground-truth) sentences
~6,600 shot-bank pairs (~3.3k per language) for in-context example… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Devanagari-OCR-ICL-Benchmark.devanagari_ocr_pretrainnepaliflow-romanized-nepali-to-devanagari-dataset
NepaliFlow Romanized Nepali to Devanagari Dataset
This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script.
Task
The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output.
Columns
prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari
completion: expected Nepali Devanagari output
Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.sanskrit-multitask-devanagari1devanagari_ocr_graphemes
Devanagari OCR Grapheme Dataset
This repository hosts a grapheme‑level OCR dataset for the Devanagari script. Each data point consists of an image of a single grapheme and its corresponding Unicode text.
Dataset Structure
.
├── data/ # Directory containing all PNG images (e.g., 00000.png, 00001.png, ...)
├── devanagari_ocr_graphemes.json # ShareGPT‑formatted JSON file
└── README.md # This file
data/ – Images are stored in… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/devanagari_ocr_graphemes.Wiki2018_Devanagari_Script_Language_IdentificationHinglish-Everyday-Conversations-1M-DevanagariDevanagari-Ecommerce-fomatted-for-llama2-chat-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.Devanagari-Ecommerce-Dataset
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Prepared by:
Aakash… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-Dataset.devanagari-ocr-datasetcc100-nepali-strictly-cleaned-devanagari-only
CC-100 Nepali — Cleaned(Devanagari Only)
Pipeline
Unicode normalisation (NFC + ftfy)
Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate)
Language ID — fastText lid.176.bin, confidence ≥ 0.7
Exact deduplication (MD5)
Near-deduplication (char 13-gram bloom filter)
98/1/1 train/val/test split, seed 42
Usage
from datasets import load_dataset
ds = load_dataset("Basanta55/cc100-nepali-strictly-cleaned-devanagari-only")
