CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Mayank022 /Devanagari-Characters-Image Devanagari Characters Image Dataset Dataset Summary The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for: Vowels (स्वर) Consonants (व्यंजन) Matra combinations (e.g., का, कि, की, कु) Hindi numerals (०-९) The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/Devanagari-Characters-Image.imageimage-classification10K<n<100K2 likes2.8k downloads1y agoHugging Face02rhythmjain30 /Devanagari-Characters-Image Devanagari Characters Image Dataset Dataset Summary The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for: Vowels (स्वर) Consonants (व्यंजन) Matra combinations (e.g., का, कि, की, कु) Hindi numerals (०-९) The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/rhythmjain30/Devanagari-Characters-Image.image-classification10K<n<100K0 likes923 downloads6mo agoHugging Face03cloudfrm-site /devanagari_pretraintext10M<n<100M0 likes211 downloads14d agoHugging Face04himalaya-ai /devanagari_pretraintext10M<n<100M0 likes185 downloads6mo agoHugging Face05vishwam-101 /devanagari-ocr-datasetimage1K<n<10K0 likes151 downloads1y agoHugging Face06Kiyo01 /Devanagari_PreTrainCorpustext10M<n<100M1 likes150 downloads2mo agoHugging Face07Yash141414 /Devanagari-Characters-Image Devanagari Characters Image Dataset Dataset Summary The Devanagari Characters Image Dataset is a high-resolution dataset designed to support research and experimentation in generative modeling, specifically for the Hindi script. It includes images for: Vowels (स्वर) Consonants (व्यंजन) Matra combinations (e.g., का, कि, की, कु) Hindi numerals (०-९) The dataset was created to address the limitations of existing Devanagari datasets, which often suffer from low resolution… See the full description on the dataset page: https://huggingface.co/datasets/Yash141414/Devanagari-Characters-Image.image-classification10K<n<100K0 likes121 downloads9mo agoHugging Face08AnjaliSarawgi /devanagari_charater_handwritten0 likes110 downloads1y agoHugging Face09kshitizgajurel /Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari Dataset Card for Dataset Name यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ। Dataset Prepared by: Manoj Kumar Baniya Aakash Kumar Thakur Manish Kathet Kshitiz Gajurel Dataset Details Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.text-generation10K<n<100K0 likes109 downloads2y agoHugging Face10rockerritesh /devanagari_and_roman_digits Dataset Card for Dataset Name The OCR Digits Dataset consists of 20,000 high-quality images of digit combinations captured under various conditions. This dataset is designed to support research in optical character recognition, particularly for multi-digit recognition tasks. Dataset Details Citation BibTeX: @dataset{SumitYadav2025OCRDigits, author = {[Sumit Yadav]}, title = {OCR Digits Dataset: A Collection of 20,000 Multi-Digit(Roman and… See the full description on the dataset page: https://huggingface.co/datasets/rockerritesh/devanagari_and_roman_digits.imageobject-detection10K<n<100K0 likes94 downloads1y agoHugging Face11himalaya-ai /devanagari-glue-ocrimage100K<n<1M0 likes84 downloads3mo agoHugging Face12himalaya-ai /devanagari_ocr_pretrain devanagari_ocr_pretrain Incrementally compiled OCR dataset for Devanagari/Nepali adaptation. Repo: himalaya-ai/devanagari_ocr_pretrain Preset: devanagari_general_ocr Raw rows use image and ocr columns plus source and language provenance. Image paths are relative to the dataset root. Generated by scripts/compile_ocr_datasets.py --upload-to-hf. imageimage-to-text10K<n<100K0 likes51 downloads4mo agoHugging Face13chronbmm /sanskrit-multitask-devanagaritext1M<n<10M0 likes49 downloads2y agoHugging Face141-800-SHARED-TASKS /LID201_Devanagari_Script_Languages_Identificationtext1M<n<10M0 likes43 downloads2y agoHugging Face15himalaya-ai /devanagari_ocr_dataset Dataset Card: Devanagari Compiled Dataset (ShareGPT) Dataset Description This is a curated compilation of 7 public Devanagari OCR datasets, filtered for script purity and repackaged into the ShareGPT conversation format for fine-tuning vision-language models like GLM-OCR. The dataset is distributed as 42 compressed image batches (to enable manageable downloads) alongside a single consolidated JSON annotation file — devanagari_ocr.json. No fixed… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/devanagari_ocr_dataset.image-to-text1M<n<10M0 likes41 downloads2mo agoHugging Face16iamkushagratomar /devanagari-ocr-synthetic-80k Devanagari OCR Synthetic 80K 80,000 synthetically rendered Devanagari (Hindi) text-line images paired with their ground-truth transcription, intended for training / fine-tuning OCR and text-recognition models (e.g. TrOCR, Donut, CRNN-CTC). Each image is a single line of Hindi text rendered with a randomly chosen font and font size. Dataset Structure column type description image Image rendered text-line image (RGB PNG, 900x64 px) text string… See the full description on the dataset page: https://huggingface.co/datasets/iamkushagratomar/devanagari-ocr-synthetic-80k.imageimage-to-text10K<n<100K0 likes41 downloads3d agoHugging Face17buddhist-nlp /pali-english-devanagari Dataset Card for "pali-english-devanagari" More Information needed text100K<n<1M0 likes40 downloads3y agoHugging Face18vrnP66 /Inhouse_Devanagaritext10K<n<100K0 likes32 downloads1y agoHugging Face19ai4bharat /IndicQA-devanagaritext1K<n<10K2 likes31 downloads2y agoHugging Face20AbhishekBhandari /Devanagari-OCR-ICL-Benchmark Devanagari Post-OCR Correction Benchmark A benchmark for evaluating post-OCR correction systems on Hindi and Marathi text rendered in Devanagari script, accompanying the paper "Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction" (Bhandari and Harit, 2026). Summary 20,000 evaluation pairs (10k Hindi, 10k Marathi) of (OCR-corrupted, ground-truth) sentences ~6,600 shot-bank pairs (~3.3k per language) for in-context example… See the full description on the dataset page: https://huggingface.co/datasets/AbhishekBhandari/Devanagari-OCR-ICL-Benchmark.texttext-generation10K<n<100K1 likes30 downloads3mo agoHugging Face21cloudfrm-site /devanagari_ocr_pretrain0 likes29 downloads14d agoHugging Face22dipeshch71 /nepaliflow-romanized-nepali-to-devanagari-dataset NepaliFlow Romanized Nepali to Devanagari Dataset This dataset contains instruction-style examples for converting Romanized Nepali words into Nepali Devanagari script. Task The task is to convert a Romanized Nepali word into its Devanagari form while returning only the Devanagari output. Columns prompt: instruction asking the model to convert a Romanized Nepali word into Devanagari completion: expected Nepali Devanagari output Size… See the full description on the dataset page: https://huggingface.co/datasets/dipeshch71/nepaliflow-romanized-nepali-to-devanagari-dataset.text10K<n<100K1 likes27 downloads3mo agoHugging Face23chronbmm /sanskrit-multitask-devanagari1text1M<n<10M0 likes26 downloads2y agoHugging Face24himalaya-ai /devanagari_ocr_graphemes Devanagari OCR Grapheme Dataset This repository hosts a grapheme‑level OCR dataset for the Devanagari script. Each data point consists of an image of a single grapheme and its corresponding Unicode text. Dataset Structure . ├── data/ # Directory containing all PNG images (e.g., 00000.png, 00001.png, ...) ├── devanagari_ocr_graphemes.json # ShareGPT‑formatted JSON file └── README.md # This file data/ – Images are stored in… See the full description on the dataset page: https://huggingface.co/datasets/himalaya-ai/devanagari_ocr_graphemes.imageimage-classification10K<n<100K1 likes26 downloads3mo agoHugging Face251-800-SHARED-TASKS /Wiki2018_Devanagari_Script_Language_Identificationtext1K<n<10K0 likes21 downloads2y agoHugging Face26AnudeepRao /Hinglish-Everyday-Conversations-1M-Devanagaritext1M<n<10M1 likes18 downloads5mo agoHugging Face27kshitizgajurel /Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset Dataset Card for Dataset Name यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ। This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-fomatted-for-llama2-chat-Dataset.texttext-generation1K<n<10K0 likes17 downloads2y agoHugging Face28kshitizgajurel /Devanagari-Ecommerce-Dataset Dataset Card for Dataset Name यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ। This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Prepared by: Aakash… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Devanagari-Ecommerce-Dataset.texttext-generation1K<n<10K1 likes15 downloads2y agoHugging Face29Malathip72 /devanagari-ocr-datasetimage1K<n<10K0 likes14 downloads5mo agoHugging Face30Basanta55 /cc100-nepali-strictly-cleaned-devanagari-only CC-100 Nepali — Cleaned(Devanagari Only) Pipeline Unicode normalisation (NFC + ftfy) Rule-based filters (length, Devanagari ratio ≥ 0.5, boilerplate) Language ID — fastText lid.176.bin, confidence ≥ 0.7 Exact deduplication (MD5) Near-deduplication (char 13-gram bloom filter) 98/1/1 train/val/test split, seed 42 Usage from datasets import load_dataset ds = load_dataset("Basanta55/cc100-nepali-strictly-cleaned-devanagari-only") tabular1M<n<10M0 likes11 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.