CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.3k downloads11mo agoHugging Face02KathirKs /fineweb-edu-hindi Fineweb-edu-hindi Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2. The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer. Hardware Resources: The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process. Code: Github: fineweb-translation Contact: If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.text100M<n<1B8 likes4.2k downloads2y agoHugging Face03zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes3.4k downloads3y agoHugging Face04cfilt /iitb-english-hindi IITB-English-Hindi Parallel Corpus About The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/iitb-english-hindi.text1M<n<10M71 likes1.6k downloads3y agoHugging Face05SPRINGLab /IndicTTS-Hindi Hindi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Hindi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.audiotext-to-speech10K<n<100K36 likes1.3k downloads2y agoHugging Face06zicsx /C4-Hindi-Cleaned Dataset Card for "C4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes1k downloads3y agoHugging Face07thoshith /hindi-english-raw-text-corpus-uncleanedtext100M<n<1B0 likes835 downloads2y agoHugging Face08Ritwika03 /syspin_hindi_mergedaudio10K<n<100K1 likes748 downloads1y agoHugging Face09pfin123 /hindi-aggregatedtext100K<n<1M2 likes650 downloads4y agoHugging Face10zicsx /mC4-Hindi-Cleaned Dataset Card for "mC4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes592 downloads3y agoHugging Face11SPRINGLab /IndicVoices-R_Hindiaudiotext-to-speech10K<n<100K11 likes492 downloads2y agoHugging Face12SPRINGLab /Hindi-1482Hrsaudio100K<n<1M5 likes472 downloads2y agoHugging Face13apjanco /hindi-ocr Dataset Card for Dataset Name Line-level text and images from https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/EGOKEI Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/apjanco/hindi-ocr.image1K<n<10K1 likes457 downloads2y agoHugging Face14Ritwika03 /hindi_karya_mergedaudio100K<n<1M0 likes399 downloads1y agoHugging Face15AsphyXIA /baarat-hindi-pretrain-datatext10M<n<100M3 likes363 downloads3y agoHugging Face16Paytmlabs /S2R_Shrutilipi_hindi Paytmlabs/S2R_Shrutilipi_hindi Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training. Viewing samples on Hugging Face The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows. To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows). Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.audioautomatic-speech-recognition100K<n<1M0 likes351 downloads6mo agoHugging Face17collabora /hindi-asr-wdsaudio1M<n<10M0 likes349 downloads1y agoHugging Face18zicsx /mC4-hindi Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.texttext-generation10M<n<100M0 likes293 downloads3y agoHugging Face19jinaai /hindi-gov-vqa_beir This is a copy of https://huggingface.co/datasets/jinaai/hindi-gov-vqa reformatted into the BEIR format. For any further information like license, please refer to the original dataset. Disclaimer This dataset may contain publicly available images or text data. All data is provided for research and educational purposes only. If you are the rights holder of any content and have concerns regarding intellectual property or copyright, please contact us at "support-data (at) jina.ai"… See the full description on the dataset page: https://huggingface.co/datasets/jinaai/hindi-gov-vqa_beir.image1K<n<10K0 likes288 downloads1y agoHugging Face20Ritwika03 /hindi_indic_voice_raudio10K<n<100K0 likes287 downloads1y agoHugging Face21Sheeba2026 /bharatvani-hindi-speech-corpusgated BharatVani Hindi Speech Corpus (150-Hour Studio Dataset) Proprietary Speech Asset • TheCreatorOS • BharatVani AI 1. Overview The BharatVani Hindi Speech Corpus is an enterprise-grade, high-fidelity Indian speech dataset engineered specifically for training sovereign neural Text-to-Speech (TTS) models, voice cloning engines, and speech foundation models in Devanagari Hindi. Audio Clips: 103,784 Verified Studio Audio Clips (24,000 Hz, 16-bit Mono… See the full description on the dataset page: https://huggingface.co/datasets/Sheeba2026/bharatvani-hindi-speech-corpus.audiotext-to-speech100K<n<1M1 likes283 downloads4d agoHugging Face22MatrixSpeechAI /All_Hindi_ASR_v1.1audio10K<n<100K0 likes280 downloads2y agoHugging Face23c3rl /IIIT-INDIC-HW-WORDS-Hindi IIIT-INDIC-HW-WORDS-Hindi Dataset containing images of hand written words in Devanagari by various humans and the corresponding text of those images. Overview The dataset, originally developed by the Centre for Visual Information Technology (CVIT) at IIIT Hyderabad, has been transformed into Parquet format to facilitate its use in modern machine learning workflows. This dataset primarily targets recognition of handwritten Hindi words and aims to advance research and… See the full description on the dataset page: https://huggingface.co/datasets/c3rl/IIIT-INDIC-HW-WORDS-Hindi.imageimage-to-text10K<n<100K5 likes277 downloads2y agoHugging Face24anirudhlakhotia /baarat-batched-hindi-pre-trainingtext1M<n<10M0 likes265 downloads3y agoHugging Face25MatrixSpeechAI /All_Hindi_ASR_v1.2audio10K<n<100K0 likes259 downloads2y agoHugging Face26meharuhanzz /OCR-Bench1000-Hindi OCR-Bench1000-Hindi 1000 synthetic printed-text line images with ground-truth transcriptions, sampled from a larger locally-held Hindi OCR training corpus. This is a benchmark/sample release, not the full training set. Data fields Field Description file_name relative path to the image (images/...) text ground-truth transcription category hindi_only / english_only / mixed / numeric_and_symbols length_bucket short / medium / long, by character count… See the full description on the dataset page: https://huggingface.co/datasets/meharuhanzz/OCR-Bench1000-Hindi.imageimage-to-text1K<n<10K0 likes243 downloads9d agoHugging Face27damerajee /long_context_hindi Dataset This dataset was filtered from AI4BHarat dataset sangraha,which is the largest high-quality, cleaned Indic language pretraining data containing 251B tokens summed up over 22 languages, extracted from curated sources, existing multilingual corpora and large scale translations. This dataset contains only Hindi as of now Information First this dataset is mainly for long context training The minimum len is 6000 and maximum len is 3754718 Getting started… See the full description on the dataset page: https://huggingface.co/datasets/damerajee/long_context_hindi.texttext-generation100K<n<1M1 likes222 downloads2y agoHugging Face28fhai50032 /Hindi_corpus-4096-packed-qtk-1.37M fhai50032/Hindi_corpus-4096-packed-qtk-1.37M Packed pretraining corpus, 1,369,441 rows x 4096 tokens = 5.609B tokens, tokenized with fhai50032/QTK-81K. Format column type notes label list<int32> exactly 4096 tokens, no padding raw_label string label decoded back to text (redundant, for inspection) There is no attention_mask column: the corpus is packed, so every position is a real token and the mask would be all ones on every row. How… See the full description on the dataset page: https://huggingface.co/datasets/fhai50032/Hindi_corpus-4096-packed-qtk-1.37M.text1M<n<10M0 likes208 downloads2mo agoHugging Face29SayantanJoker /All_Hindi_ASR_v1.1audio10K<n<100K0 likes198 downloads1y agoHugging Face30AniBirage /HindiDownloadedData2audio10K<n<100K0 likes194 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.