CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.4k downloads11mo agoHugging Face02KathirKs /fineweb-edu-hindi Fineweb-edu-hindi Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2. The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer. Hardware Resources: The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process. Code: Github: fineweb-translation Contact: If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.text100M<n<1B8 likes4.3k downloads2y agoHugging Face03zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes4.1k downloads3y agoHugging Face04hudsonburke /rat-hindlimb-mocap Rat Hindlimb Motion Capture Data Processed motion capture data from rat hindlimb gait analysis experiments. Dataset Structure processed/ ├── {subject_id}/ │ ├── markers.parquet # Marker positions (long format) │ ├── forceplates.parquet # Force plate data (long format) │ ├── events.parquet # Gait events (foot strike/off) │ └── sessions.parquet # Per-session anthropometrics └── ... Files markers.parquet: Time series of… See the full description on the dataset page: https://huggingface.co/datasets/hudsonburke/rat-hindlimb-mocap.tabular1B<n<10B1 likes3.9k downloads1mo agoHugging Face05agarwalayushi /hinglish Hinglish Concatenated Audio Dataset A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema. At a Glance Stat Value Total clips 815,171 Total Estimated Hours 2,264+ Unique speakers 6,304 Raw audio size ~243 GB Languages Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.audioautomatic-speech-recognition100K<n<1M7 likes2.1k downloads5mo agoHugging Face06AdaMLLab /HinMix HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication. We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.texttext-generation100M<n<1B1 likes1.7k downloads8mo agoHugging Face07cfilt /iitb-english-hindi IITB-English-Hindi Parallel Corpus About The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/iitb-english-hindi.text1M<n<10M71 likes1.6k downloads3y agoHugging Face08SPRINGLab /IndicTTS-Hindi Hindi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Hindi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.audiotext-to-speech10K<n<100K36 likes1.3k downloads2y agoHugging Face09zicsx /C4-Hindi-Cleaned Dataset Card for "C4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes1k downloads3y agoHugging Face10thoshith /hindi-english-raw-text-corpus-uncleanedtext100M<n<1B0 likes967 downloads2y agoHugging Face11keysun89 /hindi_data_975_original2 likes914 downloads12d agoHugging Face12suyash2739 /News_Hinglish_English News_Hinglish_English — An English ↔ Hinglish Parallel Corpus A curated parallel corpus of news-domain text in Hinglish (romanized Hindi-English code-mixed register) paired with corresponding standard English versions. Built to train and evaluate English → Hinglish translation models where existing resources (mostly conversational, e.g., CMU Hinglish DoG) don't cover the news register. DOI: 10.57967/hf/5120 · License: Apache 2.0 · Downloads: 2,500+ Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/suyash2739/News_Hinglish_English.texttranslation1K<n<10K2 likes728 downloads2mo agoHugging Face13zicsx /mC4-Hindi-Cleaned Dataset Card for "mC4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes714 downloads3y agoHugging Face14sameerbanchhor /hindi_books_collectiondocument0 likes710 downloads11mo agoHugging Face15pfin123 /hindi-aggregatedtext100K<n<1M2 likes649 downloads4y agoHugging Face16mkurman /hindawi-journals-2007-2023 Hindawi Academic Papers Dataset (CC BY 4.0 Compatible) Dataset Description This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content. Dataset Summary Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.texttext-generation100K<n<1M5 likes601 downloads1y agoHugging Face17Ritwika03 /syspin_hindi_mergedaudio10K<n<100K1 likes561 downloads1y agoHugging Face18makaveli10 /whisper-hindi-preprocessed1K<n<10K0 likes539 downloads3y agoHugging Face19SPRINGLab /Hindi-1482Hrsaudio100K<n<1M5 likes524 downloads2y agoHugging Face20SPRINGLab /IndicVoices-R_Hindiaudiotext-to-speech10K<n<100K11 likes486 downloads2y agoHugging Face21AsphyXIA /baarat-hindi-pretrain-datatext10M<n<100M3 likes478 downloads3y agoHugging Face22dianavdavidson /MUCS-Hinglish MUCS Dataset Description This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset. This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2. As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here. In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.audioautomatic-speech-recognition10K<n<100K0 likes474 downloads7mo agoHugging Face23kartikagg98 /HINMIX_hi-en Dataset Card for Hindi English Codemix Dataset - HINMIX HINMIX is a massive parallel codemixed dataset for Hindi-English code switching. See the 📚 paper on arxiv to dive deep into this synthetic codemix data generation pipeline. Dataset contains 4.2M fully parallel sentences in 6 Hindi-English forms. Further, we release gold standard codemix dev and test set manually translated by proficient bilingual annotators. Dev Set consists of 280 examples Test set consists of 2507 examples… See the full description on the dataset page: https://huggingface.co/datasets/kartikagg98/HINMIX_hi-en.texttranslation10M<n<100M6 likes468 downloads1y agoHugging Face24apjanco /hindi-ocr Dataset Card for Dataset Name Line-level text and images from https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/EGOKEI Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/apjanco/hindi-ocr.image1K<n<10K1 likes454 downloads2y agoHugging Face25FreedomIntelligence /MMLU_HindiHindi version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT. 0 likes434 downloads3y agoHugging Face26dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2audio100K<n<1M0 likes425 downloads2mo agoHugging Face27neerajaabhyankar /hindustani-raag-small hindustani-raag-small 50 raags of Hindustani classical music, chunked from YouTube recordings, each chunk carrying the hand-annotated tonic (Sa) of the recording it came from. raags 50 train clips 1810 test clips 150 clip length 20-60 s source videos 412 Columns audio — an mp3 chunk, 20-60 s, cut from a full recording. label — ClassLabel over the 50 raag names. tonic_hz — the recording's Sa in Hz, annotated by ear and snapped to a peak… See the full description on the dataset page: https://huggingface.co/datasets/neerajaabhyankar/hindustani-raag-small.audioaudio-classification1K<n<10K3 likes417 downloads15d agoHugging Face28Ritwika03 /hindi_karya_mergedaudio100K<n<1M0 likes398 downloads1y agoHugging Face29NirantK /hda_nli_hindiThis dataset is a recasted version of the Hindi Discourse Analysis Dataset used to train models for Natural Language Inference Tasks in Low-Resource Languages like Hindi.text-classification10K<n<100K0 likes380 downloads3y agoHugging Face30Paytmlabs /S2R_Shrutilipi_hindi Paytmlabs/S2R_Shrutilipi_hindi Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training. Viewing samples on Hugging Face The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows. To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows). Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.audioautomatic-speech-recognition100K<n<1M0 likes378 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.