CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01manavtabbly /hindi_audio_dataset_testaudion<1K0 likes4.4k downloads11mo agoHugging Face02KathirKs /fineweb-edu-hindi Fineweb-edu-hindi Fineweb-edu-hindi is a synthetic dataset generated by translating the Fineweb-edu to Hindi Language using IndicTrans2. The model variant used is IndicTrans2-en-indic-dist-200M. It contains about 300 Billion tokens in the Gemma-2-2b Tokenizer. Hardware Resources: The Google Cloud TPUs and the Google Cloud Platform was utilized for the dataset creation process. Code: Github: fineweb-translation Contact: If any queries or issues… See the full description on the dataset page: https://huggingface.co/datasets/KathirKs/fineweb-edu-hindi.text100M<n<1B8 likes4.3k downloads2y agoHugging Face03zicsx /mC4-Hindi-Cleaned-3.0 Dataset Card for "mC4-Hindi-Cleaned-3.0" More Information needed text1M<n<10M2 likes4.1k downloads3y agoHugging Face04hudsonburke /rat-hindlimb-mocap Rat Hindlimb Motion Capture Data Processed motion capture data from rat hindlimb gait analysis experiments. Dataset Structure processed/ ├── {subject_id}/ │ ├── markers.parquet # Marker positions (long format) │ ├── forceplates.parquet # Force plate data (long format) │ ├── events.parquet # Gait events (foot strike/off) │ └── sessions.parquet # Per-session anthropometrics └── ... Files markers.parquet: Time series of… See the full description on the dataset page: https://huggingface.co/datasets/hudsonburke/rat-hindlimb-mocap.tabular1B<n<10B1 likes3.9k downloads1mo agoHugging Face05agarwalayushi /hinglish Hinglish Concatenated Audio Dataset A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema. At a Glance Stat Value Total clips 815,171 Total Estimated Hours 2,264+ Unique speakers 6,304 Raw audio size ~243 GB Languages Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.audioautomatic-speech-recognition100K<n<1M7 likes2.1k downloads5mo agoHugging Face06AdaMLLab /HinMix HinMix (https://arxiv.org/abs/2512.18834) is a Hindi pretraining corpus containing 76 billion tokens across 60 million documents (in the minhash subset). Rather than scraping the web again, HinMix combines six publicly available Hindi datasets, applies Hindi-specific quality filtering, and performs cross-dataset deduplication. We train a 1.4B parameter language model through nanotron on 30 billion tokens to show that HinMix outperforms the previous state-of-the-art, CulturaX… See the full description on the dataset page: https://huggingface.co/datasets/AdaMLLab/HinMix.texttext-generation100M<n<1B1 likes1.7k downloads8mo agoHugging Face07cfilt /iitb-english-hindi IITB-English-Hindi Parallel Corpus About The IIT Bombay English-Hindi corpus contains parallel corpus for English-Hindi as well as monolingual Hindi corpus collected from a variety of existing sources and corpora developed at the Center for Indian Language Technology, IIT Bombay over the years. This page describes the corpus. This corpus has been used at the Workshop on Asian Language Translation Shared Task since 2016 the Hindi-to-English and English-to-Hindi… See the full description on the dataset page: https://huggingface.co/datasets/cfilt/iitb-english-hindi.text1M<n<10M71 likes1.6k downloads3y agoHugging Face08SPRINGLab /IndicTTS-Hindi Hindi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Hindi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.audiotext-to-speech10K<n<100K36 likes1.3k downloads2y agoHugging Face09zicsx /C4-Hindi-Cleaned Dataset Card for "C4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes1k downloads3y agoHugging Face10thoshith /hindi-english-raw-text-corpus-uncleanedtext100M<n<1B0 likes967 downloads2y agoHugging Face11suyash2739 /News_Hinglish_English News_Hinglish_English — An English ↔ Hinglish Parallel Corpus A curated parallel corpus of news-domain text in Hinglish (romanized Hindi-English code-mixed register) paired with corresponding standard English versions. Built to train and evaluate English → Hinglish translation models where existing resources (mostly conversational, e.g., CMU Hinglish DoG) don't cover the news register. DOI: 10.57967/hf/5120 · License: Apache 2.0 · Downloads: 2,500+ Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/suyash2739/News_Hinglish_English.texttranslation1K<n<10K2 likes728 downloads2mo agoHugging Face12zicsx /mC4-Hindi-Cleaned Dataset Card for "mC4-Hindi-Cleaned" More Information needed text1M<n<10M0 likes714 downloads3y agoHugging Face13pfin123 /hindi-aggregatedtext100K<n<1M2 likes649 downloads4y agoHugging Face14mkurman /hindawi-journals-2007-2023 Hindawi Academic Papers Dataset (CC BY 4.0 Compatible) Dataset Description This dataset contains 299,316 academic research papers from Hindawi Publishing Corporation, carefully filtered to include only papers with licenses compatible with CC BY 4.0. The dataset includes comprehensive metadata for each paper including titles, authors, journal information, publication years, DOIs, and full-text content. Dataset Summary Total Papers: 299,316 (filtered from 299… See the full description on the dataset page: https://huggingface.co/datasets/mkurman/hindawi-journals-2007-2023.texttext-generation100K<n<1M5 likes601 downloads1y agoHugging Face15Ritwika03 /syspin_hindi_mergedaudio10K<n<100K1 likes561 downloads1y agoHugging Face16SPRINGLab /Hindi-1482Hrsaudio100K<n<1M5 likes524 downloads2y agoHugging Face17SPRINGLab /IndicVoices-R_Hindiaudiotext-to-speech10K<n<100K11 likes486 downloads2y agoHugging Face18AsphyXIA /baarat-hindi-pretrain-datatext10M<n<100M3 likes478 downloads3y agoHugging Face19dianavdavidson /MUCS-Hinglish MUCS Dataset Description This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset. This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2. As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here. In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.audioautomatic-speech-recognition10K<n<100K0 likes474 downloads7mo agoHugging Face20kartikagg98 /HINMIX_hi-en Dataset Card for Hindi English Codemix Dataset - HINMIX HINMIX is a massive parallel codemixed dataset for Hindi-English code switching. See the 📚 paper on arxiv to dive deep into this synthetic codemix data generation pipeline. Dataset contains 4.2M fully parallel sentences in 6 Hindi-English forms. Further, we release gold standard codemix dev and test set manually translated by proficient bilingual annotators. Dev Set consists of 280 examples Test set consists of 2507 examples… See the full description on the dataset page: https://huggingface.co/datasets/kartikagg98/HINMIX_hi-en.texttranslation10M<n<100M6 likes468 downloads1y agoHugging Face21apjanco /hindi-ocr Dataset Card for Dataset Name Line-level text and images from https://heidata.uni-heidelberg.de/dataset.xhtml?persistentId=doi:10.11588/data/EGOKEI Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/apjanco/hindi-ocr.image1K<n<10K1 likes454 downloads2y agoHugging Face22dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2audio100K<n<1M0 likes425 downloads2mo agoHugging Face23Ritwika03 /hindi_karya_mergedaudio100K<n<1M0 likes398 downloads1y agoHugging Face24Paytmlabs /S2R_Shrutilipi_hindi Paytmlabs/S2R_Shrutilipi_hindi Hindi speech dataset prepared from ai4bharat/Shrutilipi for Ultravox training. Viewing samples on Hugging Face The hindi config stores audio inside Parquet. The website dataset viewer often cannot decode that and shows no rows. To inspect examples in the browser, open the Subset (config) drop-down and choose hindi_text_samples — text and continuation only (~2000 rows). Ultravox training should keep using subset hindi (full audio).… See the full description on the dataset page: https://huggingface.co/datasets/Paytmlabs/S2R_Shrutilipi_hindi.audioautomatic-speech-recognition100K<n<1M0 likes378 downloads6mo agoHugging Face25damerajee /long_context_hin_22ktext100K<n<1M0 likes375 downloads2y agoHugging Face26Abhishekcr448 /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M20 likes353 downloads2y agoHugging Face27collabora /hindi-asr-wdsaudio1M<n<10M0 likes347 downloads1y agoHugging Face28Asap7772 /aime-solution-hint-v6-deepscaler-respgentabular1K<n<10K0 likes345 downloads1y agoHugging Face29zicsx /mC4-hindi Dataset Card for "mC4-hindi" This dataset is a subset of the mC4 dataset, which is a multilingual colossal, cleaned version of Common Crawl's web crawl corpus. It contains natural text in 101 languages, including Hindi. This dataset is specifically focused on Hindi text, and contains a variety of different types of text, including news articles, blog posts, and social media posts. This dataset is intended to be used for training and evaluating natural language processing models for… See the full description on the dataset page: https://huggingface.co/datasets/zicsx/mC4-hindi.texttext-generation10M<n<100M0 likes337 downloads3y agoHugging Face30yiyic /Indo-Aryan-hin-urd-guj-pan_traintext1M<n<10M0 likes337 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.