CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agarwalayushi /hinglish Hinglish Concatenated Audio Dataset A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema. At a Glance Stat Value Total clips 815,171 Total Estimated Hours 2,264+ Unique speakers 6,304 Raw audio size ~243 GB Languages Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.audioautomatic-speech-recognition100K<n<1M7 likes1.9k downloads5mo agoHugging Face02suyash2739 /News_Hinglish_English News_Hinglish_English — An English ↔ Hinglish Parallel Corpus A curated parallel corpus of news-domain text in Hinglish (romanized Hindi-English code-mixed register) paired with corresponding standard English versions. Built to train and evaluate English → Hinglish translation models where existing resources (mostly conversational, e.g., CMU Hinglish DoG) don't cover the news register. DOI: 10.57967/hf/5120 · License: Apache 2.0 · Downloads: 2,500+ Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/suyash2739/News_Hinglish_English.texttranslation1K<n<10K2 likes1.1k downloads2mo agoHugging Face03dianavdavidson /MUCS-Hinglish MUCS Dataset Description This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset. This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2. As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here. In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.audioautomatic-speech-recognition10K<n<100K0 likes582 downloads7mo agoHugging Face04dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2audio100K<n<1M0 likes395 downloads2mo agoHugging Face05festvox /cmu_hinglish_dog Dataset Card for CMU Document Grounded Conversations Dataset Summary This is a collection of text conversations in Hinglish (code mixing between Hindi-English) and their corresponding English versions. Can be used for Translating between the two. The dataset has been provided by Prof. Alan Black's group from CMU. Supported Tasks and Leaderboards abstractive-mt Languages Dataset Structure Data Instances A typical data point… See the full description on the dataset page: https://huggingface.co/datasets/festvox/cmu_hinglish_dog.tabulartranslation1K<n<10K9 likes306 downloads3y agoHugging Face06dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.1audio100K<n<1M0 likes301 downloads3mo agoHugging Face07Abhishekcr448 /Hinglish-Everyday-Conversations-1M Dataset Card for Hinglish Everyday Conversations Dataset A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics. Use Model Access the model made using this dataset: Tiny-Hinglish-Chat-21M For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts. Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.texttext-generation1M<n<10M20 likes260 downloads2y agoHugging Face08dianavdavidson /indic-voices-hinglish-nospeakeroverlap-sponaudio100K<n<1M0 likes250 downloads5mo agoHugging Face09addyo07 /noisy-hinglish-asr Noisy Hinglish ASR Corpus A high-fidelity, mixed Hindi, English, Hinglish, and noise-robust ASR corpus designed for low-latency, localized voice assistant applications. Features conversational speech, heavy code-switching (English words embedded in Hindi structure), synthetically augmented desktop noise, and explicit non-speech negative frames for VAD optimization. Dataset Summary Total Samples: 28,681 audio recordings (main splits) Total Duration: 36.6 hours… See the full description on the dataset page: https://huggingface.co/datasets/addyo07/noisy-hinglish-asr.text10K<n<100K0 likes245 downloads4mo agoHugging Face10dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3audio100K<n<1M0 likes228 downloads3mo agoHugging Face11diwank /hinglish-dumpRaw merged dump of Hinglish (hi-EN) datasets.1 likes199 downloads5y agoHugging Face12dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3audio100K<n<1M0 likes179 downloads3mo agoHugging Face13ketav /hinglish-tts-data0 likes171 downloads6mo agoHugging Face14findnitai /english-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources. Sources: Hinglish TOP Dataset CMU English Dog HinGE PHINC source : 1 - Human Annotated , source : 0 - Synthetically Generated texttranslation100K<n<1M23 likes168 downloads3y agoHugging Face15dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.2audio100K<n<1M0 likes164 downloads3mo agoHugging Face16dianavdavidson /MUCS-Hinglish-traintestblindsplitaudio10K<n<100K0 likes163 downloads3mo agoHugging Face17open-llm-leaderboard-old /details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2 Dataset Card for Evaluation run of arshadshk/Mistral-Hinglish-7B-Instruct-v0.2 Dataset automatically created during the evaluation run of model arshadshk/Mistral-Hinglish-7B-Instruct-v0.2 on the Open LLM Leaderboard. The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2.0 likes152 downloads3y agoHugging Face18dianavdavidson /MUCS-Hinglish-twoaudio10K<n<100K0 likes142 downloads6mo agoHugging Face19Roshan32 /Hinglish_Dataset_instruction_and_rawtexttext-generation10K<n<100K1 likes133 downloads9mo agoHugging Face20ujs /hinglishA Hugginface version of the Hindi-English code-switched dataset from OpenSLR-104.audio10K<n<100K6 likes130 downloads3y agoHugging Face21rvv-karma /English-Hinglish-TOP English Hinglish (TOP Dataset) This dataset is generated from Hinglish-TOP Dataset. Data distribution: Train a. Human Generated - 6513 b. Synthetically generated - 170083 Validation a. Human Generated - 1390 b. Synthetically generated - 0 Test a. Human Generated - 6513 b. Synthetically generated - 0 texttranslation100K<n<1M1 likes129 downloads3y agoHugging Face22pankajbiswas6 /prism-hinglish-hate-speech PRISM - Code-Mixed Hinglish Hate-Speech Dataset Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text (RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle. Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track Summary Attribute Value Total samples (raw) 29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.texttext-classification10K<n<100K0 likes127 downloads3mo agoHugging Face23ankitdhiman /hinglish-conversations Hinglish Conversation Dataset Dataset Description This dataset contains 20403 conversations in Hinglish (Hindi-English code-mixed language). The conversations are casual dialogues that naturally mix Hindi and English, representing how many Indian users communicate in digital platforms. Dataset Structure The dataset is provided in multiple configurations to suit different use cases: Available Configurations default (turn_pairs): 203993 examples -… See the full description on the dataset page: https://huggingface.co/datasets/ankitdhiman/hinglish-conversations.text100K<n<1M1 likes122 downloads1y agoHugging Face24amulyabiradar23 /HinglishContractFlowtabular1M<n<10M0 likes117 downloads4mo agoHugging Face25HumynLabs /e-commerce-customersupport-hinglish-audio E-Commerce Customer Support Hinglish Audio Dataset Text spoken by all participants: "Mera order abhi tak nahi aaya, uska tracking kar sakte hain? Kal tak aana tha, mujhe lagta hai kahin kho gaya. Please update dein." The dataset supports training and evaluation of models in: Automatic Speech Recognition (ASR) Emotional tone classification Voice synthesis and generation Emotion-aware conversational agents Intended Uses ✅ Direct Use Training and… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/e-commerce-customersupport-hinglish-audio.audioaudio-classificationn<1K6 likes113 downloads1y agoHugging Face26tiny-aya-translate /hinglish-casual Hinglish Casual Speech 33,275 casual Hindi-English code-switched utterances (~31 GB) with audio, transcripts in both Devanagari and Latin script (utterance / utterance_latin), speaker ids, style metadata and durations. Full schema is in the YAML header above. Collected during the TinyAya programme to probe code-switched speech, which neither the FLORES-derived text nor the TTS corpora cover. It is not part of the v0.3 Stage-2 training set — that is tr-hi-mimi-encoded. from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.audioautomatic-speech-recognition10K<n<100K4 likes112 downloads2mo agoHugging Face27suyash2739 /Hinglish About Cleaned dataset from cmu_hinglish_dog [https://huggingface.co/datasets/cmu_hinglish_dog ] texttranslation1K<n<10K2 likes105 downloads2y agoHugging Face28Sujalvc /hinglish-instruct-dataset Akshar Hinglish Instruct Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns. 1. Dataset Overview Total Examples: 10,378 Base Set: 9,999 instruction-following pairs Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.texttext-generation10K<n<100K1 likes95 downloads4mo agoHugging Face29shreyansh12183 /shreyansh-hinglish-english-stem-500k 🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations. 📖 Overview In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.textquestion-answering100K<n<1M0 likes85 downloads1d agoHugging Face30theguywithblacktie /hinglish-conversations 🇮🇳 Hinglish Conversations & Instructions Dataset A high-purity, multi-subset conversational and instruction-following dataset curated for fine-tuning Large Language Models (such as LLaMA-3 / LLaMA-3.2, Mistral, and Qwen) to converse naturally, fluently, and authentically in Romanized Hinglish (code-mixed Hindi and English in Latin script). 📌 Key Highlights 100% Verified Hinglish: Every conversation turn and assistant output is strictly filtered to ensure… See the full description on the dataset page: https://huggingface.co/datasets/theguywithblacktie/hinglish-conversations.text100K<n<1M0 likes76 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.