datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hinglish
Hinglish Concatenated Audio Dataset
A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema.
At a Glance
Stat
Value
Total clips
815,171
Total Estimated Hours
2,264+
Unique speakers
6,304
Raw audio size
~243 GB
Languages
Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.News_Hinglish_English
News_Hinglish_English — An English ↔ Hinglish Parallel Corpus
A curated parallel corpus of news-domain text in Hinglish (romanized Hindi-English code-mixed register) paired with corresponding standard English versions. Built to train and evaluate English → Hinglish translation models where existing resources (mostly conversational, e.g., CMU Hinglish DoG) don't cover the news register.
DOI: 10.57967/hf/5120 · License: Apache 2.0 · Downloads: 2,500+
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/suyash2739/News_Hinglish_English.MUCS-Hinglish
MUCS
Dataset Description
This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset.
This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2.
As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here.
In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2cmu_hinglish_dog
Dataset Card for CMU Document Grounded Conversations
Dataset Summary
This is a collection of text conversations in Hinglish (code mixing between Hindi-English) and their corresponding English versions. Can be used for Translating between the two. The dataset has been provided by Prof. Alan Black's group from CMU.
Supported Tasks and Leaderboards
abstractive-mt
Languages
Dataset Structure
Data Instances
A typical data point… See the full description on the dataset page: https://huggingface.co/datasets/festvox/cmu_hinglish_dog.indic-voices-hinglish-nospeakeroverlap-spon3.1Hinglish-Everyday-Conversations-1M
Dataset Card for Hinglish Everyday Conversations Dataset
A synthetically created Hinglish-based dataset of 2 columns where every row represents a unique conversation between 2 people in Hinglish about Everyday Life Topics.
Use Model
Access the model made using this dataset: Tiny-Hinglish-Chat-21M
For more information about this model, its training process, or related resources, you can check the GitHub repository Tiny-Hinglish-Chat-21M-Scripts.
Dataset Details… See the full description on the dataset page: https://huggingface.co/datasets/Abhishekcr448/Hinglish-Everyday-Conversations-1M.indic-voices-hinglish-nospeakeroverlap-sponnoisy-hinglish-asr
Noisy Hinglish ASR Corpus
A high-fidelity, mixed Hindi, English, Hinglish, and noise-robust ASR corpus designed for low-latency, localized voice assistant applications. Features conversational speech, heavy code-switching (English words embedded in Hindi structure), synthetically augmented desktop noise, and explicit non-speech negative frames for VAD optimization.
Dataset Summary
Total Samples: 28,681 audio recordings (main splits)
Total Duration: 36.6 hours… See the full description on the dataset page: https://huggingface.co/datasets/addyo07/noisy-hinglish-asr.indic-voices-hinglish-nospeakeroverlap-spon3hinglish-dumpRaw merged dump of Hinglish (hi-EN) datasets.indic-voices-hinglish-nospeakeroverlap-spon3.3hinglish-tts-dataenglish-to-hinglishEnglish to Hinglish Dataset aggregated from publicly available datasources.
Sources:
Hinglish TOP Dataset
CMU English Dog
HinGE
PHINC
source : 1 - Human Annotated ,
source : 0 - Synthetically Generated
indic-voices-hinglish-nospeakeroverlap-spon3.2MUCS-Hinglish-traintestblindsplitdetails_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2
Dataset Card for Evaluation run of arshadshk/Mistral-Hinglish-7B-Instruct-v0.2
Dataset automatically created during the evaluation run of model arshadshk/Mistral-Hinglish-7B-Instruct-v0.2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_arshadshk__Mistral-Hinglish-7B-Instruct-v0.2.MUCS-Hinglish-twoHinglish_Dataset_instruction_and_rawhinglishA Hugginface version of the Hindi-English code-switched dataset from OpenSLR-104.English-Hinglish-TOP
English Hinglish (TOP Dataset)
This dataset is generated from Hinglish-TOP Dataset.
Data distribution:
Train a. Human Generated - 6513 b. Synthetically generated - 170083
Validation a. Human Generated - 1390 b. Synthetically generated - 0
Test a. Human Generated - 6513 b. Synthetically generated - 0
prism-hinglish-hate-speech
PRISM - Code-Mixed Hinglish Hate-Speech Dataset
Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project
Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text
(RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle.
Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track
Summary
Attribute
Value
Total samples (raw)
29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.hinglish-conversations
Hinglish Conversation Dataset
Dataset Description
This dataset contains 20403 conversations in Hinglish (Hindi-English code-mixed language). The conversations are casual dialogues that naturally mix Hindi and English, representing how many Indian users communicate in digital platforms.
Dataset Structure
The dataset is provided in multiple configurations to suit different use cases:
Available Configurations
default (turn_pairs): 203993 examples -… See the full description on the dataset page: https://huggingface.co/datasets/ankitdhiman/hinglish-conversations.HinglishContractFlowe-commerce-customersupport-hinglish-audio
E-Commerce Customer Support Hinglish Audio Dataset
Text spoken by all participants:
"Mera order abhi tak nahi aaya, uska tracking kar sakte hain? Kal tak aana tha, mujhe lagta hai kahin kho gaya. Please update dein."
The dataset supports training and evaluation of models in:
Automatic Speech Recognition (ASR)
Emotional tone classification
Voice synthesis and generation
Emotion-aware conversational agents
Intended Uses
✅ Direct Use
Training and… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/e-commerce-customersupport-hinglish-audio.hinglish-casual
Hinglish Casual Speech
33,275 casual Hindi-English code-switched utterances (~31 GB) with audio,
transcripts in both Devanagari and Latin script (utterance /
utterance_latin), speaker ids, style metadata and durations. Full schema is in
the YAML header above.
Collected during the TinyAya programme to probe code-switched speech, which
neither the FLORES-derived text nor the TTS corpora cover. It is not part of
the v0.3 Stage-2 training set — that is
tr-hi-mimi-encoded.
from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.Hinglish
About
Cleaned dataset from cmu_hinglish_dog [https://huggingface.co/datasets/cmu_hinglish_dog ]
hinglish-instruct-dataset
Akshar Hinglish Instruct
Akshar Hinglish Instruct is a high-quality, code-mixed Romanized Hindi-English (Hinglish) instruction-tuning dataset containing 10,378 dialogue pairs. It is designed to train conversational language models to understand and generate natural, domain-diverse responses in Romanized South Asian speech patterns.
1. Dataset Overview
Total Examples: 10,378
Base Set: 9,999 instruction-following pairs
Domain Expansion Subset: 379 domain-specific… See the full description on the dataset page: https://huggingface.co/datasets/Sujalvc/hinglish-instruct-dataset.shreyansh-hinglish-english-stem-500k
🇮🇳 Vigyan Indic-STEM: 500k Bilingual Hinglish & English Reasoning Corpus
Vigyan Indic-STEM 500k is a specialized, large-scale bilingual dataset created to bridge the pedagogical divide in STEM education across India. It pairs rigorous English first-principles scientific derivations with natural, conversational Hinglish (Hindi written in Roman script) explanations.
📖 Overview
In Tier-2 and Tier-3 educational institutions across India, STEM concepts (Physics… See the full description on the dataset page: https://huggingface.co/datasets/shreyansh12183/shreyansh-hinglish-english-stem-500k.hinglish-conversations
🇮🇳 Hinglish Conversations & Instructions Dataset
A high-purity, multi-subset conversational and instruction-following dataset curated for fine-tuning Large Language Models (such as LLaMA-3 / LLaMA-3.2, Mistral, and Qwen) to converse naturally, fluently, and authentically in Romanized Hinglish (code-mixed Hindi and English in Latin script).
📌 Key Highlights
100% Verified Hinglish: Every conversation turn and assistant output is strictly filtered to ensure… See the full description on the dataset page: https://huggingface.co/datasets/theguywithblacktie/hinglish-conversations.
