hinglish
Datasets
All datasets matching “hinglish”hinglish
Hinglish Concatenated Audio Dataset
A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema.
At a Glance
Stat
Value
Total clips
815,171
Total Estimated Hours
2,264+
Unique speakers
6,304
Raw audio size
~243 GB
Languages
Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.News_Hinglish_English
News_Hinglish_English — An English ↔ Hinglish Parallel Corpus
A curated parallel corpus of news-domain text in Hinglish (romanized Hindi-English code-mixed register) paired with corresponding standard English versions. Built to train and evaluate English → Hinglish translation models where existing resources (mostly conversational, e.g., CMU Hinglish DoG) don't cover the news register.
DOI: 10.57967/hf/5120 · License: Apache 2.0 · Downloads: 2,500+
Dataset summary… See the full description on the dataset page: https://huggingface.co/datasets/suyash2739/News_Hinglish_English.MUCS-Hinglish
MUCS
Dataset Description
This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset.
This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2.
As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here.
In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2cmu_hinglish_dog
Dataset Card for CMU Document Grounded Conversations
Dataset Summary
This is a collection of text conversations in Hinglish (code mixing between Hindi-English) and their corresponding English versions. Can be used for Translating between the two. The dataset has been provided by Prof. Alan Black's group from CMU.
Supported Tasks and Leaderboards
abstractive-mt
Languages
Dataset Structure
Data Instances
A typical data point… See the full description on the dataset page: https://huggingface.co/datasets/festvox/cmu_hinglish_dog.indic-voices-hinglish-nospeakeroverlap-spon3.1
