datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hinglish
Hinglish Concatenated Audio Dataset
A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema.
At a Glance
Stat
Value
Total clips
815,171
Total Estimated Hours
2,264+
Unique speakers
6,304
Raw audio size
~243 GB
Languages
Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.MUCS-Hinglish
MUCS
Dataset Description
This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset.
This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2.
As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here.
In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.hinglish-casual
Hinglish Casual Speech
33,275 casual Hindi-English code-switched utterances (~31 GB) with audio,
transcripts in both Devanagari and Latin script (utterance /
utterance_latin), speaker ids, style metadata and durations. Full schema is in
the YAML header above.
Collected during the TinyAya programme to probe code-switched speech, which
neither the FLORES-derived text nor the TTS corpora cover. It is not part of
the v0.3 Stage-2 training set — that is
tr-hi-mimi-encoded.
from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.mucs-hinglish-blindtestcleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.cleaned-asr-transcripts-hinglish
cleaned-asr-transcripts-hinglish
bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts.
This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.hinglish-code-switched-conversations-v1
Hinglish Code-Switched Conversational Dataset v1
Overview
This dataset contains structured Hinglish conversational voice data built to reflect how people actually speak in real-world interactions.
Most speech datasets are clean, scripted, or heavily processed. That works in controlled testing, but it breaks in production where speakers interrupt each other, switch languages, use regional accents, pause mid-thought, and shift context naturally.
This sample release… See the full description on the dataset page: https://huggingface.co/datasets/sonexis-ai/hinglish-code-switched-conversations-v1.hinglish-stt-tts-deepgram
Hinglish STT/TTS Speech with Deepgram Transcripts
A Hindi-English code-mixed (Hinglish) speech dataset for automatic speech
recognition (ASR) and text-to-speech (TTS) research. The dataset contains
23,543 timestamped speech segments from conversational recordings. Transcript
replacement was performed using Deepgram where a non-empty result was
available; otherwise, the original transcript was retained.
Dataset structure
Column
Type
Description
text… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hinglish-stt-tts-deepgram.
