datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Codemixed_New
Codemixed ASR Dataset
Unified collection of code-mixed ASR datasets.
mucs-hindi-english-codemix-asr
MUCS 2021 Hindi-English Code-Mixed ASR
Hindi-English code-mixed speech recognition dataset from the
MUCS 2021 (Multilingual and
Code-Switching ASR Challenges) subtask 2, released as spoken-tutorial
recordings with Hindi-English code-mixed transcripts.
Long-form recordings were sliced into per-utterance clips using the
original Kaldi-style segments/text/utt2spk alignment.
Dataset structure
split
utterances
speakers
audio hours
avg clip len
train
52,825… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/mucs-hindi-english-codemix-asr.hindi_and_english_stt_tts_codemix_data
Hindi and English STT/TTS Codemix Data
Hinglish (Hindi-English code-mixed) speech dataset for automatic speech recognition (ASR) and text-to-speech (TTS) research.
Dataset Description
Each row is a timestamped speech segment clipped from conversational Hinglish audio recordings.
Column
Type
Description
text
string
Transcript of the speech segment (Hinglish)
audio
audio (16 kHz mono)
Corresponding audio clip
duration
float32
Clip duration in seconds… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hindi_and_english_stt_tts_codemix_data.
