CoolFace
8 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01agarwalayushi /hinglish Hinglish Concatenated Audio Dataset A large-scale, cleaned and annotated speech dataset covering Hindi, Hinglish (Hindi–English code-switching), and Indian English — compiled from 14 public corpora and original custom recordings, unified into a single Parquet dataset with consistent schema. At a Glance Stat Value Total clips 815,171 Total Estimated Hours 2,264+ Unique speakers 6,304 Raw audio size ~243 GB Languages Hindi (hi), Hinglish (hi-en), Indian… See the full description on the dataset page: https://huggingface.co/datasets/agarwalayushi/hinglish.audioautomatic-speech-recognition100K<n<1M7 likes1.9k downloads5mo agoHugging Face02dianavdavidson /MUCS-Hinglish MUCS Dataset Description This dataset is a HuggingFace/Transformers compatible version of the MUCS 2021 Hinglish dataset. This dataset is part of the MUltilingual and Code-Switching ASR Challenges for Low Resource Indian Languages challenge, subtask 2. As this dataset is in Hinglish, it contains codeswitching between Hindi and English. The original dataset was found here. In addition to making the dataset compatible for Transformers, preprocessing has been applied to… See the full description on the dataset page: https://huggingface.co/datasets/dianavdavidson/MUCS-Hinglish.audioautomatic-speech-recognition10K<n<100K0 likes565 downloads7mo agoHugging Face03tiny-aya-translate /hinglish-casual Hinglish Casual Speech 33,275 casual Hindi-English code-switched utterances (~31 GB) with audio, transcripts in both Devanagari and Latin script (utterance / utterance_latin), speaker ids, style metadata and durations. Full schema is in the YAML header above. Collected during the TinyAya programme to probe code-switched speech, which neither the FLORES-derived text nor the TTS corpora cover. It is not part of the v0.3 Stage-2 training set — that is tr-hi-mimi-encoded. from… See the full description on the dataset page: https://huggingface.co/datasets/tiny-aya-translate/hinglish-casual.audioautomatic-speech-recognition10K<n<100K4 likes118 downloads2mo agoHugging Face04dianavdavidson /mucs-hinglish-blindtestaudioautomatic-speech-recognition1K<n<10K0 likes41 downloads4mo agoHugging Face05bingbangboom /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/bingbangboom/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes20 downloads5mo agoHugging Face06SaiyanSai /cleaned-asr-transcripts-hinglish cleaned-asr-transcripts-hinglish bingbangboom/cleaned-asr-transcripts-hinglish is a parallel corpus containing 14k+ pairs of raw-synthetic Hindi ASR (Automatic Speech Recognition) transcripts mapped to their clean, properly punctuated, and transliterated "Hinglish" (Romanized Hindi) counterparts. This dataset is specifically designed for ASR post-processing, transliteration models, and fine-tuning Large Language Models (LLMs) to understand and generate high-quality, conversational… See the full description on the dataset page: https://huggingface.co/datasets/SaiyanSai/cleaned-asr-transcripts-hinglish.textautomatic-speech-recognition10K<n<100K0 likes16 downloads4mo agoHugging Face07sonexis-ai /hinglish-code-switched-conversations-v1gated Hinglish Code-Switched Conversational Dataset v1 Overview This dataset contains structured Hinglish conversational voice data built to reflect how people actually speak in real-world interactions. Most speech datasets are clean, scripted, or heavily processed. That works in controlled testing, but it breaks in production where speakers interrupt each other, switch languages, use regional accents, pause mid-thought, and shift context naturally. This sample release… See the full description on the dataset page: https://huggingface.co/datasets/sonexis-ai/hinglish-code-switched-conversations-v1.audioautomatic-speech-recognitionn<1K1 likes11 downloads4mo agoHugging Face08sajalmadan0909 /hinglish-stt-tts-deepgramgated Hinglish STT/TTS Speech with Deepgram Transcripts A Hindi-English code-mixed (Hinglish) speech dataset for automatic speech recognition (ASR) and text-to-speech (TTS) research. The dataset contains 23,543 timestamped speech segments from conversational recordings. Transcript replacement was performed using Deepgram where a non-empty result was available; otherwise, the original transcript was retained. Dataset structure Column Type Description text… See the full description on the dataset page: https://huggingface.co/datasets/sajalmadan0909/hinglish-stt-tts-deepgram.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.