CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01skbose /indian-english-nptel-v0audio100K<n<1M3 likes607 downloads2y agoHugging Face02ishands /commonvoice-indian_accent CommonVoice Indian Accent Dataset Dataset Summary Total Audio Duration: 163.89 hours Number of Recordings: 110,088 Language: English (Indian Accent) Source: Mozilla CommonVoice Corpus v21 Licensing CC0 - Follows Mozilla CommonVoice dataset terms audio100K<n<1M0 likes361 downloads1y agoHugging Face03skbose /indian-english-nptel-testaudio100K<n<1M1 likes323 downloads2y agoHugging Face04deepdml /microsoft-speech-corpus-indian Microsoft Speech Corpus – Indian Languages Dataset Description This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript. Attribution required: "Data provided by Microsoft and SpeechOcean.com" ⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.audioautomatic-speech-recognition100K<n<1M3 likes306 downloads7mo agoHugging Face05WillHeld /india_accent_cvaudio10K<n<100K3 likes244 downloads3y agoHugging Face06WhissleAI /supreme-court-india-meta-speechaudio10K<n<100K0 likes234 downloads11mo agoHugging Face07En1gma02 /indian_accent_englishaudio1K<n<10K1 likes139 downloads2y agoHugging Face08Tachyeon /audio-fingerprint-indian-bench Audio Fingerprinting Benchmark on Indian Classical Music A reproducible, pre-registered benchmark of five audio-fingerprinting systems on the Saraga 1.5 corpus (Hindustani + Carnatic), plus a pre-registered training-recipe improvement to the NAFP baseline that achieves Bonferroni-significant gains on 1-second queries. v0.7  ·  closed-world retrieval  ·  5 systems  ·  6 528 evaluation cells  ·  pooled McNemar p = 3.18 × 10⁻⁶ TL;DR 5 systems benchmarked: Olaf, Dejavu… See the full description on the dataset page: https://huggingface.co/datasets/Tachyeon/audio-fingerprint-indian-bench.audioaudio-classification1K<n<10K0 likes129 downloads4mo agoHugging Face09aryanBanwala /audio-fingerprint-indian-bench Audio Fingerprinting Benchmark on Indian Classical Music A reproducible, pre-registered benchmark of five audio-fingerprinting systems on the Saraga 1.5 corpus (Hindustani + Carnatic), plus a pre-registered training-recipe improvement to the NAFP baseline that achieves Bonferroni-significant gains on 1-second queries. v0.7  ·  closed-world retrieval  ·  5 systems  ·  6 528 evaluation cells  ·  pooled McNemar p = 3.18 × 10⁻⁶ TL;DR 5 systems benchmarked: Olaf, Dejavu… See the full description on the dataset page: https://huggingface.co/datasets/aryanBanwala/audio-fingerprint-indian-bench.audioaudio-classification1K<n<10K0 likes118 downloads4mo agoHugging Face10VishalTheHuman /Common_Voice_Indian_Accentaudio10K<n<100K0 likes107 downloads2y agoHugging Face11grushaaaaa /tts-indian TTS Indian Languages Dataset Speech dataset for Text-to-Speech covering 6 Indian languages, collected and processed from YouTube. Languages & Speakers Speaker Language Gender monihara_bengali Bengali Male munir_kashmiri Kashmiri Male nandini_gujarati Gujarati Female sansri_kannada Kannada Female tamil_pokkisham Tamil Male teluguM Telugu Male Pipeline Audio was collected and processed through these stages: YouTube Download — yt-dlp… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/tts-indian.audio10K<n<100K0 likes101 downloads6mo agoHugging Face12Santhosh-kumar /Indian-Accent-Datasetaudion<1K0 likes91 downloads3y agoHugging Face13krishan23 /indian_englishaudio1K<n<10K4 likes90 downloads2y agoHugging Face14GhayasAhmed /indian_ASR_2 Dataset Card for "indian_ASR_2" More Information needed audio10K<n<100K1 likes73 downloads3y agoHugging Face15Praxel /codeswitch-pairs-lase-indian Codeswitch Pairs LASE — Indian-accent held-out corpus 1369 held-out cross-script utterance pairs from 8 ElevenLabs Indian-English Multilingual voices. Surfaces the accent-conditional finding: off-the-shelf encoders cluster Indian-accent voices closely regardless of script, while Western voices show large script-conditional gaps. Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script =… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase-indian.audioaudio-classification1K<n<10K2 likes71 downloads5mo agoHugging Face16humyn-labs /Indian-Emotional-Speech-Corpus Indian Emotional Speech Corpus Dataset Description This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry. Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points. Text spoken by all participants: (happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indian-Emotional-Speech-Corpus.audioaudio-classificationn<1K7 likes67 downloads7mo agoHugging Face17auraCodes /indian-english-hindi-tts-60min Indian English + Hindi TTS Dataset A small, heavily-curated Text-to-Speech corpus: 73.2 minutes (37.7 min Indian English, 35.5 min Hindi) of single-speaker, studio-grade clips. Every clip's audio was listened to and its transcript corrected against automated Sarvam ASR output; resulting WER against the corrected text is 0.05% (en-IN) and 0.0% (hi-IN), showing both very clean source audio and very accurate ASR. Built for the Sarvam AI ML & Speech Data Pipeline assignment using a… See the full description on the dataset page: https://huggingface.co/datasets/auraCodes/indian-english-hindi-tts-60min.audiotext-to-speechn<1K0 likes63 downloads3mo agoHugging Face18divi212 /india-supreme-court-audioaudio1K<n<10K0 likes52 downloads2y agoHugging Face19sudhanshu12 /KAI-indian-emotional-speech-corpus Indian Emotional Speech Corpus Dataset Description This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry. Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points. Text spoken by all participants: (happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/sudhanshu12/KAI-indian-emotional-speech-corpus.audioaudio-classificationn<1K0 likes50 downloads8mo agoHugging Face20MWirelabs /northeast-india-voicesgated Northeast India Voices A multilingual transcribed speech corpus covering five indigenous and regional languages of Northeast India: Khasi, Garo, Mizo, Nagamese, and Kokborok. Dataset Summary Language Family Utterances Khasi Austroasiatic 14,974 Nagamese Indo-Aryan (Creole) 7,386 Mizo Tibeto-Burman 7,250 Kokborok Tibeto-Burman 3,278 Garo Tibeto-Burman 1,612 Total 34,500 Data Collection Recorded by native speakers across… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-india-voices.audio10K<n<100K0 likes50 downloads3mo agoHugging Face21theothertom /indian_english_audio_2audion<1K0 likes49 downloads3y agoHugging Face22champTUSHARg007 /indian-tts-dataset Indian TTS Dataset A curated Text-to-Speech training dataset with high-quality audio clips, transcriptions, and emotion labels for Indian English (en-IN) and Hindi (hi-IN). Dataset Summary Metric Value Total clips 125 English (en-IN) 96 clips Hindi (hi-IN) 29 clips Total duration 28.8 minutes Sample rate 22050 Hz Format WAV (PCM 16-bit, mono) Emotion Distribution Emotion Count neutral 92 narrative 10 excited… See the full description on the dataset page: https://huggingface.co/datasets/champTUSHARg007/indian-tts-dataset.audiotext-to-speechn<1K0 likes49 downloads3mo agoHugging Face23Prasanna05 /indian-hindi-female-rawaudio1K<n<10K0 likes44 downloads10mo agoHugging Face24rojin254 /PP4-indian-music-csiaudio10K<n<100K0 likes35 downloads4mo agoHugging Face25edwixxx /indian_ted_talks_chunkedaudio1K<n<10K1 likes32 downloads2y agoHugging Face26Gbssreejith /indian-english-voiceaudio1K<n<10K1 likes32 downloads2y agoHugging Face27Sk1382 /Indian_Englsih_SSML_dataset_for_orpheus_fine_tuningaudion<1K0 likes32 downloads1y agoHugging Face28Abhi29112005 /sarvam-indian-eng-hin-tts Indian English + Hindi TTS Dataset (emotion-tagged) A curated, single-speaker-per-clip speech dataset for Text-to-Speech research, covering Indian English and Hindi. Every clip is sourced from YouTube, transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM. Total: 82 clips, 55.6 minutes Hindi: 28.8 min &nbsp;|&nbsp; Indian English: 26.8 min Audio: mono, 24 kHz, 16-bit WAV Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.audiotext-to-speechn<1K0 likes31 downloads3mo agoHugging Face29InfoBayAI /English_India_Call_Center_Audio_Dataset_Dual_ChannelgatedDataset Description: This dataset is a large-scale collection of 30,320 processed English (India) dual-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems. It consists of real-world customer and agent speech recordings collected from call center environments.… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_India_Call_Center_Audio_Dataset_Dual_Channel.audioautomatic-speech-recognitionn<1K0 likes30 downloads9d agoHugging Face30theothertom /indian_english_extendedaudio1K<n<10K1 likes26 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.