datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indian-english-nptel-v0commonvoice-indian_accent
CommonVoice Indian Accent Dataset
Dataset Summary
Total Audio Duration: 163.89 hours
Number of Recordings: 110,088
Language: English (Indian Accent)
Source: Mozilla CommonVoice Corpus v21
Licensing
CC0 - Follows Mozilla CommonVoice dataset terms
indian-english-nptel-testmicrosoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.india_accent_cvsupreme-court-india-meta-speechindian_accent_englishaudio-fingerprint-indian-bench
Audio Fingerprinting Benchmark on Indian Classical Music
A reproducible, pre-registered benchmark of five audio-fingerprinting systems on the Saraga 1.5 corpus (Hindustani + Carnatic), plus a pre-registered training-recipe improvement to the NAFP baseline that achieves Bonferroni-significant gains on 1-second queries.
v0.7 · closed-world retrieval · 5 systems · 6 528 evaluation cells · pooled McNemar p = 3.18 × 10⁻⁶
TL;DR
5 systems benchmarked: Olaf, Dejavu… See the full description on the dataset page: https://huggingface.co/datasets/Tachyeon/audio-fingerprint-indian-bench.audio-fingerprint-indian-bench
Audio Fingerprinting Benchmark on Indian Classical Music
A reproducible, pre-registered benchmark of five audio-fingerprinting systems on the Saraga 1.5 corpus (Hindustani + Carnatic), plus a pre-registered training-recipe improvement to the NAFP baseline that achieves Bonferroni-significant gains on 1-second queries.
v0.7 · closed-world retrieval · 5 systems · 6 528 evaluation cells · pooled McNemar p = 3.18 × 10⁻⁶
TL;DR
5 systems benchmarked: Olaf, Dejavu… See the full description on the dataset page: https://huggingface.co/datasets/aryanBanwala/audio-fingerprint-indian-bench.Common_Voice_Indian_Accenttts-indian
TTS Indian Languages Dataset
Speech dataset for Text-to-Speech covering 6 Indian languages, collected and processed from YouTube.
Languages & Speakers
Speaker
Language
Gender
monihara_bengali
Bengali
Male
munir_kashmiri
Kashmiri
Male
nandini_gujarati
Gujarati
Female
sansri_kannada
Kannada
Female
tamil_pokkisham
Tamil
Male
teluguM
Telugu
Male
Pipeline
Audio was collected and processed through these stages:
YouTube Download — yt-dlp… See the full description on the dataset page: https://huggingface.co/datasets/grushaaaaa/tts-indian.Indian-Accent-Datasetindian_englishindian_ASR_2
Dataset Card for "indian_ASR_2"
More Information needed
codeswitch-pairs-lase-indian
Codeswitch Pairs LASE — Indian-accent held-out corpus
1369 held-out cross-script utterance pairs from 8 ElevenLabs Indian-English Multilingual voices. Surfaces the accent-conditional finding: off-the-shelf encoders cluster Indian-accent voices closely regardless of script, while Western voices show large script-conditional gaps.
Each row is one synthesized utterance with metadata; pairs are reconstructed at evaluation time by joining on voice_id (same voice, different script =… See the full description on the dataset page: https://huggingface.co/datasets/Praxel/codeswitch-pairs-lase-indian.Indian-Emotional-Speech-Corpus
Indian Emotional Speech Corpus
Dataset Description
This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry.
Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points.
Text spoken by all participants:
(happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indian-Emotional-Speech-Corpus.indian-english-hindi-tts-60min
Indian English + Hindi TTS Dataset
A small, heavily-curated Text-to-Speech corpus: 73.2 minutes (37.7 min Indian
English, 35.5 min Hindi) of single-speaker, studio-grade clips. Every clip's audio
was listened to and its transcript corrected against automated Sarvam ASR output;
resulting WER against the corrected text is 0.05% (en-IN) and 0.0% (hi-IN), showing
both very clean source audio and very accurate ASR. Built for the Sarvam AI ML &
Speech Data Pipeline assignment using a… See the full description on the dataset page: https://huggingface.co/datasets/auraCodes/indian-english-hindi-tts-60min.india-supreme-court-audioKAI-indian-emotional-speech-corpus
Indian Emotional Speech Corpus
Dataset Description
This dataset comprises high-quality audio recordings of Indian speakers reading a standardized 50-word paragraph in four distinct emotional tones — happy, sad, surprised, and angry.
Each recording is approximately 20–25 seconds long and includes the full paragraph with tone shifts at specific points.
Text spoken by all participants:
(happy tone) Last Monday was perfect—I got the job I’d been dreaming of! I screamed… See the full description on the dataset page: https://huggingface.co/datasets/sudhanshu12/KAI-indian-emotional-speech-corpus.northeast-india-voices
Northeast India Voices
A multilingual transcribed speech corpus covering five indigenous and regional languages of Northeast India: Khasi, Garo, Mizo, Nagamese, and Kokborok.
Dataset Summary
Language
Family
Utterances
Khasi
Austroasiatic
14,974
Nagamese
Indo-Aryan (Creole)
7,386
Mizo
Tibeto-Burman
7,250
Kokborok
Tibeto-Burman
3,278
Garo
Tibeto-Burman
1,612
Total
34,500
Data Collection
Recorded by native speakers across… See the full description on the dataset page: https://huggingface.co/datasets/MWirelabs/northeast-india-voices.indian_english_audio_2indian-tts-dataset
Indian TTS Dataset
A curated Text-to-Speech training dataset with high-quality audio clips,
transcriptions, and emotion labels for Indian English (en-IN) and Hindi (hi-IN).
Dataset Summary
Metric
Value
Total clips
125
English (en-IN)
96 clips
Hindi (hi-IN)
29 clips
Total duration
28.8 minutes
Sample rate
22050 Hz
Format
WAV (PCM 16-bit, mono)
Emotion Distribution
Emotion
Count
neutral
92
narrative
10
excited… See the full description on the dataset page: https://huggingface.co/datasets/champTUSHARg007/indian-tts-dataset.indian-hindi-female-rawPP4-indian-music-csiindian_ted_talks_chunkedindian-english-voiceIndian_Englsih_SSML_dataset_for_orpheus_fine_tuningsarvam-indian-eng-hin-tts
Indian English + Hindi TTS Dataset (emotion-tagged)
A curated, single-speaker-per-clip speech dataset for Text-to-Speech research,
covering Indian English and Hindi. Every clip is sourced from YouTube,
transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM.
Total: 82 clips, 55.6 minutes
Hindi: 28.8 min | Indian English: 26.8 min
Audio: mono, 24 kHz, 16-bit WAV
Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.English_India_Call_Center_Audio_Dataset_Dual_ChannelDataset Description:
This dataset is a large-scale collection of 30,320 processed English (India) dual-channel call center audio recordings, part of a broader multilingual conversational audio collection containing approximately 3,569,083 processed call center recordings across 54 languages, designed to support the development and training of advanced speech AI and conversational AI systems.
It consists of real-world customer and agent speech recordings collected from call center environments.… See the full description on the dataset page: https://huggingface.co/datasets/InfoBayAI/English_India_Call_Center_Audio_Dataset_Dual_Channel.indian_english_extended
