datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.indian-english-hindi-tts-60min
Indian English + Hindi TTS Dataset
A small, heavily-curated Text-to-Speech corpus: 73.2 minutes (37.7 min Indian
English, 35.5 min Hindi) of single-speaker, studio-grade clips. Every clip's audio
was listened to and its transcript corrected against automated Sarvam ASR output;
resulting WER against the corrected text is 0.05% (en-IN) and 0.0% (hi-IN), showing
both very clean source audio and very accurate ASR. Built for the Sarvam AI ML &
Speech Data Pipeline assignment using a… See the full description on the dataset page: https://huggingface.co/datasets/auraCodes/indian-english-hindi-tts-60min.indian-tts-dataset
Indian TTS Dataset
A curated Text-to-Speech training dataset with high-quality audio clips,
transcriptions, and emotion labels for Indian English (en-IN) and Hindi (hi-IN).
Dataset Summary
Metric
Value
Total clips
125
English (en-IN)
96 clips
Hindi (hi-IN)
29 clips
Total duration
28.8 minutes
Sample rate
22050 Hz
Format
WAV (PCM 16-bit, mono)
Emotion Distribution
Emotion
Count
neutral
92
narrative
10
excited… See the full description on the dataset page: https://huggingface.co/datasets/champTUSHARg007/indian-tts-dataset.sarvam-indian-eng-hin-tts
Indian English + Hindi TTS Dataset (emotion-tagged)
A curated, single-speaker-per-clip speech dataset for Text-to-Speech research,
covering Indian English and Hindi. Every clip is sourced from YouTube,
transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM.
Total: 82 clips, 55.6 minutes
Hindi: 28.8 min | Indian English: 26.8 min
Audio: mono, 24 kHz, 16-bit WAV
Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.indian-tts-emotion-60min
indian-tts-emotion-60min
A small, carefully curated text-to-speech dataset: ~68 minutes of clean,
single-speaker-per-clip audio in Indian English (en-IN) and Hindi (hi-IN), sourced
from YouTube, with accurate transcriptions and per-clip emotion/style tags.
Built as a data-quality exercise: clips were filtered conservatively and a sample was
verified by listening rather than shipped straight from an automated pipeline.
Summary
Language
Clips
Duration (min)… See the full description on the dataset page: https://huggingface.co/datasets/sarthwa8/indian-tts-emotion-60min.
