Vaidik7781/sarvam-tts-dataset
Sarvam TTS Training Dataset High-quality TTS training dataset built for expressive speech synthesis. Stats Total: 377 segments | 167.9 minutes English (en-IN): 191 segments | 84.3 minutes Hindi (hi-IN): 186 segments | 83.6 minutes Rejection rate: 26.9% after full manual human review of all 483 segments Emotion Distribution neutral: 259 | calm: 36 | sad: 32 | angry: 23 | happy: 12 | surprised: 6 | fearful: 6 | excited: 3 How it was… See the full description on the dataset page: https://huggingface.co/datasets/Vaidik7781/sarvam-tts-dataset.
Sarvam TTS Training Dataset
High-quality TTS training dataset built for expressive speech synthesis.
Stats
- Total: 377 segments | 167.9 minutes
- English (en-IN): 191 segments | 84.3 minutes
- Hindi (hi-IN): 186 segments | 83.6 minutes
- Rejection rate: 26.9% after full manual human review of all 483 segments
Emotion Distribution
neutral: 259 | calm: 36 | sad: 32 | angry: 23 | happy: 12 | surprised: 6 | fearful: 6 | excited: 3
How it was built
- Audio sourced from YouTube using yt-dlp
- Extracted to 16kHz mono WAV using FFmpeg
- Transcribed using Sarvam saaras:v3 ASR (transcribe + verbatim modes)
- Language validated using Sarvam /detect-language API
- Hindi segments translated using Sarvam /translate API
- Emotion tagged using Sarvam sarvam-30b LLM
- Quality scored 0-100 per segment
- Full manual human listening review of all segments
Columns
- audio: 16kHz mono WAV
- transcript: clean transcription
- language: en-IN or hi-IN
- emotion: 8 emotion tags
- style: speaking style
- quality_score: 0-100 composite score
- english_translation: English gloss for Hindi segments
- manually_corrected: whether human reviewed
Pipeline Code
https://github.com/Vaidik-7781/sarvam-tts-pipeline
