CoolFace
Datasetpublic

Vaidik7781/sarvam-tts-dataset

Sarvam TTS Training Dataset High-quality TTS training dataset built for expressive speech synthesis. Stats Total: 377 segments | 167.9 minutes English (en-IN): 191 segments | 84.3 minutes Hindi (hi-IN): 186 segments | 83.6 minutes Rejection rate: 26.9% after full manual human review of all 483 segments Emotion Distribution neutral: 259 | calm: 36 | sad: 32 | angry: 23 | happy: 12 | surprised: 6 | fearful: 6 | excited: 3 How it was… See the full description on the dataset page: https://huggingface.co/datasets/Vaidik7781/sarvam-tts-dataset.

sourceHugging Faceupdated 4mo agoView on Hugging Face
0likes11downloads
Dataset Card

Sarvam TTS Training Dataset

High-quality TTS training dataset built for expressive speech synthesis.

Stats

  • Total: 377 segments | 167.9 minutes
  • English (en-IN): 191 segments | 84.3 minutes
  • Hindi (hi-IN): 186 segments | 83.6 minutes
  • Rejection rate: 26.9% after full manual human review of all 483 segments

Emotion Distribution

neutral: 259 | calm: 36 | sad: 32 | angry: 23 | happy: 12 | surprised: 6 | fearful: 6 | excited: 3

How it was built

  • Audio sourced from YouTube using yt-dlp
  • Extracted to 16kHz mono WAV using FFmpeg
  • Transcribed using Sarvam saaras:v3 ASR (transcribe + verbatim modes)
  • Language validated using Sarvam /detect-language API
  • Hindi segments translated using Sarvam /translate API
  • Emotion tagged using Sarvam sarvam-30b LLM
  • Quality scored 0-100 per segment
  • Full manual human listening review of all segments

Columns

  • audio: 16kHz mono WAV
  • transcript: clean transcription
  • language: en-IN or hi-IN
  • emotion: 8 emotion tags
  • style: speaking style
  • quality_score: 0-100 composite score
  • english_translation: English gloss for Hindi segments
  • manually_corrected: whether human reviewed

Pipeline Code

https://github.com/Vaidik-7781/sarvam-tts-pipeline