datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
audio_assignmentSarvam_Assignment
Sarvam Bilingual Dataset: Indian English + Malayalam
A curated bilingual speech dataset for model training, covering Indian English (en-IN) and Malayalam (ml-IN). Built using Sarvam AI's speech APIs — Batch STT diarization, Saaras v3 ASR, and the sarvam-105b LLM — with an emphasis on audio quality, clean speaker segmentation, and rich per-clip metadata.
GitHub: gourilaxmi/Sarvam_assignment
Dataset Summary
Language
Clips
Total Duration
Avg Duration
Avg SNR… See the full description on the dataset page: https://huggingface.co/datasets/gouri005/Sarvam_Assignment.sarvam-ai-assignment
Indic TTS Dataset — English + Hindi (single-speaker)
A small, high-quality text-to-speech dataset of 120 single-speaker clips
(~55.7 min) in Indian English and Hindi, built from YouTube
sources with the Sarvam AI APIs. Each clip is a
~30-second mono / 16 kHz, loudness-normalized segment with an accurate
transcript and an emotion/style tag.
Composition
Language
Clips
Duration
Indian English (en-IN)
60
28.0 min
Hindi (hi-IN)
60
27.7 min
Total
120
55.7… See the full description on the dataset page: https://huggingface.co/datasets/yadynesh/sarvam-ai-assignment.Shubham_sarvam_tts_assignment
Sarvam Indian TTS Dataset — 63 Minutes of Indian English & Hindi Speech
A curated, annotated speech dataset for Text-to-Speech (TTS) model training, built as part of the Sarvam AI ML & Speech Data Pipeline internship screening assignment. Contains 63 minutes of clean, single-speaker audio split across Indian English (en-IN) and Hindi (hi-IN), with accurate transcriptions and LLM-generated emotion/style annotations.
Dataset Statistics
Split
Clips
Duration… See the full description on the dataset page: https://huggingface.co/datasets/ss8816/Shubham_sarvam_tts_assignment.
