datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.contextual_asr_benchmark
Synthetic Contextual ASR Benchmark (Indic)
Dataset Summary
This dataset is a Synthetic Contextual Automatic Speech Recognition (ASR) benchmark designed to evaluate and improve speech recognition systems in voice bot scenarios. It focuses on context-aware transcription, where the ASR model can leverage conversation history and agent prompts to better transcribe user responses.
The dataset covers the top 10 Indian languages, providing a diverse linguistic landscape for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/contextual_asr_benchmark.Sarvam_Assignment
Sarvam Bilingual Dataset: Indian English + Malayalam
A curated bilingual speech dataset for model training, covering Indian English (en-IN) and Malayalam (ml-IN). Built using Sarvam AI's speech APIs — Batch STT diarization, Saaras v3 ASR, and the sarvam-105b LLM — with an emphasis on audio quality, clean speaker segmentation, and rich per-clip metadata.
GitHub: gourilaxmi/Sarvam_assignment
Dataset Summary
Language
Clips
Total Duration
Avg Duration
Avg SNR… See the full description on the dataset page: https://huggingface.co/datasets/gouri005/Sarvam_Assignment.sarvam-tts-in-te-en
Indian English + Telugu Single-Speaker TTS Dataset (emotion-tagged)
Clean audio clips sourced from YouTube, transcribed with Sarvam ASR, segmented with
diarization, and labeled with emotion/style tags. Built as a data-quality / curation exercise.
"Single-speaker" means each clip contains exactly one speaker (verified by
diarization and speaker-embedding similarity). The dataset spans 9 distinct speakers
total (4 English, 5 Telugu), tracked via speaker_id.
Contents… See the full description on the dataset page: https://huggingface.co/datasets/AkCodes23/sarvam-tts-in-te-en.audiollm-evalsThis evaluation set contains ~100 questions in both text and audio format in Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu. We use this dataset internally at Sarvam to evaluate the performance of our audio models. We open-source this data to enable the research community to replicate the results mentioned in our Shuka blog.
By deisgn, the questions are sometimes vague, and the audio has noise and other inconsistencies, to measure the robustness of… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/audiollm-evals.sarvam-tts-dataset
Sarvam-style TTS Training Dataset (Telugu + Indian English)
A small, high-curation single-speaker-per-clip TTS dataset (~80 minutes, 193 clips) of
Telugu and Indian English speech. Every clip ships an accurate transcript and a
rich, Parler-TTS-style natural-language style/emotion description. Built as a take-home
where the emphasis is data quality and curation judgment, not pipeline code — "listen
to the data" was the core principle.
Clips: 193 (~25s each, always cut at silence… See the full description on the dataset page: https://huggingface.co/datasets/theshaikasad/sarvam-tts-dataset.sarvam-dub-benchmark-set
Sarvam Dubbing Benchmark Dataset
Dataset Description
Multilingual evaluation dataset for real-time dubbing and voice cloning benchmarking with focus on speaker similarity preservation across same-lingual and cross-lingual scenarios.
This dataset was used to benchmark production dubbing systems. Internal evaluations showed higher speaker similarity than ElevenLabs v3 and Cartesia Sonic under an identical scoring protocol.
Check out the Sarvam Dub blog for more… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/sarvam-dub-benchmark-set.sarvam-indian-eng-hin-tts
Indian English + Hindi TTS Dataset (emotion-tagged)
A curated, single-speaker-per-clip speech dataset for Text-to-Speech research,
covering Indian English and Hindi. Every clip is sourced from YouTube,
transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM.
Total: 82 clips, 55.6 minutes
Hindi: 28.8 min | Indian English: 26.8 min
Audio: mono, 24 kHz, 16-bit WAV
Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.sarvam-tts-dataset
Sarvam Indian TTS Dataset
A curated Text-to-Speech dataset containing Indian English and Hindi
single-speaker audio segments, built as part of the Sarvam AI hiring assignment.
Dataset Summary
Property
Value
Total duration
78.8 minutes
English (en-IN)
57.5 minutes
Hindi (hi-IN)
21.3 minutes
Total segments
175
Human reviewed
99.4%
Mean quality score
92.1/100
Mean SNR
49.0 dB
Mean ASR confidence
0.900
Dataset Card… See the full description on the dataset page: https://huggingface.co/datasets/Avishii0309/sarvam-tts-dataset.octopus-tts-sarvam-stylesarvam-tts-dataset
Sarvam TTS Dataset: Indian English and Hindi
A curated speech dataset for TTS model training, containing 396 clips in Indian English (en-IN) and Hindi (hi-IN) totalling 57.3 minutes (3440 seconds).
Pipeline source code: https://github.com/Ayush147258/sarvam-tts-dataset
Dataset Summary
Stat
Value
Total clips
396
Total duration
57.3 minutes (3440 seconds)
Clip duration range
5.0s to 26.5s
Mean clip duration
8.7s
Median clip duration
7.5s… See the full description on the dataset page: https://huggingface.co/datasets/ayush712145/sarvam-tts-dataset.sarvam-indian-tts-60min
Sarvam Indian TTS Dataset
Curated Indian English and Hindi/code-mixed speech segments for Text-to-Speech dataset preparation, built for the Sarvam AI assignment. Audio is sourced from manually reviewed YouTube videos; each row includes source metadata for auditability.
Dataset Stats
Split
Clips
Duration
Avg proxy SNR (dB)
Indian English (en-IN)
81
30.70 min
39.34
Hindi/code-mixed (hi-IN)
107
30.41 min
39.54
Total
188
61.11 min
39.45
All final… See the full description on the dataset page: https://huggingface.co/datasets/Rushabh3/sarvam-indian-tts-60min.sarvam-tts-dataset
Sarvam TTS Dataset
A curated dataset of clean, single-speaker audio clips in Indian English and Hindi,
assembled for training and evaluating text-to-speech (TTS) and speech recognition models.
Each clip is 45–60 seconds, loudness-normalized to broadcast standard, and tagged with a
primary emotion label by a human reviewer.
Dataset Summary
Total clips
72
Total duration
62.5 minutes
Languages
Indian English (EN), Hindi (HI)
Created
2026-06-18… See the full description on the dataset page: https://huggingface.co/datasets/nitya2405/sarvam-tts-dataset.Shubham_sarvam_tts_assignment
Sarvam Indian TTS Dataset — 63 Minutes of Indian English & Hindi Speech
A curated, annotated speech dataset for Text-to-Speech (TTS) model training, built as part of the Sarvam AI ML & Speech Data Pipeline internship screening assignment. Contains 63 minutes of clean, single-speaker audio split across Indian English (en-IN) and Hindi (hi-IN), with accurate transcriptions and LLM-generated emotion/style annotations.
Dataset Statistics
Split
Clips
Duration… See the full description on the dataset page: https://huggingface.co/datasets/ss8816/Shubham_sarvam_tts_assignment.sarvam-ai-assignment
Indic TTS Dataset — English + Hindi (single-speaker)
A small, high-quality text-to-speech dataset of 120 single-speaker clips
(~55.7 min) in Indian English and Hindi, built from YouTube
sources with the Sarvam AI APIs. Each clip is a
~30-second mono / 16 kHz, loudness-normalized segment with an accurate
transcript and an emotion/style tag.
Composition
Language
Clips
Duration
Indian English (en-IN)
60
28.0 min
Hindi (hi-IN)
60
27.7 min
Total
120
55.7… See the full description on the dataset page: https://huggingface.co/datasets/yadynesh/sarvam-ai-assignment.sarvam-tts-dataset
Sarvam TTS Training Dataset
High-quality TTS training dataset built for expressive speech synthesis.
Stats
Total: 377 segments | 167.9 minutes
English (en-IN): 191 segments | 84.3 minutes
Hindi (hi-IN): 186 segments | 83.6 minutes
Rejection rate: 26.9% after full manual human review of all 483 segments
Emotion Distribution
neutral: 259 | calm: 36 | sad: 32 | angry: 23 | happy: 12 | surprised: 6 | fearful: 6 | excited: 3
How it was… See the full description on the dataset page: https://huggingface.co/datasets/Vaidik7781/sarvam-tts-dataset.sarvam-tts-indian-60min
Sarvam TTS Indian 60min Dataset
Overview
164-clip, 63.85-minute single-speaker Indian speech dataset
for TTS training. Contains Indian English (en-IN) and
Hindi (hi-IN) audio with rich emotion and style annotations.
Audio Format
Format: WAV, 22050 Hz, mono
Loudness: Normalized to -23 LUFS (EBU R128)
Duration per clip: 10–35 seconds
Dataset Statistics
Total clips: 164
Total duration: 63.85 minutes
English (en-IN): 81 clips, 31.94… See the full description on the dataset page: https://huggingface.co/datasets/SidakChhabra/sarvam-tts-indian-60min.asr-yt-longform-eval-data-privatesarvam-dub-benchmark-setHF_sarvam_tts_dataset
