datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.Sarvam_Assignment
Sarvam Bilingual Dataset: Indian English + Malayalam
A curated bilingual speech dataset for model training, covering Indian English (en-IN) and Malayalam (ml-IN). Built using Sarvam AI's speech APIs — Batch STT diarization, Saaras v3 ASR, and the sarvam-105b LLM — with an emphasis on audio quality, clean speaker segmentation, and rich per-clip metadata.
GitHub: gourilaxmi/Sarvam_assignment
Dataset Summary
Language
Clips
Total Duration
Avg Duration
Avg SNR… See the full description on the dataset page: https://huggingface.co/datasets/gouri005/Sarvam_Assignment.sarvam-indian-eng-hin-tts
Indian English + Hindi TTS Dataset (emotion-tagged)
A curated, single-speaker-per-clip speech dataset for Text-to-Speech research,
covering Indian English and Hindi. Every clip is sourced from YouTube,
transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM.
Total: 82 clips, 55.6 minutes
Hindi: 28.8 min | Indian English: 26.8 min
Audio: mono, 24 kHz, 16-bit WAV
Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.sarvam-tts-dataset
Sarvam TTS Dataset: Indian English and Hindi
A curated speech dataset for TTS model training, containing 396 clips in Indian English (en-IN) and Hindi (hi-IN) totalling 57.3 minutes (3440 seconds).
Pipeline source code: https://github.com/Ayush147258/sarvam-tts-dataset
Dataset Summary
Stat
Value
Total clips
396
Total duration
57.3 minutes (3440 seconds)
Clip duration range
5.0s to 26.5s
Mean clip duration
8.7s
Median clip duration
7.5s… See the full description on the dataset page: https://huggingface.co/datasets/ayush712145/sarvam-tts-dataset.sarvam-tts-dataset
Sarvam TTS Dataset
A curated dataset of clean, single-speaker audio clips in Indian English and Hindi,
assembled for training and evaluating text-to-speech (TTS) and speech recognition models.
Each clip is 45–60 seconds, loudness-normalized to broadcast standard, and tagged with a
primary emotion label by a human reviewer.
Dataset Summary
Total clips
72
Total duration
62.5 minutes
Languages
Indian English (EN), Hindi (HI)
Created
2026-06-18… See the full description on the dataset page: https://huggingface.co/datasets/nitya2405/sarvam-tts-dataset.sarvam-ai-assignment
Indic TTS Dataset — English + Hindi (single-speaker)
A small, high-quality text-to-speech dataset of 120 single-speaker clips
(~55.7 min) in Indian English and Hindi, built from YouTube
sources with the Sarvam AI APIs. Each clip is a
~30-second mono / 16 kHz, loudness-normalized segment with an accurate
transcript and an emotion/style tag.
Composition
Language
Clips
Duration
Indian English (en-IN)
60
28.0 min
Hindi (hi-IN)
60
27.7 min
Total
120
55.7… See the full description on the dataset page: https://huggingface.co/datasets/yadynesh/sarvam-ai-assignment.
