datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.trn-indcnfr-hi-pilot
trn-indcnfr-hi-pilot — Teacher-Output Cache for Ternary-ASR Distillation
Append-only teacher-output cache used to distill a ternary Hindi ASR student.
It stores, per audio clip, ONLY:
row_id (a deterministic source-shard/position pointer),
top-k (k=64) CTC logits (vocab id + log-prob per kept entry, blank always kept),
selected encoder hidden states (last-3 blocks: layers 14 / 15 / 16), fp16.
No audio and no ground-truth transcripts are stored or redistributed.… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/trn-indcnfr-hi-pilot.indic_tts_ml
Indic TTS Malayalam Speech Corpus
The Malayalam subset of Indic TTS Corpus, taken from
this Kaggle database. The corpus contains
one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given
in the repository.
indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
IndicCMix
IndicCMix
Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph.
This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/lakshay1234t/indic-diarbench.IndicContextEval
IndicContextEval
A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Code and resources: https://github.com/AI4Bharat/IndicContextEval
Dataset at a glance
Languages
Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu
Speakers
555
Duration
55.93 h
Utterances
16,884
Domains
23 professional domains
Speech styles
Read, Extempore
Prompt levels
L0–L6 (7 levels)… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicContextEval.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.Indic_ASR_Eval
Indic ASR Eval
A curated evaluation set for Indic-language automatic speech recognition.
100 samples are sampled (seed = 42) from each (source dataset × language)
cell of seven public Indic ASR corpora. Each source corpus is published
as its own dataset config with a single test split, at 16 kHz.
Rows: 6,169 across 7 configs
Total audio: ~13.3 hours
Sampling rate: 16 kHz (mono)
Split: test (single split in every config)
Configs
Config
Rows
Notes
kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.sravaani-indic-diarbench-oracle-v1
SraVaani Indic DiarBench Oracle ASR
This is a portable evaluation-only oracle-turn view derived from
sarvamai/indic-diarbench at the
immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only.
Do not use these turns for fine-tuning if Indic DiarBench will remain an
external benchmark. Training on this export contaminates the test set.
Configurations
Config
Test rows
Audio hours
Purpose
primary
2,417
3.8233
Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.reazonspeech-v2-quality-index
ReazonSpeech v2 Quality Index
Quality metadata for 21,932,215 ReazonSpeech v2 utterances.
It joins the following two source analyses by exact audio path:
ayousanz/reazon-speech-v2-all-speechMOS-analyze/audio_analysis_results_speechMOS.json
ayousanz/reazon-speech-v2-all-WAND-SNR-analyze/reazonspeech-all-wada-snr.json
The source repositories are not modified and this repository does not contain
the source audio.
Validation
Check
Count
SpeechMOS rows
21… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/reazonspeech-v2-quality-index.indian-english-hindi-tts-60min
Indian English + Hindi TTS Dataset
A small, heavily-curated Text-to-Speech corpus: 73.2 minutes (37.7 min Indian
English, 35.5 min Hindi) of single-speaker, studio-grade clips. Every clip's audio
was listened to and its transcript corrected against automated Sarvam ASR output;
resulting WER against the corrected text is 0.05% (en-IN) and 0.0% (hi-IN), showing
both very clean source audio and very accurate ASR. Built for the Sarvam AI ML &
Speech Data Pipeline assignment using a… See the full description on the dataset page: https://huggingface.co/datasets/auraCodes/indian-english-hindi-tts-60min.librispeech_asr_individualLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz,
prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read
audiobooks from the LibriVox project, and has been carefully segmented and aligned.87asr-context-induced-leakage
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
Overview
SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language.
The dataset contains total of 2283 training, 368 validation and 838 test samples.
Data Sample:
{'id': 1053,
'num_samples': 226560,
'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav',
'audio': {'path': 'train/10000267636955490843.wav',
'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Indic-LLM-Labs/Fleurs-Kn.sundanese-spontaneous-conversation-sample
Sundanese Spontaneous Conversation — Free Sample
Unscripted two-speaker Sundanese (su-ID) conversation recorded in
West Java, Indonesia. This is a free 30-minute sample of a larger
commercially licensable corpus.
Why this exists
Sundanese has approximately 40 million native speakers, yet
spontaneous conversational data is almost absent:
Resource
Type
Volume
Common Voice Spontaneous 4.0
Spontaneous, 2 speakers
0.62 h
OpenSLR SLR36 / SLR44
Read speech
—… See the full description on the dataset page: https://huggingface.co/datasets/Indospeech/sundanese-spontaneous-conversation-sample.indian-tts-dataset
Indian TTS Dataset
A curated Text-to-Speech training dataset with high-quality audio clips,
transcriptions, and emotion labels for Indian English (en-IN) and Hindi (hi-IN).
Dataset Summary
Metric
Value
Total clips
125
English (en-IN)
96 clips
Hindi (hi-IN)
29 clips
Total duration
28.8 minutes
Sample rate
22050 Hz
Format
WAV (PCM 16-bit, mono)
Emotion Distribution
Emotion
Count
neutral
92
narrative
10
excited… See the full description on the dataset page: https://huggingface.co/datasets/champTUSHARg007/indian-tts-dataset.indictelephony-bench
IndicTelephony-Bench v1.0
An open benchmark for speech recognition on Indian telephone speech: 25,393
human-curated utterances (30.2 hours) in nine languages, recorded
over live phone lines and released exactly as the line delivered them, at 8 kHz.
Most of it is code-mixed, much of it is short, and the business vocabulary a voice
agent acts on is tagged in the references.
At a glance
Utterances
25,393
Audio
30.18 hours, 8 kHz mono 16-bit PCM WAV… See the full description on the dataset page: https://huggingface.co/datasets/ConvoZenAI/indictelephony-bench.Indic-subtitler-audio_evals
Indic_audio_evals
As part of this project. We are evaluating our performance of various ASR models as well
in a benchmarking dataset, we have created in various languages. This benchmarking dataset
is more alligned to real-world use-cases rather than having any academic datasets.
About Dataset
Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals
This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.indic-audio-natural-conversations-sample
Dataset Card for Indic Audio Natural Conversations Sample Dataset
Dataset Details
Dataset Description
The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.Indonesian-Speech-Dataset
🎧 Indonesian Speech Dataset
The Indonesian Speech Dataset is a high-quality speech audio dataset designed to deliver structured and scalable audio data for AI-powered voice systems. It contains 162 hours of audio data across 821 files, provided in MP3 and WAV formats, with a total size of 210 MB. This well-curated audio dataset ensures balanced and representative voice data, with 51% female and 49% male speakers, and a broad age distribution from 18 to 50+ years. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/Speech-data/Indonesian-Speech-Dataset.sarvam-indian-eng-hin-tts
Indian English + Hindi TTS Dataset (emotion-tagged)
A curated, single-speaker-per-clip speech dataset for Text-to-Speech research,
covering Indian English and Hindi. Every clip is sourced from YouTube,
transcribed with Sarvam Saaras v3, and emotion-tagged via acoustic cues + a Sarvam LLM.
Total: 82 clips, 55.6 minutes
Hindi: 28.8 min | Indian English: 26.8 min
Audio: mono, 24 kHz, 16-bit WAV
Single speaker per clip, clean (no background music / overlapping speakers)… See the full description on the dataset page: https://huggingface.co/datasets/Abhi29112005/sarvam-indian-eng-hin-tts.indic-audio-dialog-sample
Dataset Card for Indic Dialog Sample Dataset
Dataset Details
Dataset Description
The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.
