datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.indic_asr
Indic ASR Unified Dataset
Unified collection of Indian language ASR datasets for pretraining.
Stats
Total hours: 10,278
Total samples: 4,732,705
Languages: 1
Audio: 16kHz mono (mixed flac/mp3/wav)
Languages
Language
Hours
Samples
hi2
10,278
4,732,705
Usage
from datasets import load_dataset
# Load all languages (streaming)
ds = load_dataset("aman-hf/indic_asr", streaming=True, split="train")
# Load specific language
ds_hi =… See the full description on the dataset page: https://huggingface.co/datasets/aman-hf/indic_asr.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.trn-indcnfr-hi-pilot
trn-indcnfr-hi-pilot — Teacher-Output Cache for Ternary-ASR Distillation
Append-only teacher-output cache used to distill a ternary Hindi ASR student.
It stores, per audio clip, ONLY:
row_id (a deterministic source-shard/position pointer),
top-k (k=64) CTC logits (vocab id + log-prob per kept entry, blank always kept),
selected encoder hidden states (last-3 blocks: layers 14 / 15 / 16), fp16.
No audio and no ground-truth transcripts are stored or redistributed.… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/trn-indcnfr-hi-pilot.indic_tts_ml
Indic TTS Malayalam Speech Corpus
The Malayalam subset of Indic TTS Corpus, taken from
this Kaggle database. The corpus contains
one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given
in the repository.
ePark_zu_yu_duan_wen_indigenous_language_essays
FormosanBank publication status
This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card.
FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays
Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.librivox-indonesia
Dataset Card for LibriVox Indonesia 1.0
Dataset Summary
The LibriVox Indonesia dataset consists of MP3 audio and a corresponding text file we generated from the public
domain audiobooks LibriVox. We collected only languages in Indonesia for this dataset.
The original LibriVox audiobooks or sound files' duration varies from a few minutes to a few hours. Each audio
file in the speech dataset now lasts from a few seconds to a maximum of 20 seconds.
We converted the… See the full description on the dataset page: https://huggingface.co/datasets/indonesian-nlp/librivox-indonesia.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
IndicCMix
IndicCMix
Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph.
This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.Indonesian-ASR-11-Class-Dataset
Indonesian ASR 11-Class Dataset
Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts.
Dataset summary
Audio files: 104,500 WAV files
Real/human recordings: 104,368
Synthetic repair files: 132
Sentence classes: 11 Indonesian sentence categories
Canonical balanced sentence slots: 209 (11 categories × 19 retained slots)
Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs*
Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.microsoft-speech-corpus-indian
Microsoft Speech Corpus – Indian Languages
Dataset Description
This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript.
Attribution required: "Data provided by Microsoft and SpeechOcean.com"
⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.Sudan-MM
Sudan-MM: A Multimodal Dataset of Sudanese Arabic
Sudan-MM is the first publicly available multimodal dataset for Sudanese Arabic (السودانية), a low-resource dialect with no prior paired image-caption, video-caption, or voice-caption data. It was produced through a competitive shared task held in 2025, where five teams collected and annotated media depicting everyday Sudanese life.
Each item in the dataset pairs a visual or video recording with:
a written caption in Modern Standard… See the full description on the dataset page: https://huggingface.co/datasets/IndabaXSudan/Sudan-MM.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/lakshay1234t/indic-diarbench.indicvoices_hi_tagged_transcripts
Dataset Card for indicvoices_hi_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.IndicContextEval
IndicContextEval
A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Code and resources: https://github.com/AI4Bharat/IndicContextEval
Dataset at a glance
Languages
Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu
Speakers
555
Duration
55.93 h
Utterances
16,884
Domains
23 professional domains
Speech styles
Read, Extempore
Prompt levels
L0–L6 (7 levels)… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicContextEval.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.indicvoices_pa_tagged_transcripts
Dataset Card for indicvoices_pa_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.Indic_ASR_Eval
Indic ASR Eval
A curated evaluation set for Indic-language automatic speech recognition.
100 samples are sampled (seed = 42) from each (source dataset × language)
cell of seven public Indic ASR corpora. Each source corpus is published
as its own dataset config with a single test split, at 16 kHz.
Rows: 6,169 across 7 configs
Total audio: ~13.3 hours
Sampling rate: 16 kHz (mono)
Split: test (single split in every config)
Configs
Config
Rows
Notes
kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.sravaani-indic-diarbench-oracle-v1
SraVaani Indic DiarBench Oracle ASR
This is a portable evaluation-only oracle-turn view derived from
sarvamai/indic-diarbench at the
immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only.
Do not use these turns for fine-tuning if Indic DiarBench will remain an
external benchmark. Training on this export contaminates the test set.
Configurations
Config
Test rows
Audio hours
Purpose
primary
2,417
3.8233
Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.telugu-tech-indicf5-custom-voice
🎙️ Telugu Tech IndicF5 Custom Voice Dataset
A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models.
All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.librivox-indonesia
Dataset Card for LibriVox Indonesia 1.0
Dataset Summary
The LibriVox Indonesia dataset consists of MP3 audio and a corresponding text file we generated from the public
domain audiobooks LibriVox. We collected only languages in Indonesia for this dataset.
The original LibriVox audiobooks or sound files' duration varies from a few minutes to a few hours. Each audio
file in the speech dataset now lasts from a few seconds to a maximum of 20 seconds.
We converted the… See the full description on the dataset page: https://huggingface.co/datasets/cahya/librivox-indonesia.reazonspeech-v2-quality-index
ReazonSpeech v2 Quality Index
Quality metadata for 21,932,215 ReazonSpeech v2 utterances.
It joins the following two source analyses by exact audio path:
ayousanz/reazon-speech-v2-all-speechMOS-analyze/audio_analysis_results_speechMOS.json
ayousanz/reazon-speech-v2-all-WAND-SNR-analyze/reazonspeech-all-wada-snr.json
The source repositories are not modified and this repository does not contain
the source audio.
Validation
Check
Count
SpeechMOS rows
21… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/reazonspeech-v2-quality-index.indian-english-hindi-tts-60min
Indian English + Hindi TTS Dataset
A small, heavily-curated Text-to-Speech corpus: 73.2 minutes (37.7 min Indian
English, 35.5 min Hindi) of single-speaker, studio-grade clips. Every clip's audio
was listened to and its transcript corrected against automated Sarvam ASR output;
resulting WER against the corrected text is 0.05% (en-IN) and 0.0% (hi-IN), showing
both very clean source audio and very accurate ASR. Built for the Sarvam AI ML &
Speech Data Pipeline assignment using a… See the full description on the dataset page: https://huggingface.co/datasets/auraCodes/indian-english-hindi-tts-60min.librispeech_asr_individualLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz,
prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read
audiobooks from the LibriVox project, and has been carefully segmented and aligned.87asr-context-induced-leakage
When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR
Overview
SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language.
The dataset contains total of 2283 training, 368 validation and 838 test samples.
Data Sample:
{'id': 1053,
'num_samples': 226560,
'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav',
'audio': {'path': 'train/10000267636955490843.wav',
'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Indic-LLM-Labs/Fleurs-Kn.
