datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.indic_asr
Indic ASR Unified Dataset
Unified collection of Indian language ASR datasets for pretraining.
Stats
Total hours: 10,278
Total samples: 4,732,705
Languages: 1
Audio: 16kHz mono (mixed flac/mp3/wav)
Languages
Language
Hours
Samples
hi2
10,278
4,732,705
Usage
from datasets import load_dataset
# Load all languages (streaming)
ds = load_dataset("aman-hf/indic_asr", streaming=True, split="train")
# Load specific language
ds_hi =… See the full description on the dataset page: https://huggingface.co/datasets/aman-hf/indic_asr.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.IndicSynth
IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset
A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research
🏆 Outstanding Paper Award, ACL 2025
🧠 Overview
IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.indic_tts_ml
Indic TTS Malayalam Speech Corpus
The Malayalam subset of Indic TTS Corpus, taken from
this Kaggle database. The corpus contains
one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given
in the repository.
IndicCMix
IndicCMix
Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph.
This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.indic-dialect-asr
Indic Dialect ASR Dataset
A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples.
Usage
from datasets import load_dataset
# Load a specific language
ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train")
Features
audio: 16kHz WAV audio
sentence: Transcription text
language: Language name
source: Source dataset
indicvoices_hi_tagged_transcripts
Dataset Card for indicvoices_hi_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026)
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/lakshay1234t/indic-diarbench.IndicContextEval
IndicContextEval
A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Code and resources: https://github.com/AI4Bharat/IndicContextEval
Dataset at a glance
Languages
Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu
Speakers
555
Duration
55.93 h
Utterances
16,884
Domains
23 professional domains
Speech styles
Read, Extempore
Prompt levels
L0–L6 (7 levels)… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicContextEval.indicvoices_pa_tagged_transcripts
Dataset Card for indicvoices_pa_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.indic-diarbench
Indic DiarBench
A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio.
Dataset Summary
Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.Indic-High-Fidelity-MultiSpeaker-ASR
Dataset Overview
This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages.
The dataset includes:
Paired audio + timestamped transcripts
Natural, non-scripted conversational speech
Dual-speaker interactions
Segment-level speaker annotations
Regionally diverse accents
Audio Specifications
Format: WAV (PCM 16-bit)
Sampling Rate: 16 kHz
Channel: Mono
Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.Indic_ASR_Eval
Indic ASR Eval
A curated evaluation set for Indic-language automatic speech recognition.
100 samples are sampled (seed = 42) from each (source dataset × language)
cell of seven public Indic ASR corpora. Each source corpus is published
as its own dataset config with a single test split, at 16 kHz.
Rows: 6,169 across 7 configs
Total audio: ~13.3 hours
Sampling rate: 16 kHz (mono)
Split: test (single split in every config)
Configs
Config
Rows
Notes
kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.indicvoices_bn_tagged_transcripts
Dataset Card for indicvoices_bn_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.sravaani-indic-diarbench-oracle-v1
SraVaani Indic DiarBench Oracle ASR
This is a portable evaluation-only oracle-turn view derived from
sarvamai/indic-diarbench at the
immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only.
Do not use these turns for fine-tuning if Indic DiarBench will remain an
external benchmark. Training on this export contaminates the test set.
Configurations
Config
Test rows
Audio hours
Purpose
primary
2,417
3.8233
Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.telugu-tech-indicf5-custom-voice
🎙️ Telugu Tech IndicF5 Custom Voice Dataset
A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models.
All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance.
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language.
The dataset contains total of 2283 training, 368 validation and 838 test samples.
Data Sample:
{'id': 1053,
'num_samples': 226560,
'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav',
'audio': {'path': 'train/10000267636955490843.wav',
'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Indic-LLM-Labs/Fleurs-Kn.Indic-subtitler-audio_evals
Indic_audio_evals
As part of this project. We are evaluating our performance of various ASR models as well
in a benchmarking dataset, we have created in various languages. This benchmarking dataset
is more alligned to real-world use-cases rather than having any academic datasets.
About Dataset
Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals
This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.indic-audio-natural-conversations-sample
Dataset Card for Indic Audio Natural Conversations Sample Dataset
Dataset Details
Dataset Description
The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.telugu-tech-indicf5-custom-voice-v2
🎙️ Telugu Technical Custom Voice Dataset
A high-quality, single-speaker Telugu tech speech dataset designed for fine-tuning text-to-speech (TTS) models like IndicF5-TTS, F5-TTS, XTTS v2, VITS, and ElevenLabs Voice Cloning.
📊 Dataset Overview
Total Clips: 676 WAV files
Total Audio Duration: 70.61 minutes (1.18 hours / 4,236.54 seconds)
Total Disk Size: 1.14 GB
Average Clip Duration: 6.26 seconds (ranging 2.0s – 15.0s, optimal for TTS attention alignment)
Audio… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice-v2.Indic_Hindi-English_Parallel_Speech
Dataset Access Information
This dataset is provided for research and academic purposes. Access to the dataset is gated, and users must request permission before downloading.
Dataset Summary
This repository contains the Hindi–English Speech-to-Speech Translation (S2ST) dataset introduced in the paper:
Benchmarking Hindi-to-English Direct Speech-to-Speech Translation with Synthetic Data
The dataset is designed to support research on direct speech-to-speech translation… See the full description on the dataset page: https://huggingface.co/datasets/mahendraphd/Indic_Hindi-English_Parallel_Speech.indic-audio-dialog-sample
Dataset Card for Indic Dialog Sample Dataset
Dataset Details
Dataset Description
The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.indicvoices_mr_tagged_transcripts
Dataset Card for indicvoices_mr_tagged_transcripts
Dataset Description
This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks.
Languages
The dataset is primarily in Hindi.
Data Collection
The dataset was collected through automated processes and manual transcription.
Dataset Structure
The dataset contains:
Audio files (.wav format)
Transcriptions
Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_mr_tagged_transcripts.common_voice_17_indic_te
Ekacare/Common Voice 17 Indic Te
Dataset Description
This dataset contains 62 samples organized across multiple splits.
The dataset includes audio data.
Splits
train: 62 samples
Dataset Creation
This dataset was created using StreamableDatasetManager on 2025-10-10T19:51:47.900631.
Data Fields
The dataset includes the following columns:
client_id: String data
audio: Audio data (16kHz sampling rate)
md5_audio: String data
duration:… See the full description on the dataset page: https://huggingface.co/datasets/AdityK2409/common_voice_17_indic_te.indic-text-audio-sample
Dataset Card for Indic Text Audio Sample Dataset
Dataset Details
Dataset Description
The IndicTextAudioSample Dataset is a multilingual, text-speech pair sample dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte
Language(s) (NLP): hi, ta, te, pa, ml, kn, bn, gu, mr
License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-text-audio-sample.indic-tts-sample-snac-encoded
Dataset Card for Indic TTS Sample SNAC Encoded Dataset
Dataset Details
Dataset Description
The IndicTTSSampleSNACEncoded Dataset is a multilingual, text-speech pair sample dataset. It features ~135 hours of human-voiced recordings of transcripts in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi, across multiple speakers and other metadata.
Curated by: snorbyte
Funded by: snorbyte
Shared by: snorbyte… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-tts-sample-snac-encoded.IndicContextEvalNOTE: This is a duplicate repo of "https://huggingface.co/datasets/ai4bharat/IndicContextEval" - visit the reference dataset - for any new updates made after Jul 30, 2026.
IndicContextEval
A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages
Code and resources: https://github.com/AI4Bharat/IndicContextEval
Dataset at a glance
Languages
Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/IndicContextEval.
