CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vdivyasharma /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.audioaudio-classification1M<n<10M14 likes5.9k downloads8mo agoHugging Face02grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.2k downloads7mo agoHugging Face03ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes2.5k downloads14d agoHugging Face04aman-hf /indic_asr Indic ASR Unified Dataset Unified collection of Indian language ASR datasets for pretraining. Stats Total hours: 10,278 Total samples: 4,732,705 Languages: 1 Audio: 16kHz mono (mixed flac/mp3/wav) Languages Language Hours Samples hi2 10,278 4,732,705 Usage from datasets import load_dataset # Load all languages (streaming) ds = load_dataset("aman-hf/indic_asr", streaming=True, split="train") # Load specific language ds_hi =… See the full description on the dataset page: https://huggingface.co/datasets/aman-hf/indic_asr.automatic-speech-recognition10M<n<100M0 likes2.1k downloads7mo agoHugging Face05sarvamai /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.audioautomatic-speech-recognition1K<n<10K18 likes1.4k downloads1mo agoHugging Face06mrunmai18 /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.audioaudio-classification1M<n<10M0 likes985 downloads1mo agoHugging Face07thennal /indic_tts_ml Indic TTS Malayalam Speech Corpus The Malayalam subset of Indic TTS Corpus, taken from this Kaggle database. The corpus contains one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given in the repository. audiotext-to-speech1K<n<10K6 likes343 downloads4y agoHugging Face08ai4bharat /IndicCMixgated IndicCMix Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph. This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.audiotranslationn<1K1 likes311 downloads5mo agoHugging Face09Vinidapooh /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M0 likes308 downloads10d agoHugging Face10WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes222 downloads2y agoHugging Face11lakshay1234t /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/lakshay1234t/indic-diarbench.audioautomatic-speech-recognition1K<n<10K0 likes221 downloads26d agoHugging Face12ai4bharat /IndicContextEval IndicContextEval A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages Code and resources: https://github.com/AI4Bharat/IndicContextEval Dataset at a glance Languages Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu Speakers 555 Duration 55.93 h Utterances 16,884 Domains 23 professional domains Speech styles Read, Extempore Prompt levels L0–L6 (7 levels)… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicContextEval.audioautomatic-speech-recognition10K<n<100K3 likes194 downloads3mo agoHugging Face13WhissleAI /indicvoices_pa_tagged_transcripts Dataset Card for indicvoices_pa_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes142 downloads2y agoHugging Face14manojkumarcs /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.audioautomatic-speech-recognition1K<n<10K0 likes114 downloads1mo agoHugging Face15humyn-labs /Indic-High-Fidelity-MultiSpeaker-ASR Dataset Overview This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages. The dataset includes: Paired audio + timestamped transcripts Natural, non-scripted conversational speech Dual-speaker interactions Segment-level speaker annotations Regionally diverse accents Audio Specifications Format: WAV (PCM 16-bit) Sampling Rate: 16 kHz Channel: Mono Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.audioautomatic-speech-recognitionn<1K1 likes107 downloads6mo agoHugging Face16ayush-shunyalabs /Indic_ASR_Eval Indic ASR Eval A curated evaluation set for Indic-language automatic speech recognition. 100 samples are sampled (seed = 42) from each (source dataset × language) cell of seven public Indic ASR corpora. Each source corpus is published as its own dataset config with a single test split, at 16 kHz. Rows: 6,169 across 7 configs Total audio: ~13.3 hours Sampling rate: 16 kHz (mono) Split: test (single split in every config) Configs Config Rows Notes kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.audioautomatic-speech-recognition1K<n<10K1 likes99 downloads5mo agoHugging Face17WhissleAI /indicvoices_bn_tagged_transcripts Dataset Card for indicvoices_bn_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_bn_tagged_transcripts.audioautomatic-speech-recognitionn<1K0 likes78 downloads2y agoHugging Face18abhirl /sravaani-indic-diarbench-oracle-v1 SraVaani Indic DiarBench Oracle ASR This is a portable evaluation-only oracle-turn view derived from sarvamai/indic-diarbench at the immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only. Do not use these turns for fine-tuning if Indic DiarBench will remain an external benchmark. Training on this export contaminates the test set. Configurations Config Test rows Audio hours Purpose primary 2,417 3.8233 Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.audioautomatic-speech-recognition10K<n<100K0 likes75 downloads1mo agoHugging Face19Noothi /telugu-tech-indicf5-custom-voice 🎙️ Telugu Tech IndicF5 Custom Voice Dataset A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models. All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.audiotext-to-speech100K<n<1M0 likes67 downloads1mo agoHugging Face20Indic-LLM-Labs /Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language. The dataset contains total of 2283 training, 368 validation and 838 test samples. Data Sample: {'id': 1053, 'num_samples': 226560, 'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav', 'audio': {'path': 'train/10000267636955490843.wav', 'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Indic-LLM-Labs/Fleurs-Kn.audioautomatic-speech-recognition1K<n<10K0 likes61 downloads3y agoHugging Face21kurianbenoy /Indic-subtitler-audio_evals Indic_audio_evals As part of this project. We are evaluating our performance of various ASR models as well in a benchmarking dataset, we have created in various languages. This benchmarking dataset is more alligned to real-world use-cases rather than having any academic datasets. About Dataset Dataset Link in HuggingFace: kurianbenoy/Indic-subtitler-audio_evals This dataset contains audio file in .wav format and video file in .mp4. The respective groundtruth will be… See the full description on the dataset page: https://huggingface.co/datasets/kurianbenoy/Indic-subtitler-audio_evals.audioautomatic-speech-recognitionn<1K2 likes48 downloads2y agoHugging Face22snorbyte /indic-audio-natural-conversations-samplegated Dataset Card for Indic Audio Natural Conversations Sample Dataset Dataset Details Dataset Description The IndicAudioNaturalConversations Dataset is a multilingual, multichannel, source-separated conversational speech dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP):… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-natural-conversations-sample.audioaudio-to-audion<1K1 likes38 downloads1y agoHugging Face23Noothi /telugu-tech-indicf5-custom-voice-v2 🎙️ Telugu Technical Custom Voice Dataset A high-quality, single-speaker Telugu tech speech dataset designed for fine-tuning text-to-speech (TTS) models like IndicF5-TTS, F5-TTS, XTTS v2, VITS, and ElevenLabs Voice Cloning. 📊 Dataset Overview Total Clips: 676 WAV files Total Audio Duration: 70.61 minutes (1.18 hours / 4,236.54 seconds) Total Disk Size: 1.14 GB Average Clip Duration: 6.26 seconds (ranging 2.0s – 15.0s, optimal for TTS attention alignment) Audio… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice-v2.text-to-speechn<1K0 likes37 downloads2mo agoHugging Face24mahendraphd /Indic_Hindi-English_Parallel_Speechgated Dataset Access Information This dataset is provided for research and academic purposes. Access to the dataset is gated, and users must request permission before downloading. Dataset Summary This repository contains the Hindi–English Speech-to-Speech Translation (S2ST) dataset introduced in the paper: Benchmarking Hindi-to-English Direct Speech-to-Speech Translation with Synthetic Data The dataset is designed to support research on direct speech-to-speech translation… See the full description on the dataset page: https://huggingface.co/datasets/mahendraphd/Indic_Hindi-English_Parallel_Speech.audio-to-audio100K<n<1M2 likes35 downloads5mo agoHugging Face25snorbyte /indic-audio-dialog-samplegated Dataset Card for Indic Dialog Sample Dataset Dataset Details Dataset Description The IndicAudioDialog Sample Dataset is a multilingual, multichannel, source-separated, conversational speech dataset. It features human-voiced recordings of dialogues translated into 9 Indian languages (Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi) using GPT-4.1. The dataset contains over 30 hours of high-quality audio, recorded by native… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-audio-dialog-sample.audioaudio-to-audio1K<n<10K1 likes23 downloads1y agoHugging Face26WhissleAI /indicvoices_mr_tagged_transcripts Dataset Card for indicvoices_mr_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_mr_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes21 downloads2y agoHugging Face27AdityK2409 /common_voice_17_indic_te Ekacare/Common Voice 17 Indic Te Dataset Description This dataset contains 62 samples organized across multiple splits. The dataset includes audio data. Splits train: 62 samples Dataset Creation This dataset was created using StreamableDatasetManager on 2025-10-10T19:51:47.900631. Data Fields The dataset includes the following columns: client_id: String data audio: Audio data (16kHz sampling rate) md5_audio: String data duration:… See the full description on the dataset page: https://huggingface.co/datasets/AdityK2409/common_voice_17_indic_te.audioautomatic-speech-recognitionn<1K0 likes21 downloads11mo agoHugging Face28snorbyte /indic-text-audio-samplegated Dataset Card for Indic Text Audio Sample Dataset Dataset Details Dataset Description The IndicTextAudioSample Dataset is a multilingual, text-speech pair sample dataset. It features human-voiced recordings of dialogues in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte Language(s) (NLP): hi, ta, te, pa, ml, kn, bn, gu, mr License: CC BY 4.0… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-text-audio-sample.audioaudio-classification10K<n<100K0 likes8 downloads1y agoHugging Face29snorbyte /indic-tts-sample-snac-encodedgated Dataset Card for Indic TTS Sample SNAC Encoded Dataset Dataset Details Dataset Description The IndicTTSSampleSNACEncoded Dataset is a multilingual, text-speech pair sample dataset. It features ~135 hours of human-voiced recordings of transcripts in nine Indian languages: Hindi, Tamil, Telugu, Punjabi, Malayalam, Kannada, Bengali, Gujarati, and Marathi, across multiple speakers and other metadata. Curated by: snorbyte Funded by: snorbyte Shared by: snorbyte… See the full description on the dataset page: https://huggingface.co/datasets/snorbyte/indic-tts-sample-snac-encoded.textaudio-to-audio10K<n<100K3 likes7 downloads1y agoHugging Face30Minutor /IndicContextEvalgatedNOTE: This is a duplicate repo of "https://huggingface.co/datasets/ai4bharat/IndicContextEval" - visit the reference dataset - for any new updates made after Jul 30, 2026. IndicContextEval A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages Code and resources: https://github.com/AI4Bharat/IndicContextEval Dataset at a glance Languages Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu… See the full description on the dataset page: https://huggingface.co/datasets/Minutor/IndicContextEval.audioautomatic-speech-recognition10K<n<100K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.