CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vdivyasharma /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.audioaudio-classification1M<n<10M14 likes5.2k downloads9mo agoHugging Face02grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.6k downloads8mo agoHugging Face03ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes3.4k downloads17d agoHugging Face04aman-hf /indic_asr Indic ASR Unified Dataset Unified collection of Indian language ASR datasets for pretraining. Stats Total hours: 10,278 Total samples: 4,732,705 Languages: 1 Audio: 16kHz mono (mixed flac/mp3/wav) Languages Language Hours Samples hi2 10,278 4,732,705 Usage from datasets import load_dataset # Load all languages (streaming) ds = load_dataset("aman-hf/indic_asr", streaming=True, split="train") # Load specific language ds_hi =… See the full description on the dataset page: https://huggingface.co/datasets/aman-hf/indic_asr.automatic-speech-recognition10M<n<100M0 likes1.8k downloads7mo agoHugging Face05sarvamai /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.audioautomatic-speech-recognition1K<n<10K18 likes1.4k downloads2mo agoHugging Face06mrunmai18 /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.audioaudio-classification1M<n<10M0 likes1k downloads2mo agoHugging Face07sulabhkatiyar /trn-indcnfr-hi-pilot trn-indcnfr-hi-pilot — Teacher-Output Cache for Ternary-ASR Distillation Append-only teacher-output cache used to distill a ternary Hindi ASR student. It stores, per audio clip, ONLY: row_id (a deterministic source-shard/position pointer), top-k (k=64) CTC logits (vocab id + log-prob per kept entry, blank always kept), selected encoder hidden states (last-3 blocks: layers 14 / 15 / 16), fp16. No audio and no ground-truth transcripts are stored or redistributed.… See the full description on the dataset page: https://huggingface.co/datasets/sulabhkatiyar/trn-indcnfr-hi-pilot.tabularautomatic-speech-recognitionn<1K0 likes404 downloads2d agoHugging Face08thennal /indic_tts_ml Indic TTS Malayalam Speech Corpus The Malayalam subset of Indic TTS Corpus, taken from this Kaggle database. The corpus contains one male and one female speaker, with a 2:1 ratio of samples due to missing files for the female speaker. The license is given in the repository. audiotext-to-speech1K<n<10K6 likes373 downloads4y agoHugging Face09FormosanBank /ePark_zu_yu_duan_wen_indigenous_language_essays FormosanBank publication status This audio is associated with XML published in the public FormosanBank corpus and uses the same license recorded in that XML: CC BY-NC-SA 4.0. View the published XML. Publication approval is recorded on the corresponding FormosanBank Basecamp card. FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays Commercial AI Use is prohibited without prior written permission. See the FormosanBank Terms of Use and AI Use Addendum.… See the full description on the dataset page: https://huggingface.co/datasets/FormosanBank/ePark_zu_yu_duan_wen_indigenous_language_essays.audioautomatic-speech-recognition1K<n<10K0 likes356 downloads2mo agoHugging Face10indonesian-nlp /librivox-indonesia Dataset Card for LibriVox Indonesia 1.0 Dataset Summary The LibriVox Indonesia dataset consists of MP3 audio and a corresponding text file we generated from the public domain audiobooks LibriVox. We collected only languages in Indonesia for this dataset. The original LibriVox audiobooks or sound files' duration varies from a few minutes to a few hours. Each audio file in the speech dataset now lasts from a few seconds to a maximum of 20 seconds. We converted the… See the full description on the dataset page: https://huggingface.co/datasets/indonesian-nlp/librivox-indonesia.automatic-speech-recognition1K<n<10K15 likes341 downloads2y agoHugging Face11Vinidapooh /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M0 likes320 downloads14d agoHugging Face12ai4bharat /IndicCMixgated IndicCMix Most Indic NLP data assumes people write in one script and one language at a time. Real chat looks nothing like that. You get Hindi words in Roman letters, English verbs in the middle of a Tamil sentence, and the same person switching scripts halfway through a paragraph. This dataset is an attempt to cover that actual messiness. For every English sentence, you get three different Indic renderings of it: one code-mixed in the native script, one clean native-script… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicCMix.audiotranslationn<1K1 likes307 downloads5mo agoHugging Face13Atika88 /Indonesian-ASR-11-Class-Dataset Indonesian ASR 11-Class Dataset Public Hugging Face repository for an Indonesian ASR corpus and its paper-supporting benchmark artifacts. Dataset summary Audio files: 104,500 WAV files Real/human recordings: 104,368 Synthetic repair files: 132 Sentence classes: 11 Indonesian sentence categories Canonical balanced sentence slots: 209 (11 categories × 19 retained slots) Public speaker labels: M1..M12, F1..F8, plus synthetic labels Ms*/Fs* Audio format: 16 kHz… See the full description on the dataset page: https://huggingface.co/datasets/Atika88/Indonesian-ASR-11-Class-Dataset.tabularautomatic-speech-recognition100K<n<1M0 likes306 downloads17d agoHugging Face14deepdml /microsoft-speech-corpus-indian Microsoft Speech Corpus – Indian Languages Dataset Description This dataset is a redistribution of the Microsoft Speech Corpus (Indian Languages) containing conversational and phrasal speech training and test data for Telugu, Tamil, and Gujarati languages. Each entry includes an audio recording and its corresponding transcript. Attribution required: "Data provided by Microsoft and SpeechOcean.com" ⚠️ License: This data is provided for research purposes only. Commercial… See the full description on the dataset page: https://huggingface.co/datasets/deepdml/microsoft-speech-corpus-indian.audioautomatic-speech-recognition100K<n<1M3 likes302 downloads7mo agoHugging Face15IndabaXSudan /Sudan-MM Sudan-MM: A Multimodal Dataset of Sudanese Arabic Sudan-MM is the first publicly available multimodal dataset for Sudanese Arabic (السودانية), a low-resource dialect with no prior paired image-caption, video-caption, or voice-caption data. It was produced through a competitive shared task held in 2025, where five teams collected and annotated media depicting everyday Sudanese life. Each item in the dataset pairs a visual or video recording with: a written caption in Modern Standard… See the full description on the dataset page: https://huggingface.co/datasets/IndabaXSudan/Sudan-MM.audioimage-to-text1K<n<10K2 likes266 downloads4mo agoHugging Face16lakshay1234t /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/lakshay1234t/indic-diarbench.audioautomatic-speech-recognition1K<n<10K0 likes234 downloads29d agoHugging Face17WhissleAI /indicvoices_hi_tagged_transcripts Dataset Card for indicvoices_hi_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_hi_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes225 downloads2y agoHugging Face18ai4bharat /IndicContextEval IndicContextEval A Benchmark for Evaluating Context Utilisation in Audio Large Language Models Across 8 Indic Languages Code and resources: https://github.com/AI4Bharat/IndicContextEval Dataset at a glance Languages Hindi, Bengali, Telugu, Marathi, Gujarati, Malayalam, Odia, Urdu Speakers 555 Duration 55.93 h Utterances 16,884 Domains 23 professional domains Speech styles Read, Extempore Prompt levels L0–L6 (7 levels)… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicContextEval.audioautomatic-speech-recognition10K<n<100K3 likes166 downloads3mo agoHugging Face19manojkumarcs /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for Indian languages. All annotations are human-corrected with time-aligned, speaker-attributed transcriptions. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/manojkumarcs/indic-diarbench.audioautomatic-speech-recognition1K<n<10K0 likes150 downloads2mo agoHugging Face20WhissleAI /indicvoices_pa_tagged_transcripts Dataset Card for indicvoices_pa_tagged_transcripts Dataset Description This dataset contains audio files and their corresponding transcriptions in Hindi for automatic speech recognition (ASR) tasks. Languages The dataset is primarily in Hindi. Data Collection The dataset was collected through automated processes and manual transcription. Dataset Structure The dataset contains: Audio files (.wav format) Transcriptions Duration information… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/indicvoices_pa_tagged_transcripts.audioautomatic-speech-recognition1K<n<10K0 likes139 downloads2y agoHugging Face21humyn-labs /Indic-High-Fidelity-MultiSpeaker-ASR Dataset Overview This dataset contains high-quality multi-speaker conversational audio recordings curated for Automatic Speech Recognition (ASR) research across multiple Indic languages. The dataset includes: Paired audio + timestamped transcripts Natural, non-scripted conversational speech Dual-speaker interactions Segment-level speaker annotations Regionally diverse accents Audio Specifications Format: WAV (PCM 16-bit) Sampling Rate: 16 kHz Channel: Mono Speech… See the full description on the dataset page: https://huggingface.co/datasets/humyn-labs/Indic-High-Fidelity-MultiSpeaker-ASR.audioautomatic-speech-recognitionn<1K1 likes114 downloads7mo agoHugging Face22ayush-shunyalabs /Indic_ASR_Eval Indic ASR Eval A curated evaluation set for Indic-language automatic speech recognition. 100 samples are sampled (seed = 42) from each (source dataset × language) cell of seven public Indic ASR corpora. Each source corpus is published as its own dataset config with a single test split, at 16 kHz. Rows: 6,169 across 7 configs Total audio: ~13.3 hours Sampling rate: 16 kHz (mono) Split: test (single split in every config) Configs Config Rows Notes kathbath… See the full description on the dataset page: https://huggingface.co/datasets/ayush-shunyalabs/Indic_ASR_Eval.audioautomatic-speech-recognition1K<n<10K1 likes97 downloads5mo agoHugging Face23abhirl /sravaani-indic-diarbench-oracle-v1 SraVaani Indic DiarBench Oracle ASR This is a portable evaluation-only oracle-turn view derived from sarvamai/indic-diarbench at the immutable revision 92877bad8aab6e598167d91c6ee02aa8ca6ede09. It contains Hindi and Telugu only. Do not use these turns for fine-tuning if Indic DiarBench will remain an external benchmark. Training on this export contaminates the test set. Configurations Config Test rows Audio hours Purpose primary 2,417 3.8233 Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhirl/sravaani-indic-diarbench-oracle-v1.audioautomatic-speech-recognition10K<n<100K0 likes85 downloads1mo agoHugging Face24Noothi /telugu-tech-indicf5-custom-voice 🎙️ Telugu Tech IndicF5 Custom Voice Dataset A 100% verified, clean, single-speaker Telugu Speech & Voice dataset specially formatted and phonetically cleaned for training and fine-tuning ai4bharat/IndicF5 and neural Text-to-Speech (TTS) models. All English technical terms, numbers, acronyms, and ASR mishearings have been converted into native Telugu phonetic script, cleaned of noise/brackets, and validated for optimal IndicF5 fine-tuning performance. 📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/Noothi/telugu-tech-indicf5-custom-voice.audiotext-to-speech100K<n<1M0 likes80 downloads2mo agoHugging Face25cahya /librivox-indonesia Dataset Card for LibriVox Indonesia 1.0 Dataset Summary The LibriVox Indonesia dataset consists of MP3 audio and a corresponding text file we generated from the public domain audiobooks LibriVox. We collected only languages in Indonesia for this dataset. The original LibriVox audiobooks or sound files' duration varies from a few minutes to a few hours. Each audio file in the speech dataset now lasts from a few seconds to a maximum of 20 seconds. We converted the… See the full description on the dataset page: https://huggingface.co/datasets/cahya/librivox-indonesia.automatic-speech-recognition1K<n<10K2 likes78 downloads3y agoHugging Face26ayousanz /reazonspeech-v2-quality-index ReazonSpeech v2 Quality Index Quality metadata for 21,932,215 ReazonSpeech v2 utterances. It joins the following two source analyses by exact audio path: ayousanz/reazon-speech-v2-all-speechMOS-analyze/audio_analysis_results_speechMOS.json ayousanz/reazon-speech-v2-all-WAND-SNR-analyze/reazonspeech-all-wada-snr.json The source repositories are not modified and this repository does not contain the source audio. Validation Check Count SpeechMOS rows 21… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/reazonspeech-v2-quality-index.tabularautomatic-speech-recognition10M<n<100M0 likes66 downloads4d agoHugging Face27auraCodes /indian-english-hindi-tts-60min Indian English + Hindi TTS Dataset A small, heavily-curated Text-to-Speech corpus: 73.2 minutes (37.7 min Indian English, 35.5 min Hindi) of single-speaker, studio-grade clips. Every clip's audio was listened to and its transcript corrected against automated Sarvam ASR output; resulting WER against the corrected text is 0.05% (en-IN) and 0.0% (hi-IN), showing both very clean source audio and very accurate ASR. Built for the Sarvam AI ML & Speech Data Pipeline assignment using a… See the full description on the dataset page: https://huggingface.co/datasets/auraCodes/indian-english-hindi-tts-60min.audiotext-to-speechn<1K0 likes65 downloads3mo agoHugging Face28Splend1dchan /librispeech_asr_individualLibriSpeech is a corpus of approximately 1000 hours of read English speech with sampling rate of 16 kHz, prepared by Vassil Panayotov with the assistance of Daniel Povey. The data is derived from read audiobooks from the LibriVox project, and has been carefully segmented and aligned.87audioautomatic-speech-recognition100K<n<1M2 likes64 downloads3y agoHugging Face29maikezu /asr-context-induced-leakage When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR Overview SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar… See the full description on the dataset page: https://huggingface.co/datasets/maikezu/asr-context-induced-leakage.audioautomatic-speech-recognition1K<n<10K0 likes63 downloads4mo agoHugging Face30Indic-LLM-Labs /Fleurs-KnThis is a filtered version of the Fleurs dataset only containing samples of Kannada language. The dataset contains total of 2283 training, 368 validation and 838 test samples. Data Sample: {'id': 1053, 'num_samples': 226560, 'path': '/home/ravi.naik/.cache/huggingface/datasets/downloads/extracted/e7c8b501d4e6892673b6dc291d42de48e7987b0d2aa6471066a671f686224ed1/10000267636955490843.wav', 'audio': {'path': 'train/10000267636955490843.wav', 'array': array([ 0. , 0.… See the full description on the dataset page: https://huggingface.co/datasets/Indic-LLM-Labs/Fleurs-Kn.audioautomatic-speech-recognition1K<n<10K0 likes62 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.