CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M115 likes25k downloads3mo agoHugging Face02ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M44 likes7.8k downloads2y agoHugging Face03vdivyasharma /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/vdivyasharma/IndicSynth.audioaudio-classification1M<n<10M14 likes5.2k downloads9mo agoHugging Face04grushaaaaa /indic-dialect-asr Indic Dialect ASR Dataset A multilingual ASR dataset covering 30 Indic dialect/languages with 2.8M+ samples. Usage from datasets import load_dataset # Load a specific language ds = load_dataset("grushaaaaa/indic-dialect-asr", "assamese", split="train") Features audio: 16kHz WAV audio sentence: Transcription text language: Language name source: Source dataset audioautomatic-speech-recognition1M<n<10M4 likes3.6k downloads8mo agoHugging Face05ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes3.4k downloads16d agoHugging Face06psk /indic-tts-966h Indic-TTS-966h Six-language Indian TTS corpus: ~966 hours of paired speech and text, 24 kHz mono WAV clips with sentence-level transcripts in native scripts (natural English code-switching preserved). Subset Clips Hours bengali 18,343 94.9 malayalam 30,548 192.5 marathi 34,327 213.4 punjabi 28,083 161.8 tamil 26,817 171.1 telugu 21,923 132.8 Columns: audio (24 kHz mono), file_name, transcript. One config per language: from datasets import load_dataset… See the full description on the dataset page: https://huggingface.co/datasets/psk/indic-tts-966h.audio100K<n<1M6 likes2.5k downloads2mo agoHugging Face07insanalamin /quran-terjemahan-indonesia-audio-by-ayah Audio Terjemahan Al-Qur'an Indonesia per Ayat Dataset ini berisi audio terjemahan Al-Qur'an bahasa Indonesia yang dipisahkan per ayat. Audio dibuat menggunakan Gemini TTS dari teks terjemahan Al-Qur'an bahasa Indonesia. Setiap file audio mewakili satu ayat. Isi Dataset 6.236 file audio WAV 6.236 file audio M4A/AAC terkompresi untuk aplikasi mobile Bahasa Indonesia Satu file audio untuk setiap ayat Format WAV dan M4A/AAC Disusun berdasarkan nomor surah dan nomor… See the full description on the dataset page: https://huggingface.co/datasets/insanalamin/quran-terjemahan-indonesia-audio-by-ayah.audio1K<n<10K0 likes2.2k downloads2mo agoHugging Face08sarvamai /indic-diarbench Indic DiarBench A multilingual joint diarization and ASR benchmark for Indian languages, spanning all 22 scheduled languages of India with approximately 108 hours of natural multi-speaker audio. Paper: Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages (Interspeech 2026) Dataset Summary Indic DiarBench is a conversational speech benchmark designed to evaluate speaker-attributed ASR in realistic multi-speaker settings for… See the full description on the dataset page: https://huggingface.co/datasets/sarvamai/indic-diarbench.audioautomatic-speech-recognition1K<n<10K18 likes1.4k downloads2mo agoHugging Face09SherryT997 /IndicTTS-Deepfake-Challenge-Data IndicTTS Deepfake Detection Challenge Participants will use the SherryT997/IndicTTS-Deepfake-Challenge-Data dataset, hosted on Hugging Face. This dataset consists of train and test splits and contains speech samples in 16 Indian languages, along with metadata for each audio clip. 🚀 Dataset to Use: SherryT997/IndicTTS-Deepfake-Challenge-Data This is the official dataset for the challenge and must be used for training and evaluation. 📌 Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SherryT997/IndicTTS-Deepfake-Challenge-Data.audio10K<n<100K1 likes1.3k downloads2y agoHugging Face10SPRINGLab /IndicTTS-Hindi Hindi Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Hindi monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Hindi Total Duration: ~10.33 hours (Male: 5.16 hours, Female: 5.18 hours) Audio Format: WAV Sampling Rate: 48000Hz… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS-Hindi.audiotext-to-speech10K<n<100K36 likes1.3k downloads2y agoHugging Face11SPRINGLab /IndicTTS-Englishaudio100K<n<1M2 likes1.1k downloads2y agoHugging Face12mrunmai18 /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/mrunmai18/IndicSynth.audioaudio-classification1M<n<10M0 likes1k downloads2mo agoHugging Face13RidheshBhati /Indic-total-New-TTS-Merge Indic Total TTS Merge Merged TTS dataset with 13 Indic languages. All audio clips are >= 3.0 seconds duration. Languages assamese, bengali, english, gujarati, hindi, kannada, malayalam, marathi, nepali, odia, punjabi, tamil, telugu Columns audio: Audio data text: Transcript text duration: Duration in seconds (all >= 3.0s) language: Language name audio100K<n<1M1 likes885 downloads7mo agoHugging Face14ThivyanRR /indic_monovoiceaudio100K<n<1M1 likes867 downloads1y agoHugging Face15vnahata /IndicDiarBench-speaker-retrieval Indic DiarBench speaker retrieval (MTEB) Indic DiarBench reshaped for speaker retrieval across the 22 scheduled languages of India: given a clip of one speaker, find other clips of that same speaker. Source: sarvamai/indic-diarbench at revision 92877ba, cc-by-4.0, official test split. Turns are cut by their annotated times, restricted to 2 to 15 seconds, and turns overlapping a different speaker are dropped. identity pairs the recording session with the speaker, because speaker… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/IndicDiarBench-speaker-retrieval.audioaudio-classification1K<n<10K0 likes848 downloads24d agoHugging Face16PrakashPask /tts_indicaudio10K<n<100K1 likes835 downloads10mo agoHugging Face17skbose /indian-english-nptel-v0audio100K<n<1M3 likes627 downloads2y agoHugging Face18makaveli10 /indic-superb-whisperaudio1K<n<10K0 likes619 downloads3y agoHugging Face19Praha-Labs /indic-Malayalam-PDaudio10K<n<100K0 likes610 downloads1y agoHugging Face20SPRINGLab /IndicTTS_Bengali Bengali Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Bengali monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Bengali Total Duration: ~15.06 hours (Male: 10.05 hours, Female: 5.01 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Bengali.audiotext-to-speech10K<n<100K3 likes547 downloads2y agoHugging Face21athaze /indicV-cleanedaudio100K<n<1M0 likes544 downloads2mo agoHugging Face22SPRINGLab /IndicTTS_Gujarati task_categories: - text-to-speech language: - gj pretty_name: Gujarati Indic TTS dataset size_categories: - n<1K Gujarati Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Gujarati monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Gujarati… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Gujarati.audio1K<n<10K3 likes516 downloads2y agoHugging Face23SPRINGLab /IndicVoices-R_Hindiaudiotext-to-speech10K<n<100K11 likes486 downloads2y agoHugging Face24dianavdavidson /indic-voices-hinglish-nospeakeroverlap-spon3.3-acronyms-fixed2audio100K<n<1M0 likes425 downloads2mo agoHugging Face25SPRINGLab /IndicTTS_Telugu Telugu Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Telugu monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Telugu Total Duration: ~8.74 hours (Male: 4.47 hours, Female: 4.27 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Telugu.audiotext-to-speech1K<n<10K7 likes423 downloads2y agoHugging Face26SPRINGLab /IndicTTS_Tamil Tamil Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Tamil monolingual recordings from both male and female speakers. The dataset contains high-quality speech recordings with corresponding text transcriptions, making it suitable for text-to-speech (TTS) research and development. Dataset Details Language: Tamil Total Duration: ~20.33 hours (Male: 10.3 hours, Female: 10.03 hours) Audio Format: WAV Sampling Rate:… See the full description on the dataset page: https://huggingface.co/datasets/SPRINGLab/IndicTTS_Tamil.audiotext-to-speech1K<n<10K4 likes409 downloads2y agoHugging Face27octava /indonesian-voice-transcription-1.4.9a-raudio10K<n<100K0 likes407 downloads2y agoHugging Face28agufsamudra /tts-indo agufsamudra/tts-indo agufsamudra/tts-indo is a preprocessed Indonesian speech dataset designed for training Text-to-Speech (TTS) models. This dataset is derived from the original Dataset TTS Indo available on Kaggle. Dataset Details Number of Examples: 114,036 Dataset Size: ~4GB Audio Sampling Rate: 16,000 Hz Features: audio: WAV audio recordings text: Transcription of the audio Data Structure Each sample in the dataset contains: audio: A dictionary… See the full description on the dataset page: https://huggingface.co/datasets/agufsamudra/tts-indo.audiotext-to-speech100K<n<1M8 likes402 downloads1y agoHugging Face29bc7ec356 /synthetic-speech-indicaudio100K<n<1M1 likes399 downloads5mo agoHugging Face30BH-Builds /indic-audio indic-audio A multi-speaker synthetic speech dataset for Hindi, Indian English, and Hinglish (Hindi in Latin script): 5,160 clips, 7.1 hours, 15 voices. Built to train goonj-1-82M, an edge-sized Indian-language TTS model. Summary Clips 5,160 Total audio 7.13 h (avg 5.0 s/clip) Voices 15 (14 named personas + 1 Hindi "language bed") Languages Hindi (Devanagari), Indian English, Hinglish (romanized) Format 44.1 kHz; WAV (bed_hindi) and MP3… See the full description on the dataset page: https://huggingface.co/datasets/BH-Builds/indic-audio.audiotext-to-speechn<1K0 likes399 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.