CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01fixie-ai /common_voice_17_0audio10M<n<100M18 likes197k downloads2y agoHugging Face02ropedia-ai /xperience-10mgated ⚠️ Important: If you have already submitted an access request but have not completed the required DocuSign agreement, your request will remain pending. Please complete signing and we will grant access once verified. Interactive Intelligence from Human Xperience Xperience-10M Dataset Summary Xperience-10M is a large-scale egocentric multimodal dataset of human experience for embodied AI, robotics, world models, and spatial… See the full description on the dataset page: https://huggingface.co/datasets/ropedia-ai/xperience-10m.3dvideo-classification1M<n<10M249 likes97k downloads5mo agoHugging Face03fixie-ai /covost2This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included in the dataset, which makes usage easier and allows browsing the samples using HF Dataset Viewer. The limitation of this method is that all audio samples of the EN_XX subsets are duplicated, as such the size of the dataset is larger. As such, not all the data is included: Only the validation and test subsets are available. From the XX_EN subsets, only fr, es, and zh-CN are included. audio1M<n<10M5 likes63k downloads2y agoHugging Face04ai4bharat /IndicVoicesgated IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates [23 December 2025] We now have 11,200 hours of transcribed data! 🎉 Overview INDICVOICES is a dataset of natural and spontaneous speech containing a total of 23.7K hours of read (8%), extempore (76%) and conversational (15%) audio from 51K speakers covering 400+ Indian districts and 22 languages. Of these 23.7K hours, 11.2K hours have… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/IndicVoices.audio1M<n<10M114 likes25k downloads3mo agoHugging Face05AISHELL /AISHELL-3AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be used to train multi-speaker Text-to-Speech (TTS) systems.The corpus contains roughly 85 hours of emotion-neutral recordings spoken by 218 native Chinese mandarin speakers and total 88035 utterances. Their auxiliary attributes such as gender, age group and native accents are explicitly marked and provided in the corpus. Accordingly, transcripts in… See the full description on the dataset page: https://huggingface.co/datasets/AISHELL/AISHELL-3.audiotext-to-speech10K<n<100K14 likes24k downloads3y agoHugging Face06shenyunhang /AISHELL-4 AISHELL-4 Identifier: SLR111 Summary: A Free Mandarin Multi-channel Meeting Speech Corpus, provided by Beijing Shell Shell Technology Co.,Ltd Category: Speech License: CC BY-SA 4.0 Downloads (use a mirror closer to you): train_L.tar.gz [7.0G] &nbsp; ( Training set of large room, 8-channel microphone array speech ) &nbsp; Mirrors: [US] &nbsp; [EU] &nbsp; [CN] &nbsp; train_M.tar.gz [25G] &nbsp; ( Training set of medium room, 8-channel microphone array speech ) &nbsp;… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-4.audio10K<n<100K2 likes14k downloads2y agoHugging Face07malaysia-ai /malaysian-youtube Malaysian Youtube Malaysian and Singaporean youtube channels, total up to 60k audio files with total 18.7k hours. URLs data at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube/data Notebooks at https://github.com/mesolitica/malaya-speech/tree/master/data/youtube How to load the data efficiently? import pandas as pd import json from datasets import Audio from torch.utils.data import DataLoader, Dataset chunks = 30 sr = 16000 class Train(Dataset):… See the full description on the dataset page: https://huggingface.co/datasets/malaysia-ai/malaysian-youtube.audio10K<n<100K5 likes12k downloads2y agoHugging Face08qyang1021 /AIR-Bench-Dataset AIR-Bench Arxiv: https://arxiv.org/html/2402.07729v1This is the AIR-Bench dataset download page.AIR-Bench encompasses two dimensions: foundation and chat benchmarks. The former consists of 19 tasks with approximately 19k single-choice questions. The latter one contains 2k instances of open-ended question-and-answer data.For how to run AIR-Bench, Please refer to AIR-Bench github page(https://github.com/OFA-Sys/AIR-Bench)(will be public soon). Data Sources(All come from… See the full description on the dataset page: https://huggingface.co/datasets/qyang1021/AIR-Bench-Dataset.audioquestion-answeringn<1K8 likes11k downloads2y agoHugging Face09ai4bharat /Rasagated Rasa: Towards Building an Expressive Multilingual Text-To-Speech Dataset for Indian Languages Funded by: Bhashini, Ministry of Electronics and Information Technology, Government of IndiaSupported by: EkStep Foundation and Nilekani Philanthropies Overview We introduce Rasa, the first high-quality multilingual expressive Text-to-Speech (TTS) dataset for any Indian language. It comprises a minimum of 20 hours per speaker with a target of covering a female and male… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Rasa.audiotext-to-speech1M<n<10M58 likes10k downloads4mo agoHugging Face10moonshine-ai /audio_samples_1kaudio0 likes9.1k downloads6mo agoHugging Face11ai4bharat /indicvoices_rgated IndicVoices-R: Multilingual, Multi-Speaker Speech Corpus for Indian TTS Dataset Summary IndicVoices-R (IV-R) is the largest multilingual Indian text-to-speech (TTS) dataset derived from an automatic speech recognition (ASR) dataset. It contains 1,704 hours of high-quality speech from 10,496 speakers across 22 Indian languages. This dataset is designed to enhance the development of robust Indian TTS models by providing diverse speaker demographics, natural… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/indicvoices_r.audiotext-to-speech100K<n<1M44 likes7.4k downloads2y agoHugging Face12AISHELL /AISHELL-4audio2 likes5.8k downloads3y agoHugging Face13hanamizuki-ai /genshin-voice-v3.3-mandarin Dataset Card for Genshin Voice Dataset Description Dataset Summary The Genshin Voice dataset is a text-to-voice dataset of different Genshin Impact characters unpacked from the game. Languages The text in the dataset is in Mandarin. Dataset Creation Source Data Initial Data Collection and Normalization The data was obtained by unpacking the Genshin Impact game. Who are the source language producers? The… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/genshin-voice-v3.3-mandarin.audiotext-to-speech10K<n<100K41 likes5.2k downloads4y agoHugging Face14krutrim-ai-labs /VoiceAgentBench VoiceAgentBench This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978). VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.audio1K<n<10K9 likes5k downloads7mo agoHugging Face15ai4bharat /Kathbathgated Kathbath Kathbath is an human-labeled ASR dataset containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India Languages Bengali Gujarati Kannada Hindi Malayalam Marathi Odia Punjabi Sanskrit Tamil Telugu Urdu Licensing Information The IndicSUPERB dataset is released under this licensing scheme: We do not own any of the raw text used in creating this dataset. The text data… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Kathbath.audio100K<n<1M28 likes4.5k downloads2y agoHugging Face16shenyunhang /AISHELL-3 AISHELL-3 Identifier: SLR93 Summary: Mandarin data, provided by Beijing Shell Shell Technology Co., Ltd. Category: Speech License: Apache License v.2.0 Downloads (use a mirror closer to you): data_aishell3.tgz [19G] &nbsp; (speech data and transcripts ) &nbsp; Mirrors: [US] &nbsp; [EU] &nbsp; [CN] &nbsp; About this resource:AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It can be… See the full description on the dataset page: https://huggingface.co/datasets/shenyunhang/AISHELL-3.audio10K<n<100K2 likes3.6k downloads2y agoHugging Face17tarteel-ai /everyayah﷽ Dataset Card for Tarteel AI's EveryAyah Dataset Dataset Summary This dataset is a collection of Quranic verses and their transcriptions, with diacritization, by different reciters. Supported Tasks and Leaderboards [Needs More Information] Languages The audio is in Arabic. Dataset Structure Data Instances A typical data point comprises the audio file audio, and its transcription called text. The duration… See the full description on the dataset page: https://huggingface.co/datasets/tarteel-ai/everyayah.audioautomatic-speech-recognition100K<n<1M42 likes3.2k downloads5d agoHugging Face18zeroweight-ai /ZeroSpeech ZeroSpeech A large synthetic Vietnamese speech corpus for ASR training: 9,867,987 utterances / 26,896 hours, spoken by 199,265 distinct voices, generated with ZeroTTS from web and conversational text. Every clip is 16 kHz mono FLAC, 1–30 s, paired with the exact text it was synthesized from. Fields field type description audio Audio(16 kHz) the waveform, FLAC-encoded text string the transcript — the exact string given to the TTS source string which… See the full description on the dataset page: https://huggingface.co/datasets/zeroweight-ai/ZeroSpeech.audioautomatic-speech-recognition1M<n<10M0 likes2.9k downloads24d agoHugging Face19mundo-ai /turn-benchmark-devgated TurnBench - Dev Set TurnBench is a benchmark for evaluating conversational turn-taking: end-of-turn and interruption detection on real annotated two-speaker conversations. This repository contains the development split: 38 English conversations, about 7.3 hours of audio, packaged as one row per conversation. Each row contains two time-aligned per-speaker audio streams plus three independent annotator tracks per speaker.… See the full description on the dataset page: https://huggingface.co/datasets/mundo-ai/turn-benchmark-dev.audiovoice-activity-detectionn<1K8 likes2.7k downloads1mo agoHugging Face20tarteel-ai /EA-DIaudio100K<n<1M8 likes2.6k downloads4y agoHugging Face21disco-eth /AIMEfrom datasets import load_dataset dataset = load_dataset('disco-eth/AIME') AIME: AI Music Evaluation Dataset The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo. The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset. The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset. The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.audio1K<n<10K9 likes2.5k downloads2y agoHugging Face22EZMONYI /music-ai-human-test-audio Interpretable AI and Human Music Evaluation Archive Research audio and versioned experiment outputs for an English graduation thesis. The audio archive is incomplete. Completed experiments and verified partial audio publications must not be confused with whole-project delivery completion. No blanket license is assigned to this mixed-source archive. Completed experiments and thesis The BC extension, expanded YuE Native30 evaluation, locked YuE Native30 scoring… See the full description on the dataset page: https://huggingface.co/datasets/EZMONYI/music-ai-human-test-audio.audio1K<n<10K0 likes2.5k downloads10d agoHugging Face23Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes2.5k downloads4mo agoHugging Face24ai4bharat /Shrutilipigated Shrutilipi Overview Shrutilipi is a labelled ASR corpus obtained by mining parallel audio and text pairs at the document scale from All India Radio news bulletins for 12 Indian languages: Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Sanskrit, Tamil, Telugu, Urdu. The corpus has over 6400 hours of data across all languages. This work is funded by Bhashini, MeitY and Nilekani Philanthropies Usage The datasets library… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Shrutilipi.audio1M<n<10M17 likes2.3k downloads2y agoHugging Face25fixie-ai /gigaspeechaudio10M<n<100M12 likes2.2k downloads2y agoHugging Face26pipecat-ai /smart-turn-data-v3.2-trainTraining dataset for Smart Turn v3.2. Thank you to the following contributors whose audio samples are included in this dataset: The Pipecat team Liva AI: https://www.theliva.ai/ Midcentury: https://www.midcentury.xyz/ MundoAI: https://mundoai.world/ Also, thank you to the following people for the CC-0 background noise sample data which has been used in this dataset: https://freesound.org/people/4team/sounds/214995/ https://freesound.org/people/tomhannen/sounds/698090/… See the full description on the dataset page: https://huggingface.co/datasets/pipecat-ai/smart-turn-data-v3.2-train.audio100K<n<1M11 likes2.2k downloads9mo agoHugging Face27deepsu /himalaya-ai-stt-datasetaudio10K<n<100K2 likes2.2k downloads9d agoHugging Face28ai-sage /TimeGround-1M TimeGround-1M Synthetic English audio dataset for time-aware speech understanding, covering temporal localization, temporal description, and timed summaries. Data Filtering We use 14k hours of audio from YODAS2 English shards, selected from a 24k-hour source pool after language- and silence-ratio filtering. Synthetic annotations were generated for three time-grounded tasks, then filtered through LLM-based verification, deterministic validity checks, and… See the full description on the dataset page: https://huggingface.co/datasets/ai-sage/TimeGround-1M.audioquestion-answering100K<n<1M7 likes2.1k downloads2mo agoHugging Face29laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face30awakening-ai /ReactNet ResponseNet ResponseNet is a large-scale dyadic video dataset designed for Online Multimodal Conversational Response Generation (OMCRG). It fills the gap left by existing datasets by providing high-resolution, split-screen recordings of both speaker and listener, separate audio channels, and word‑level textual annotations for both participants. Paper If you use this dataset, please cite: ResponseNet: A High‑Resolution Dyadic Video Dataset for Online Multimodal… See the full description on the dataset page: https://huggingface.co/datasets/awakening-ai/ReactNet.audio1K<n<10K3 likes2k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.