CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01awsaf49 /epic_kitchens_100 Motivation The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here. Source You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts Citation @INPROCEEDINGS{Damen2018EPICKITCHENS, title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset}, author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/awsaf49/epic_kitchens_100.videovoice-activity-detectionn<1K25 likes12k downloads2y agoHugging Face02kapturecx /bolAIndiagated bolAIndia Human-side speech from production call recordings, cut into utterance-level chunks by a two-engine VAD (Silero + TEN) and transcribed by third-party ASR providers. Each row keeps the transcript, the provider's confidence, and full provenance back to the source recording. Sources One config per transcription system, so their output stays separable. config (source_id) provider model hours rows shards vendor-a vendor-a undisclosed 420.03 480774… See the full description on the dataset page: https://huggingface.co/datasets/kapturecx/bolAIndia.audioautomatic-speech-recognition10M<n<100M1 likes12k downloads14m agoHugging Face03KRAFTON /Raon-OpenTTS-Pool Raon-OpenTTS-Pool Technical Report Raon-OpenTTS-Pool is a large-scale open English speech corpus for text-to-speech (TTS) training, constructed from 8 publicly available speech corpora and a set of web-sourced recordings. It is the training data behind Raon-OpenTTS, an open TTS model that performs on par with state-of-the-art closed-data systems. 615K hours of speech audio 239.7M speech segments 11 source datasets aggregated into a unified format All… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Pool.texttext-to-speech100M<n<1B43 likes9.8k downloads4mo agoHugging Face04KuofengGao /ADU-Bench ADU-Bench: Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models We provide the code for evaluation on Github. If you use ADU-Bench in your project, please kindly cite: @article{gao2024benchmarking, title={Benchmarking Open-ended Audio Dialogue Understanding for Large Audio-Language Models}, author={Gao, Kuofeng and Xia, Shu-Tao and Xu, Ke and Torr, Philip and Gu, Jindong}, journal={arXiv preprint arXiv:2412.05167}, year={2024} } audion<1K0 likes8.9k downloads2y agoHugging Face05KBLab /rixvox-v2 RixVox-v2: A Swedish parliamentary speech dataset RixVox-v2 is a parliamentary speech dataset spanning nearly 23000 hours of speech. The dataset was built by matching and force aligning speeches in parliamentary protocols to media recordings of debates. Each observation contains metadata about the speaker's name, gender, district, role, party affiliation, and the date the speech was given. We include identifiers for protocols, speeches and speakers that allow linking observations in… See the full description on the dataset page: https://huggingface.co/datasets/KBLab/rixvox-v2.audioautomatic-speech-recognition1M<n<10M12 likes6.7k downloads1y agoHugging Face06huseyin-karaca /fasttaudio1M<n<10M0 likes6k downloads12d agoHugging Face07khursani8 /maa Multilingual MFA-Aligned Speech Dataset (UNDER DEVELOPMENT) A large-scale multilingual speech dataset with word-level and phoneme-level alignments produced using the Montreal Forced Aligner (MFA). Dataset Description This dataset consolidates multiple speech corpora across various languages, all processed through MFA to provide precise phoneme and word alignments. Each sample includes the original audio, transcript, and detailed timing information for both words and… See the full description on the dataset page: https://huggingface.co/datasets/khursani8/maa.audioautomatic-speech-recognition10M<n<100M7 likes5.7k downloads7mo agoHugging Face08ai4bharat /Kathbathgated Kathbath Kathbath is an human-labeled ASR dataset containing 1,684 hours of labelled speech data across 12 Indian languages from 1,218 contributors located in 203 districts in India Languages Bengali Gujarati Kannada Hindi Malayalam Marathi Odia Punjabi Sanskrit Tamil Telugu Urdu Licensing Information The IndicSUPERB dataset is released under this licensing scheme: We do not own any of the raw text used in creating this dataset. The text data… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Kathbath.audio100K<n<1M28 likes5.5k downloads2y agoHugging Face09krutrim-ai-labs /VoiceAgentBench VoiceAgentBench This repository contains dataset for VoiceAgentBench, a large-scale speech benchmark introduced in “VoiceAgentBench: Are Voice Assistants Ready for Agentic Tasks?” (arXiv:2510.07978). VoiceAgentBench is designed to evaluate end-to-end speech-based agents in realistic, tool-driven settings. Unlike prior speech benchmarks that focus on transcription, intent detection, and speech question answering, this benchmark targets agentic reasoning from speech input, requiring… See the full description on the dataset page: https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench.audio1K<n<10K9 likes4.9k downloads7mo agoHugging Face10aranemini /central-kurdish-pseudolabel Central Kurdish → English Pseudo-Labeled Speech Translation Corpus Dataset Summary This repository contains a large-scale pseudo-labeled speech translation corpus for Central Kurdish (Sorani Kurdish). The dataset was automatically generated using a pipeline composed of: Speech segmentation Automatic Speech Recognition (ASR) Machine Translation (MT) The objective is to provide training data for end-to-end Speech-to-Text Translation (S2TT) in a language with very… See the full description on the dataset page: https://huggingface.co/datasets/aranemini/central-kurdish-pseudolabel.audioautomatic-speech-recognition1M<n<10M2 likes4.2k downloads3mo agoHugging Face11Digital-Divide-Data /khmer-speech-dataset Khmer ASR Cultural Dataset 727.94 hours of manually curated speech-text pairs by native speakers in the Khmer language about Cambodian cultural topics. On average, each recording is 8 seconds. Speaker metadata (gender, age group, and origin city) is provided. Language: Khmer (khm). Source(s): Native speakers from Cambodia (5 females, 7 males). The utterances were manually generated based on topics and subtopics listed in metadata. Domain(s): Cultural domain, with a total of 61… See the full description on the dataset page: https://huggingface.co/datasets/Digital-Divide-Data/khmer-speech-dataset.audioautomatic-speech-recognition100K<n<1M26 likes3.8k downloads3mo agoHugging Face12ksmashhero /IndicSynth IndicSynth: Indian Multilingual Audio Deepfake Detection & Anti-Spoofing Dataset A Large-Scale Multilingual Synthetic Speech Dataset for Low-Resource Indian Languages to facilitate audio deepfake detection and anti-spoofing research 🏆 Outstanding Paper Award, ACL 2025 🧠 Overview IndicSynth is a novel multilingual synthetic speech dataset designed to advance multilingual audio deepfake detection (ADD) and anti-spoofing research. It covers 12 low-resource Indian… See the full description on the dataset page: https://huggingface.co/datasets/ksmashhero/IndicSynth.audioaudio-classification1M<n<10M0 likes3.4k downloads16d agoHugging Face13Winderth /pl-kwsaudio100K<n<1M2 likes3.3k downloads3mo agoHugging Face14Ken-Z /Latin-Audio Dataset Summary Vox Classica is a Latin speech corpus of ~73 hours of audio, segmented into short audio clips by sentence. Vox Classica is a large-scale, ML-ready dataset of human-read Classical Latin. It was designed to address the absence of a publicly available human-read Latin corpus large enough for model training. Alignment and curation: Kaiyuan Zhao Language: Latin (Classical) Uses This dataset is built for training and evaluating speech processing models… See the full description on the dataset page: https://huggingface.co/datasets/Ken-Z/Latin-Audio.audiotext-to-speech10K<n<100K8 likes3.2k downloads2mo agoHugging Face15kwatcharasupat /dnr-v3-multilingualaudio0 likes3k downloads11mo agoHugging Face16ghanaopenai /kasem-speech-text-parallel This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/. Kasem Speech-Text Parallel Dataset Dataset Description This dataset contains 75990 parallel speech-text pairs for Kasem, a language spoken primarily in Ghana. The dataset consists of audio recordings paired with their… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/kasem-speech-text-parallel.audioautomatic-speech-recognition10K<n<100K0 likes2.9k downloads3mo agoHugging Face17kadirnar /voicehub-arena-seed-tts-eval VoiceHub Arena — full English Seed-TTS-Eval 35,904 synthesized WAV files: 33 model families × the same 1,088 target texts. The campaign completed on 15 September 2026 on one NVIDIA A100-SXM4 40 GB. All 198 shards and every WAV SHA256 were verified after backup. Interactive leaderboard and all audio samples · Source repository (access required). Contents audio_shards/<model>.tar: 33 WebDataset shards, each containing 1,088 original WAVs and matching JSON metadata.… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/voicehub-arena-seed-tts-eval.audiotext-to-speech10K<n<100K0 likes2.8k downloads8d agoHugging Face18Digital-Divide-Data /Kamba-ASR-Data-Subset-484H Kamba ASR Data Subset 484H Kamba speech dataset for automatic speech recognition. audioautomatic-speech-recognition100K<n<1M0 likes2.8k downloads1mo agoHugging Face19severyn-k /isolated-guitar-chords Isolated Guitar Chords Dataset Overview This dataset contains isolated guitar chord recordings designed for audio classification tasks such as chord recognition and real-time music analysis. The data was intentionally recorded in realistic acoustic conditions, including minor background sounds, to improve robustness during inference in real-world environments. Recording Setup Instrument: 🎸 Fender FA-15 3/4 Acoustic Guitar Recording method: Manual recording… See the full description on the dataset page: https://huggingface.co/datasets/severyn-k/isolated-guitar-chords.audioaudio-classification3 likes2.5k downloads10mo agoHugging Face20kresnik /zeroth_korean Zeroth-Korean Dataset Introduction The Zeroth-Korean dataset is a publicly available speech dataset created for Korean automatic speech recognition (ASR) research and development. This dataset is distributed under the CC BY 4.0 license, allowing anyone to use it freely. The goal of the Zeroth project is to make Korean speech recognition more widely accessible. Dataset Overview Total Data: Approximately 51.6 hours of training data and 1.2 hours of test data… See the full description on the dataset page: https://huggingface.co/datasets/kresnik/zeroth_korean.audio10K<n<100K23 likes2.4k downloads2y agoHugging Face21Shirali /ISSAI_KSC_335RS_v_1_1 Dataset Card for "ISSAI_KSC_335RS_v_1_1" Kazakh Speech Corpus (KSC) Identifier: SLR102 Summary: A crowdsourced open-source Kazakh speech corpus developed by ISSAI (330 hours) Category: Speech License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): ISSAI_KSC_335RS_v1.1_flac.tar.gz [19G] (speech, transcripts and metadata ) Mirrors: [US] [EU] [CN] About this resource: A crowdsourced open-source speech corpus for the Kazakh language. The KSC… See the full description on the dataset page: https://huggingface.co/datasets/Shirali/ISSAI_KSC_335RS_v_1_1.audioautomatic-speech-recognition100K<n<1M3 likes2.4k downloads4y agoHugging Face22kyutai /DailyTalkContiguous DailyTalkContiguous This repo contains a concatenated version of the DailyTalk dataset (official repo). Rather than having separate files for each speaker's turn, this uses a stereo file for each conversation. The two speakers in a conversation are put separately on the left and right channels. The dataset is annotated with word level timestamps. The original DailyTalk dataset and baseline code are freely available for academic use with CC-BY-SA 4.0 license, this dataset uses the… See the full description on the dataset page: https://huggingface.co/datasets/kyutai/DailyTalkContiguous.audio21 likes2.4k downloads2y agoHugging Face23Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes2.3k downloads4mo agoHugging Face24ky552 /cszs_fr_enThis dataset contains the French-English track of the benchmark from ICASSP 2024: Zero Resource Code-Switched Speech Benchmark Using Speech Utterance Pairs for Multiple Spoken Languages.Though the benchmark is originally designed to assess the semantic and syntactic abilities of the speech foundation models, you can also use this dataset for code-switching ASR. If you find this dataset helpful, please consider to cite the following paper: @INPROCEEDINGS{10446737, author={Huang, Kuan-Po and… See the full description on the dataset page: https://huggingface.co/datasets/ky552/cszs_fr_en.audio100K<n<1M0 likes2.2k downloads2y agoHugging Face25kensho /spgispeechgated Dataset Card for SPGISpeech Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Terms of Usage… See the full description on the dataset page: https://huggingface.co/datasets/kensho/spgispeech.audioautomatic-speech-recognition1M<n<10M42 likes2.2k downloads8mo agoHugging Face26kanuli1983 /japanese-listening-voicevox-backupaudio0 likes2k downloads19d agoHugging Face27KevinJustin /FormulaBank-28K FormulaBank-28K FormulaBank-28K is a deterministic procedural-audio corpus for audio representation pre-training. It contains 28,000 mono clips at 16 kHz, each exactly 10.24 seconds long. Configuration Version: C125-I224-v1 Formula classes: 125 Renderings per class: 224 Total clips: 28,000 Audio format: lossless 24-bit FLAC Source: frozen AudioPG-Atomic-H7-C224-R0-Clean FormulaBank manifest Each formula class specifies an acoustic rendering rule. Each rendering… See the full description on the dataset page: https://huggingface.co/datasets/KevinJustin/FormulaBank-28K.audiofeature-extraction10K<n<100K0 likes2k downloads20d agoHugging Face28KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face29kensho /SPGISpeech2.0gated Dataset Card for SPGISpeech 2.0 Dataset Details Dataset Overview We are excited to present SPGISpeech 2.0, a dataset suitable for speaker-tagged transcription in the financial domain. SPGISpeech 2.0 improves the diversity of applicable modeling tasks while maintaining the core characteristic of the original SPGISpeech dataset: audio snippets and their corresponding fully formatted text transcriptions, usable for end-to-end automatic speech recognition (ASR).… See the full description on the dataset page: https://huggingface.co/datasets/kensho/SPGISpeech2.0.audioautomatic-speech-recognition100K<n<1M4 likes1.9k downloads5mo agoHugging Face30kohei0209 /mls_hq_urgent_track1audio100K<n<1M0 likes1.9k downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.