CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01H-Liu1997 /BEAT2audio1K<n<10K11 likes23k downloads3y agoHugging Face02verstar /MRSAudio MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive spatial environments, where auditory cues enable sound source localization in three-dimensional space. Despite the critical role of spatial audio in immersive technologies such as VR/AR, most existing multimodal datasets provide only monaural audio, which limits the development of spatial audio generation and understanding. To address… See the full description on the dataset page: https://huggingface.co/datasets/verstar/MRSAudio.audio100K<n<1M7 likes11k downloads11mo agoHugging Face03Hui519 /WildElder WILDELDER: A CHINESE ELDERLY SPEECH DATASET FROM THE WILD WITH FINE-GRAINED MANUAL ANNOTATIONS Paper: https://huggingface.co/papers/2510.09344Code: https://github.com/NKU-HLT/WildElder WildElder is a speech dataset focused on elderly scenarios. It contains raw audio and corresponding text annotations and can be used for ASR, speaker-related tasks, and front-/back-end speech processing research. The data was collected and cleaned from real-world environments to preserve diversity and… See the full description on the dataset page: https://huggingface.co/datasets/Hui519/WildElder.audioautomatic-speech-recognition10K<n<100K3 likes4.4k downloads5mo agoHugging Face04plnguyen2908 /AV-SpeakerBench AV-SpeakerBench Audiovisual QA benchmark with speaker-aware questions and aligned clips. This drop includes trimmed segments (audio-only, visual-only, audiovisual) plus annotations to probe fine-grained AV reasoning. Project page: https://plnguyen2908.github.io/AV-SpeakerBench-project-page/ Code & benchmarks: https://github.com/plnguyen2908/AV-SpeakerBench Paper: https://arxiv.org/abs/2512.02231 Files test.csv - original annotations and metadata with clip paths… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AV-SpeakerBench.audioquestion-answering1K<n<10K2 likes4.1k downloads9mo agoHugging Face05Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes3.9k downloads10mo agoHugging Face06Kukedlc /suno-ai-music-dataset Suno AI Music Dataset (Multi-Genre Curated) A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research. This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.audioaudio-classificationn<1K29 likes2.5k downloads4mo agoHugging Face07softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face08m-hamza-mughal /beat2-additional-annotations BEAT2 Official Release + Additional Annotations This is a fork of H-Liu1997/BEAT2 that adds annotations contributed by the RAG-Gesture (CVPR 2025) and MIBURI (CVPR 2026) projects. The base BEAT2-English data (motion, audio, TextGrids, semantic labels, pretrained motion-autoencoder weights) is inherited verbatim from upstream; the additional annotations from RAG-Gesture and MIBURI are pushed on top. Citations If you use only the original BEAT2 dataset, please cite… See the full description on the dataset page: https://huggingface.co/datasets/m-hamza-mughal/beat2-additional-annotations.audio1K<n<10K0 likes2.3k downloads3mo agoHugging Face09hzhongresearch /ahead_ds Another HEaring AiD DataSet (AHEAD-DS) Another HEaring AiD DataSet (AHEAD-DS) is an audio dataset labelled with audiologically relevant scene categories for hearing aids. Website Paper Code Dataset AHEAD-DS Dataset AHEAD-DS unmixed Models Description of data All files are encoded as single channel WAV, 16 bit signed, sampled at 16 kHz with 10 seconds per recording. Category Training Validation Testing All cocktail_party 934 134 266 1334 interfering_speakers… See the full description on the dataset page: https://huggingface.co/datasets/hzhongresearch/ahead_ds.audioaudio-classification1K<n<10K0 likes2k downloads8mo agoHugging Face10KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face11BAAI /Chinese-LiPS Chinese-LiPS: A Chinese audio-visual speech recognition dataset with Lip-reading and Presentation Slides ⭐ Introduction The Chinese-LiPS dataset is a multimodal dataset designed for audio-visual speech recognition (AVSR) in Mandarin Chinese. This dataset combines speech, video, and textual transcriptions to enhance automatic speech recognition (ASR) performance, especially in educational and instructional scenarios. 🚀 Dataset Details Total Duration:… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Chinese-LiPS.audioautomatic-speech-recognition10K<n<100K12 likes1.5k downloads10mo agoHugging Face12ai4bharat /MANGO MANGO: A Corpus of Human Ratings for Speech MANGO (MUSHRA Assessment corpus using Native listeners and Guidelines to understand human Opinions at scale) is the first large-scale dataset designed for evaluating Text-to-Speech (TTS) systems in Indian languages. Key Features: 255,150 human ratings of TTS-generated outputs and ground-truth human speech. Covers two major Indian languages: Hindi & Tamil, and English. Based on the MUSHRA (Multiple Stimuli with Hidden Reference… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/MANGO.audiotext-to-speech10K<n<100K6 likes1.3k downloads1y agoHugging Face13Silasimo /SynthGT SynthGT A Synthetic Solo-Singing Dataset for Singing-Oriented Forced Alignment Authors Silas Antonisen, Iván López-Espejo Associated paper submitted to IEEE Transactions on Audio, Speech and Language Processing. Overview SynthGT (Synthetic Ground Truth) is a synthetic English solo-singing dataset containing 4,900 singing performances with automatically generated phoneme boundary annotations. The dataset was created through music… See the full description on the dataset page: https://huggingface.co/datasets/Silasimo/SynthGT.audioautomatic-speech-recognition1K<n<10K1 likes1.3k downloads2mo agoHugging Face14MahiA /UrbanSound8K UrbanSound8K This is an audio classification dataset for Sound Event Classification. Classes = 10   ,   Split = Ten-Fold Structure audios folder contains audio files. csv_files folder contains CSV files for ten-fold cross-validation. To perform cross-validation on fold 1, train_1.csv will be used for the training split and test_1.csv for the testing split, with the same pattern followed for the other folds. To perform training and testing witout cross-validation, use… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/UrbanSound8K.audio10K<n<100K2 likes1.3k downloads2y agoHugging Face15beshiribrahim /tigre-hubert-dataaudio1K<n<10K0 likes1.2k downloads2mo agoHugging Face16Pekkpo /AVSR-Vietnamese-Datasetaudio1K<n<10K2 likes1.1k downloads3mo agoHugging Face17fawzanaramam /the-truthaudio100K<n<1M0 likes926 downloads2y agoHugging Face18strikersoft /strikerData 🎧 StrikerData Overview StrikerData is an audio dataset developed by Strikersoft for research and development in audio and speech technologies.It contains human speech, environmental noise, and other sound types. The dataset is available for non-commercial use only, except for the company Strikersoft. Category Percentage of Total Dataset Clean human speech 20% Distorted speech 15% Human-made noise 15% Non-human noise 50% ⚖️ License… See the full description on the dataset page: https://huggingface.co/datasets/strikersoft/strikerData.audio10K<n<100K2 likes831 downloads8mo agoHugging Face19nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes802 downloads10mo agoHugging Face20Mohammed01 /ArFakegated ArFake-Dataset ARFAKE: A Robust Framework for Multi-Dialect Arabic Speech Spoofing Detection Benchmark ARFAKE is the first end-to-end benchmark for Arabic speech spoofing detection across multiple dialects. The framework systematically generates synthetic Arabic speech, evaluates its intelligibility and realism, constructs a large-scale spoofing dataset, trains robust detectors, and evaluates generalization across both unseen generators and unseen dialects.… See the full description on the dataset page: https://huggingface.co/datasets/Mohammed01/ArFake.audio10K<n<100K2 likes779 downloads3mo agoHugging Face21egcortes /asr-jargon-specialized-vocabulary A Dataset for Evaluating ASR on Specialized Vocabulary Novel synthetic datasets from the paper "A Dataset for Evaluating ASR on Specialized Vocabulary" (LREC 2026). Code and reproduction scripts: https://github.com/eduardogc8/ASR-Jargon-Dataset-Code Configs Config Language Description synthetic_terms_en English Utterances embedding entirely novel, 100% OOV, LLM-generated technical terms synthetic_terms_pt Portuguese Portuguese equivalent… See the full description on the dataset page: https://huggingface.co/datasets/egcortes/asr-jargon-specialized-vocabulary.audioautomatic-speech-recognition10K<n<100K0 likes569 downloads2mo agoHugging Face22tsinghua-ee /QualiSpeech QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions 📄 Paper: https://arxiv.org/abs/2503.20290 QualiSpeech is a comprehensive English-language speech quality assessment dataset designed to go beyond traditional numerical scores. It introduces detailed natural language comments with reasoning, capturing low-level speech perception aspects such as noise, distortion, continuity, speed, naturalness, listening effort, and overall… See the full description on the dataset page: https://huggingface.co/datasets/tsinghua-ee/QualiSpeech.audioaudio-text-to-text10K<n<100K25 likes529 downloads1y agoHugging Face23elliottash /doppelganger Doppelganger: Sound Effects and Their Synthetic Twins Benchmark for matching a synthetic sound effect to the real recording it was generated from. Paper: https://arxiv.org/abs/2607.04337 · Code: https://github.com/elliottash/doppelganger · models: https://huggingface.co/elliottash/doppelganger Contents sao_pairs/<CatID>/<instance_id>.wav — Stable-Audio-Open audio-conditioned synthetic twins (the main UCS corpus, one twin per verified real clip).… See the full description on the dataset page: https://huggingface.co/datasets/elliottash/doppelganger.audioaudio-classification1K<n<10K0 likes472 downloads2mo agoHugging Face24committa /serena-synthetic-it-28h Qwen3-TTS Italian Synthetic Speech (27h) Synthetic Italian single-speaker speech dataset for TTS training (e.g. Piper), generated with Qwen3-TTS-1.7B-Base in voice-cloning mode. ~29.5k clips, ~27 hours, 22.05 kHz mono WAV, Piper-ready metadata. Dataset summary Property Value Clips (train / val) 26,523 / 2,947 Total duration ~27.3 h (98,099 s) Sample rate 22,050 Hz mono, 16-bit WAV Loudness Normalized to -23 LUFS, silence-trimmed Language Italian… See the full description on the dataset page: https://huggingface.co/datasets/committa/serena-synthetic-it-28h.audiotext-to-speech10K<n<100K1 likes421 downloads1mo agoHugging Face25ggfox00000 /dia-earning21-all Earnings 21 The Earnings 21 dataset ( also referred to as earnings21 ) is a 39-hour corpus of earnings calls containing entity dense speech from nine different financial sectors. This corpus is intended to benchmark automatic speech recognition (ASR) systems in the wild with special attention towards named entity recognition (NER). This work has been recently accepted to Interspeech 2021! File Format Overview In the following section, we provide an overview of the file… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-earning21-all.audion<1K0 likes355 downloads5mo agoHugging Face26anonymous-nsc-author /Neapolitan-Spoken-Corpus Neapolitan Spoken Corpus (NSC) A corpus of read Neapolitan speech for ASR evaluation, with a validated Neapolitan–Italian lexicon, LOSO fine-tuning splits, trained LoRA adapters, metric implementations, per-clip results, and error annotations. This release supersedes the earlier 141-clip single-speaker version of this repository. The earlier release corresponds to Speaker S1 of the present corpus; the old audioData/ and transcripts.csv are replaced by data/audio/ and… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-nsc-author/Neapolitan-Spoken-Corpus.audioautomatic-speech-recognitionn<1K4 likes314 downloads2mo agoHugging Face27itruonghai /EK100 Motivation The actual download link is very slow, including the academic torrent. Therefore, to spare fellow community members from this misery, I am uploading the dataset here. Source You can fnd the original source to download the dataset: https://github.com/epic-kitchens/epic-kitchens-download-scripts Citation @INPROCEEDINGS{Damen2018EPICKITCHENS, title={Scaling Egocentric Vision: The EPIC-KITCHENS Dataset}, author={Damen, Dima and Doughty, Hazel and… See the full description on the dataset page: https://huggingface.co/datasets/itruonghai/EK100.tabularvoice-activity-detection100K<n<1M0 likes289 downloads5mo agoHugging Face28MahiA /CREMA-D CREMA-D This is an audio classification dataset for Emotion Recognition. Classes = 6   ,   Split = Train-Test Structure audios folder contains audio files. train.csv for training split and test.csv for the testing split. Download import os import huggingface_hub audio_datasets_path = "DATASET_PATH/Audio-Datasets" if not os.path.exists(audio_datasets_path): print(f"Given {audio_datasets_path=} does not exist. Specify a valid path ending with 'Audio-Datasets'… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/CREMA-D.audio1K<n<10K0 likes278 downloads2y agoHugging Face29shadow-wxh /VoiceCommandAudioThis is mainly used for fine tune "VoiceCommand" a speech congnition MOD dedicated for SilentHunter game series audioautomatic-speech-recognitionn<1K1 likes275 downloads2y agoHugging Face30facebook /EgoAVU_data [CVPR2026 HIGHLIGHT] EgoAVU, [ICASSP2026 Oral] Exploring Audio Hallucination in Egocentric Video Understanding Official Implementation of EgoAVU: Egocentric Audio-Visual Understanding and Exploring Audio Hallucination in Egocentric Video Understanding See our github for the code and setup instructions. Check out our homepage, paper (CVPR) and paper (ICASSP) for more information. We introduce EgoAVU, a scalable and automated data engine to enable egocentric audio–visual… See the full description on the dataset page: https://huggingface.co/datasets/facebook/EgoAVU_data.tabularquestion-answering1M<n<10M14 likes260 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.