CoolFace
29 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01softcatala /wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons. License identifiers are normalized to cc-zero, cc-by-4.0, cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self. This provides a richer alternative to Common Voice. Characteristics of the dataset: One or multiple speakers Different accents Different domain texts 761 audio files We found this dataset useful for audio tasks such as: Language detection Evaluation of STT systems New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.audioautomatic-speech-recognitionn<1K0 likes2.3k downloads2mo agoHugging Face02Hezep /AudioMarathon 🎵 AudioMarathon: A Comprehensive Benchmark for Long-Context Audio Understanding and Efficient Inference in Multimodal LLMs Abstract AudioMarathon is a large-scale, multi-task audio understanding benchmark designed to systematically evaluate audio language models' capabilities in processing and comprehending long-form audio content. It provides a diverse set of 10 tasks built upon three pillars: long-context audio inputs with durations ranging from 90.0 to 300.0… See the full description on the dataset page: https://huggingface.co/datasets/Hezep/AudioMarathon.audioaudio-classification1K<n<10K4 likes1.8k downloads11mo agoHugging Face03nvidia /Nemotron-Content-Safety-Audio-Dataset Nemotron Content Safety Audio Dataset Dataset Description The Nemotron Content Safety Audio Dataset is a multimodal extension of the Nemotron Content Safety Dataset V2 (Aegis 2.0), comprising 1,928 audio files generated from the test set prompts. This dataset enables multimodal AI safety research by providing spoken versions of adversarial and safety-critical prompts across 23 violation categories. LANGUAGE: All prompts are in English. However, the audio files were… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Content-Safety-Audio-Dataset.audioaudio-classification1K<n<10K5 likes953 downloads10mo agoHugging Face04paodigitalhub /pao-audio-dataset 🎙️ Pa'O Audio Dataset ပအိုဝ်ႏ အငေါဝ်း အဆင်ႏဗာႏ ရွမ်ခြွဉ်းဗူႏ 📌 Project Summary The Pa'O Audio Dataset is an open-source initiative created to facilitate the development of speech technologies and Artificial Intelligence tools for the Pa'O language (ပအိုဝ်ႏဘာႏသာႏငေါဝ်းငွါ). Pa'O is primarily spoken in Shan State and other regions of Myanmar. As a low-resource language in the AI landscape, this dataset provides audio recordings and corresponding… See the full description on the dataset page: https://huggingface.co/datasets/paodigitalhub/pao-audio-dataset.audioautomatic-speech-recognitionn<1K1 likes235 downloads20h agoHugging Face05plnguyen2908 /AudioVisual-Benchmark-Evaluation AudioVisual Benchmark Evaluation — evaluation subsets Item-id lists for the audio-visual benchmark subsets used in our reported evaluation tables. Layout <benchmark>/eval_subset.csv item ids evaluated in the paper <benchmark>/media_index.csv id -> media filename(s) <benchmark>/media/ the media files those ids refer to eval_subset.csv holds a single id column keyed to the source benchmark (question_id, idx, or index). media/ contains exactly the… See the full description on the dataset page: https://huggingface.co/datasets/plnguyen2908/AudioVisual-Benchmark-Evaluation.audiomultiple-choice10K<n<100K0 likes212 downloads24d agoHugging Face06AudioMarathon /AudioMarathon AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings. The release package in this directory is organized around 11 benchmark tasks spanning meeting summarization, automatic speech recognition, reading comprehension, authenticity detection, music genre classification, acoustic scene classification, emotion recognition, spoken named entity reasoning, sound event detection, speaker gender… See the full description on the dataset page: https://huggingface.co/datasets/AudioMarathon/AudioMarathon.audioaudio-classificationn<1K0 likes142 downloads5mo agoHugging Face07igorriti /ambience-audio Ambience audio dataset Overview This dataset was generated by scraping videos from prominent YouTube channels focused on ambient audio. The dataset includes a collection of videos that feature various ambient sounds, such as nature sounds, relaxing music, and environmental noises. For each video, essential metadata was extracted, and a caption was generated using an AI model to enhance the discoverability of the content. This dataset can be useful in various applications… See the full description on the dataset page: https://huggingface.co/datasets/igorriti/ambience-audio.image1K<n<10K6 likes98 downloads2y agoHugging Face08Chengxiang1122 /mcl-mmcl-audiocapsaudio10K<n<100K0 likes98 downloads7mo agoHugging Face09mueller91 /human-perception-audio-deepfake-2026 Human Audio Deepfake Perception 2026 A large-scale listening study evaluating how well humans detect modern audio deepfakes. The dataset contains 35,532 deepfake-detection judgments from 1,768 anonymous participants across 138 TTS and voice-conversion systems, collected via a publicly accessible online listening game in 2025–2026. This is the successor to the 2021 ASVspoof-2019 perception study (Müller, Pizzi & Williams, 2022) and extends the same paradigm to modern systems… See the full description on the dataset page: https://huggingface.co/datasets/mueller91/human-perception-audio-deepfake-2026.textaudio-classification10K<n<100K4 likes73 downloads4mo agoHugging Face10Sodkhuu /Mongolian_audiosaudio1K<n<10K1 likes70 downloads1y agoHugging Face11freococo /9000hours_voa_burmese_audio Overview VOA Burmese radio news archive covering Morning (နံနက် ၅:၃၀ – ၆:၃၀) and Evening (ညပိုင်း ၉:၀၀ – ၁၀:၀၀) programmes for every calendar day from 2012-09-16 → 2025-06-09. Metric Value Hours / rows 9 159 Files per day 2 (morning, evening) Typical file size 15 – 50 MB Licence Public-domain (VOA staff recordings, U.S. 17 U.S.C. § 105) This dataset upgrades Burmese from low-resource to mid-resource status for speech research, enabling self-supervised… See the full description on the dataset page: https://huggingface.co/datasets/freococo/9000hours_voa_burmese_audio.textautomatic-speech-recognition1K<n<10K1 likes50 downloads1y agoHugging Face12FatimahEmadEldin /Arabic-Emotional-Audio-Dataset-Baved BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging) A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits. Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.audioaudio-classification1K<n<10K0 likes49 downloads5mo agoHugging Face13cdactvm /save_audio_punjabi Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cdactvm/save_audio_punjabi.audion<1K0 likes25 downloads2y agoHugging Face14JavisVerse /JavisData-Audioaudio100K<n<1M0 likes24 downloads1y agoHugging Face15Reihaneh /audio_dataset Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Reihaneh/audio_dataset.audion<1K0 likes20 downloads3y agoHugging Face16ilyaslbern7347 /darija-hotel-audioaudion<1K0 likes20 downloads5mo agoHugging Face17Olivia714 /audiocapsaudio1K<n<10K0 likes13 downloads2y agoHugging Face18abdelhaqueidali /Tizuzaf-Audio-Datasetaudioautomatic-speech-recognitionn<1K1 likes12 downloads4mo agoHugging Face19Nifesimih /emi-respiratory-triage-audioaudion<1K0 likes9 downloads2mo agoHugging Face20cdactvm /save_audioaudion<1K0 likes8 downloads2y agoHugging Face21raghad23 /eou_AudioTextaudio100K<n<1M0 likes8 downloads10mo agoHugging Face22vadithiyan /AudioProjaudio10K<n<100K0 likes7 downloads1y agoHugging Face23ThisUsernameAlreadyExistsAlreadyExists /real-or-fake-audiogatedaudioaudio-classification1K<n<10K0 likes7 downloads5mo agoHugging Face24cdactvm /save_audio_punjabi2 Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cdactvm/save_audio_punjabi2.audion<1K0 likes6 downloads2y agoHugging Face25cdactvm /save_audio_tamil_numbers Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/cdactvm/save_audio_tamil_numbers.audion<1K0 likes5 downloads2y agoHugging Face26Aashish17405 /my-audio-dataset Audio Transcription Dataset This dataset contains audio file paths and their corresponding transcriptions for automatic speech recognition (ASR) tasks. Dataset Description This dataset is structured for audio transcription tasks with two main columns: audio: Audio file paths (type: audio) transcript: Text transcriptions (type: text) Files audio_dataset.csv: Main dataset file containing audio paths and transcriptions Dataset Structure audio… See the full description on the dataset page: https://huggingface.co/datasets/Aashish17405/my-audio-dataset.textautomatic-speech-recognitionn<1K0 likes5 downloads1y agoHugging Face27mxchiefx /Sheng_Audioaudio0 likes3 downloads10mo agoHugging Face28ThisUsernameAlreadyExistsAlreadyExists /aitf-dfk3-neutral-audiosgatedaudio1K<n<10K0 likes3 downloads5mo agoHugging Face29ThisUsernameAlreadyExistsAlreadyExists /aitf-dfk3-audios-datasetgatedaudioaudio-classification1K<n<10K0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.