CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01avalab /Allo-AVAaudion>1T3 likes8.3k downloads2y agoHugging Face02nccratliri /vad-human-ava-speech Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard H.… See the full description on the dataset page: https://huggingface.co/datasets/nccratliri/vad-human-ava-speech.audio3 likes859 downloads3y agoHugging Face03Reza2kn /nasle-mana-clean-chunked-30s-avasanj Nasl-e-Mana Clean Persian Speech — corrected 30-second chunks Corrected, provenance-preserving audio chunks collected from the Nasl-e-Mana magazine website, generated on 2026-08-30. This release supersedes the earlier unreliable proportional-mapping chunk export; that older release was not used here. Splits Split Rows Audio Columns labeled 4,981 41.41 hours audio, label to_transcribe 11,127 92.72 hours audio The labeled split contains the… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/nasle-mana-clean-chunked-30s-avasanj.audioautomatic-speech-recognition10K<n<100K1 likes672 downloads17d agoHugging Face04avalonai /librivox_small_aud_onlyThis dataset is sampled from LibriVox audio (135+ hours). The purpose of this dataset is for training neural audio codecs or autoencoder to reconstruct high-fidelity audio data. audio100K<n<1M0 likes350 downloads1y agoHugging Face05ZhangHanXD /AvaMERGaudio4 likes220 downloads1y agoHugging Face06lulidong /AVAGDaudio1 likes183 downloads10mo agoHugging Face07ggfox00000 /dia-AvaAvd-test AVA-AVD — test split (audio-visual speaker diarization in the wild) Copie du split test d'AVA-AVD (Xu et al., ACM MM 2022), un corpus de diarisation dans des conditions "in the wild" construit sur AVA Active Speaker. Chaque clip (~5 min) est extrait des vidéos AVA à des offsets précis définis par le repo officiel, puis l'audio est re-synchronisé : les timestamps des RTTMs ont été recalculés en temps clip-local (soustrait min_start du RTTM d'origine) pour être directement utilisables… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/dia-AvaAvd-test.audiovoice-activity-detectionn<1K0 likes148 downloads5mo agoHugging Face08argmaxinc /ava-avdaudion<1K1 likes119 downloads2y agoHugging Face09ShoukanLabs /Kalliope-FalconH1-Avadec-LJSpeechaudio10K<n<100K0 likes113 downloads3mo agoHugging Face10mipal /AVATAR AVATAR: What’s Making That Sound Right Now? Video-centric Audio-Visual Localization AVATAR stands for Audio-Visual localizAtion benchmark for a spatio-TemporAl peRspective in video. AVATAR is a benchmark dataset designed to evaluate video-centric audio-visual localization (AVL) in complex and dynamic real-world scenarios.Unlike previous benchmarks that rely on static image-level annotations and assume simplified conditions, AVATAR offers high-resolution temporal annotations over… See the full description on the dataset page: https://huggingface.co/datasets/mipal/AVATAR.video1B<n<10B1 likes70 downloads11mo agoHugging Face11jobs-git /Allo-AVAaudion<1K0 likes62 downloads2y agoHugging Face12meghtedari /AvationATCaudio10K<n<100K1 likes25 downloads3y agoHugging Face13archivartaunik /liudmila-rubleuskaia-avantury-studyezusa-vyrvicha Авантуры студыёзуса Вырвіча Metadata Author: Людміла Рублеўская Title: Авантуры студыёзуса Вырвіча Narrator: Source Group: Дзіцячыя Source: Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into smaller folders. Target maximum… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/liudmila-rubleuskaia-avantury-studyezusa-vyrvicha.audion<1K0 likes9 downloads4mo agoHugging Face14Tundragoon /avak-englishaudio1K<n<10K0 likes8 downloads8mo agoHugging Face15ahyoun /khmer-voice-avataraudion<1K0 likes8 downloads3mo agoHugging Face160x3 /vad-human-ava-speech Positive Transfer Of The Whisper Speech Transformer To Human And Animal Voice Activity Detection We proposed WhisperSeg, utilizing the Whisper Transformer pre-trained for Automatic Speech Recognition (ASR) for both human and animal Voice Activity Detection (VAD). For more details, please refer to our paper Positive Transfer of the Whisper Speech Transformer to Human and Animal Voice Activity Detection Nianlong Gu, Kanghwi Lee, Maris Basha, Sumit Kumar Ram, Guanghao You, Richard H.… See the full description on the dataset page: https://huggingface.co/datasets/0x3/vad-human-ava-speech.audion<1K0 likes6 downloads4mo agoHugging Face17avatarTester /mydatasetaudio1K<n<10K0 likes5 downloads8mo agoHugging Face18archivartaunik /liudmila-rubleuskaia-avantury-prantsisha-vyrvicha-shkaliara-i-shpega Авантуры Пранціша Вырвіча, шкаляра і шпега Metadata Author: Людміла Рублеўская Title: Авантуры Пранціша Вырвіча, шкаляра і шпега Narrator: Source Group: Дзіцячыя Source: Notes The original audio files are preserved as-is: no conversion; no re-encoding; no filename changes inside each split folder, except removing one common top-level archive folder when present. To avoid Hugging Face Dataset Viewer scan-size errors, the dataset is split into… See the full description on the dataset page: https://huggingface.co/datasets/archivartaunik/liudmila-rubleuskaia-avantury-prantsisha-vyrvicha-shkaliara-i-shpega.audion<1K0 likes3 downloads4mo agoHugging Face19avaeziaiteam /chunked-cherrypicked-booksgatedaudio10K<n<100K0 likes3 downloads2mo agoHugging Face20SaifPunjwani /Allo-AVA Allo-AVA: A Large-Scale Multimodal Dataset for Allocentric Avatar Animation Overview Allo-AVA (Allocentric Audio-Visual Avatar) is a large-scale multimodal dataset designed for research and development in avatar animation. It focuses on generating natural and contextually appropriate gestures from text and audio inputs in an allocentric (third-person) perspective. The dataset addresses the scarcity of high-quality, synchronized multimodal data capturing the intricate… See the full description on the dataset page: https://huggingface.co/datasets/SaifPunjwani/Allo-AVA.audio0 likes2 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.