CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face02krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face03anyspeech /ipapack_plus_train_3audio1M<n<10M0 likes1.1k downloads1y agoHugging Face04project-riz /osu-beatmaps osu! Beatmaps Dataset (WebDataset) A collection of ranked/loved osu! beatmaps with audio and chart data, in WebDataset format. Dataset Variants Variant Audio Format Description original MP3/OGG/WAV Full quality original audio files compressed 64kbps Mono Opus Compressed audio for smaller download from datasets import load_dataset # Load original audio variant ds = load_dataset("project-riz/osu-beatmaps", "original", streaming=True) # Load compressed… See the full description on the dataset page: https://huggingface.co/datasets/project-riz/osu-beatmaps.audioaudio-classification10K<n<100K4 likes1k downloads8mo agoHugging Face05noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes743 downloads6mo agoHugging Face06laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes670 downloads1y agoHugging Face07FYQ12138 /log_prompt_dataaudio100K<n<1M0 likes520 downloads3mo agoHugging Face08hkchengrex /MMAudio-precomputed-results Precomputed results for MMAudio Results from four model variants of MMAudio. All results are in the .flac format with lossless compression. A cache folder contains the feature caches computed by the evaluation script. Code: https://github.com/hkchengrex/MMAudio Evaluation: https://github.com/hkchengrex/av-benchmark VGGSound Contains the VGGSound test set results. There are 15220 videos, collected with our best effort. Not all videos in the test sets are available… See the full description on the dataset page: https://huggingface.co/datasets/hkchengrex/MMAudio-precomputed-results.audio10K<n<100K0 likes454 downloads2y agoHugging Face09projecte-aina /parlament_parla_v3 Dataset Card for ParlamentParla v3 - Speech Corpus of Catalan Parliamentary Sessions A speech corpus composed of Catalan Parliamentary Sessions.The v3 and last version of the corpus includes both clean and other quality segments, divided into short segments (less than 30 seconds) and long segments (more than 30 seconds). The total dataset encompasses 1059h 48m 04s of speech, including 945h 51m 06s for the short segments and 113h 56m 58s for the long segments, with a total of… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/parlament_parla_v3.audioautomatic-speech-recognition100K<n<1M1 likes439 downloads2y agoHugging Face10laion /laions_got_talent_previewaudio1K<n<10K1 likes380 downloads2y agoHugging Face11ufal /parczech4speech-segmented ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based on parliamentary recordings and official transcripts. This sentence-segmented variant is designed for speech recognition and synthesis tasks, offering clean audio-text alignment and reliable segment boundaries. It is derived from the ParCzech 4.0 corpus and AudioPSP 24.01 audio collection. Using WhisperX and Wav2Vec 2.0… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-segmented.audioautomatic-speech-recognition100K<n<1M1 likes336 downloads1y agoHugging Face12laion /talent_plus_rl_groups_of_50_with_audiobox_scoresaudio1M<n<10M0 likes293 downloads10mo agoHugging Face13laion /timbre-audio-caption-pairsaudio100K<n<1M2 likes212 downloads9mo agoHugging Face14laion /freesound-commercially-permissive-subset-with-captionsaudio100K<n<1M2 likes204 downloads11mo agoHugging Face15JavisVerse /MM-PreTrain JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation [HomePage] [Paper] [GitHub] TL;DR We introduce JavisGPT, a multimodal LLM that can understand audiovisual inputs and simultaneously generate synchronized sounding videos in a unified model. We also curate the JavisInst-Omni dataset to facilitate instruction-tuning for comprehension and generation on sounding videos. 📰 News [2025.12.30] 🚀 We release the training… See the full description on the dataset page: https://huggingface.co/datasets/JavisVerse/MM-PreTrain.audio100K<n<1M0 likes197 downloads9mo agoHugging Face16freococo /115hours_pvtv_myanmar_asr 115 Hours PVTV Myanmar ASR This dataset contains 156,262 audio-transcript pairs of spoken Burmese, totaling approximately 115.31 hours. The audio segments were extracted from publicly available YouTube videos published by PVTV and aligned using subtitle timestamps. Dedication This dataset would not exist without the persistent voices of PVTV editors, journalists, narrators, and production teams, who continue to speak to the people under difficult conditions. PVTV is the… See the full description on the dataset page: https://huggingface.co/datasets/freococo/115hours_pvtv_myanmar_asr.audioautomatic-speech-recognition100K<n<1M2 likes164 downloads1y agoHugging Face17anyspeech /ipapack_plus_5audio1M<n<10M0 likes98 downloads2y agoHugging Face18krishnakalyan3 /laion-audio-preview-splitaudio1M<n<10M2 likes88 downloads2y agoHugging Face19PinnHe /CompA-R-BackupOriginally from https://huggingface.co/papers/2406.11768, we downloaded it from Google Drive and converted it to HuggingFace, as the train_audio portion was observed to have disappeared. audio10K<n<100K0 likes88 downloads10mo agoHugging Face20Vano04 /MELD-Preprocessed MELD Preprocessed for SER This dataset is the manually preprocessed audio only version of MELD, only audio IDs, utterance transcriptions, dialogue IDs and Utterance IDs were extracted. S. Poria, D. Hazarika, N. Majumder, G. Naik, R. Mihalcea, E. Cambria. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation. (2018) Chen, S.Y., Hsu, C.C., Kuo, C.C. and Ku, L.W. EmotionLines: An Emotion Corpus of Multi-Party Conversations. arXiv preprint arXiv:1802.08379… See the full description on the dataset page: https://huggingface.co/datasets/Vano04/MELD-Preprocessed.audio10K<n<100K0 likes80 downloads10mo agoHugging Face21ufal /parczech4speech-unsegmented ParCzech4Speech (Unsegmented Variant) Dataset Summary ParCzech4Speech (Unsegmented Variant) is a large-scale Czech speech dataset derived from parliamentary recordings and official transcripts. This variant captures continuous speech segments without enforcing sentence boundaries, making it well-suited for real-world streaming ASR scenarios and speech modeling tasks that benefit from natural discourse flow. The dataset is created using a combination of WhisperX and… See the full description on the dataset page: https://huggingface.co/datasets/ufal/parczech4speech-unsegmented.audioautomatic-speech-recognition1M<n<10M1 likes78 downloads1y agoHugging Face22Reord-AI /kenya-philippines-twospeaker-english-dialoguegated Kenya/Philippines English Dialogue Two-speaker dialogues in English, recorded on split tracks. Changelog Jan 2026: v1 release - vad-segmented WebRTC tracks Specs Speakers: >150; ~15 PH, remaining KE Total duration: ~65 hours Files sample rate: 48kHz Actual sample rate: TBD Language: English (PH, KE accents) Topics: day-to-day conversation Collection method The dataset is built to capture the variety in the Kenyan accent. The Philippino interviewers… See the full description on the dataset page: https://huggingface.co/datasets/Reord-AI/kenya-philippines-twospeaker-english-dialogue.audio10K<n<100K2 likes58 downloads8mo agoHugging Face23Darknsu /voxceleb2-40k-part1-preprocess-all-files-separateaudio100K<n<1M0 likes58 downloads5mo agoHugging Face24anyspeech /ipapack_plus_3audio100K<n<1M0 likes57 downloads2y agoHugging Face25anyspeech /ipapack_plus_6audio1M<n<10M0 likes51 downloads2y agoHugging Face26EarthSpeciesProject /animalspeak-pseudovox AnimalSpeak Pseudovox Train-Unseen This dataset contains the train-unseen split of AnimalSpeak Pseudovox. Each example is a short, silence-trimmed, single-vocalization WAV clip plus compact per-clip metadata. It does not include generated conversations, captions, QA pairs, or MCQ answers. Rows: 346,907 Shards: 18 Maximum rows per shard: 20,000 Files data-20k/train-*.tar: WebDataset-style shards containing audio/<audio_name> WAV entries. metadata.parquet: one row per… See the full description on the dataset page: https://huggingface.co/datasets/EarthSpeciesProject/animalspeak-pseudovox.audioaudio-classification100K<n<1M0 likes47 downloads5mo agoHugging Face27anyspeech /ipapack_plus_1audio100K<n<1M0 likes42 downloads2y agoHugging Face28PuristanLabs1 /urdu-turn-detection-audio-v2 🗣️ Urdu Turn Detection (Audio Dataset V2) This is the official dataset for the model [PuristanLabs1/urdu-turn-v2](https://huggingface.co/PuristanLabs1/urdu-turn-v2), a high precision, low latency system for detecting the end of a conversational turn in Urdu speech. It contains 11,479 audio clips (balanced between Complete and Incomplete) specifically designed to train robust models for realtime Voice AI applications like "Smart Turn" or "Barge-in" detection. 🚀 How… See the full description on the dataset page: https://huggingface.co/datasets/PuristanLabs1/urdu-turn-detection-audio-v2.audioaudio-classification10K<n<100K0 likes42 downloads9mo agoHugging Face29umoubuton /paimonaudio10K<n<100K2 likes41 downloads3y agoHugging Face30ndandanov /polyglot-modelsThis dataset repository shall serve as a mirror hosting models for polyglot. Availability Please note that currently the only languages, which have all models, are English (en) and Bulgarian (bg). Other languages may have partial support. Adding missing models In case you have previously downloaded polyglot language models, which are not available in this repo, please open a Pull Request. License All rights belong to the original authors. Please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/ndandanov/polyglot-models.audion<1K1 likes38 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.