CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face02laion /soundscapesaudio10M<n<100M7 likes16k downloads1y agoHugging Face03laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.7k downloads2y agoHugging Face04laion /BVD-A-10M-URLs LAION-BVD — 10M Audio Clip URLs This repository contains the metadata and captions for ~10 million audio clips randomly sampled from BVD-V-55M for large-scale audio pre-training. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated laion/BVD-A-10M repository. Dataset structure One row per audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-10M-URLs.tabulartext-to-audio10M<n<100M1 likes9.1k downloads28d agoHugging Face05laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face06laion /laions_got_talent_rawaudio10K<n<100K7 likes3.4k downloads2y agoHugging Face07laion /BVD-A-1.7M-URLs LAION-BVD — 1.7M Audio Clip URLs This repository contains the metadata and captions for ~1.7 million audio clips taken from BVD-V-55M and sampled for uniqueness of the source video, so that the subset maximises source diversity rather than clip count. The audio itself is not included in this repository — every clip is described by the URL of its source video plus the start_time/end_time offsets needed to reproduce it. The corresponding clip files are available in the gated… See the full description on the dataset page: https://huggingface.co/datasets/laion/BVD-A-1.7M-URLs.tabulartext-to-audio1M<n<10M1 likes2.6k downloads28d agoHugging Face08benjamin-paine /freesound-laion-640k About this Repository This repository is a re-upload of the FreeSound.org dataset as curated by LAION for the larger LAION-Audio-630k dataset, with the following changes: Limited columns to only the audio and basic metadata. Incorporated necessary information for licensing and attribution. Removed ambiguously licensed samples, amounting to around 1,000 total samples. What about download links? Links were ommitted for the sake of size, as they can be constructed from… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k.audioaudio-to-audio100K<n<1M13 likes2.5k downloads2y agoHugging Face09laion /laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants. Updated Composition Voices and Languages English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.audio1M<n<10M6 likes2.5k downloads1y agoHugging Face10laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.2k downloads10mo agoHugging Face11laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes2.1k downloads11mo agoHugging Face12benjamin-paine /freesound-laion-640k-commercial-16khz-full About this Repository This repository is the training split of the complete FreeSound LAION 640k dataset, limited only to licenses that permit commercial works, resampled to 16khz using torchaudio.transforms.Resample. This is ideal for use cases where a variety of audio is desired but fidelity and labels are unnecessary, such as background audio for augmenting other datasets. Dataset Versions You are looking at the full dataset which contains 403,146 unique sounds… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k-commercial-16khz-full.audioaudio-to-audio100K<n<1M2 likes1.9k downloads2y agoHugging Face13laion /majestrino-dataaudio1M<n<10M1 likes1.7k downloads6mo agoHugging Face14laion /synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. https://huggingface.co/datasets/sleeping-ai/Vocal-burst We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories. It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts. audio100K<n<1M6 likes1.6k downloads2y agoHugging Face15laion /laions_got_talent_enhanced_no_metadataaudio10K<n<100K0 likes1.2k downloads2y agoHugging Face16laion /emolia-thinking-balanced-buckets Emolia-Thinking — Balanced Per-Dimension Bucket Subset A balanced, per-dimension bucket subset of VoiceNet/emolia-thinking, derived from that dataset's zero-shot VoiceNet-dimension labels. For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a roughly equal number of clips from each ordinal bucket (0–6), so that downstream training / probing sees a balanced distribution along each axis instead of the strongly skewed natural distribution. How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.audioaudio-classification100K<n<1M0 likes1k downloads2mo agoHugging Face17laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes945 downloads1y agoHugging Face18laion /laions_got_talent_german_bicodecaudio100K<n<1M0 likes923 downloads2y agoHugging Face19laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes743 downloads1y agoHugging Face20laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes671 downloads1y agoHugging Face21laion /common-voice-subset-for-clapaudion<1K1 likes524 downloads9mo agoHugging Face22laion /voiceclap-data VoiceCLAP Data The audio + dense-caption mixture used to train laion/voiceclap-small and laion/voiceclap-large. Each tar shard is a WebDataset of paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata) samples. Captions and structured attribute annotations are produced automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5, and a thinking-mode reasoning model that scores emotion under the EmoNet taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.audioaudio-classification1M<n<10M0 likes521 downloads5mo agoHugging Face23BidirLM /laion_audio_contrastive ⚠️ Part of the BidirLM-Omni Collection > This dataset is a specific modality sub-sample of the corpus used to train the BidirLM-Omni models. Looking for the full training mixture? > If you want to access the complete, balanced 1.8M sample omnimodal dataset (integrating text, image, audio), please visit the global integration hub here:👉 BidirLM/BidirLM-Omni-Contrastive 📜 Citation If you use this processed dataset or the broader BidirLM mixture in your research, please cite… See the full description on the dataset page: https://huggingface.co/datasets/BidirLM/laion_audio_contrastive.text100K<n<1M0 likes479 downloads4mo agoHugging Face24laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes382 downloads1y agoHugging Face25laion /voice-acting-data-annotated Voice Acting Data - Annotated Post-processed version of laion/voice-acting-data. Processing Pipeline RE-USE Speech Enhancement (nvidia/RE-USE) - Applied to non-singing samples for noise reduction LavaSR Super Resolution (YatharthS/LavaSR) - Audio bandwidth extension to 48kHz Whisper Turbo ASR - Full transcript with word-level timestamps Scene Split - Audio split at CUT TO: transition into two parts (Part 1 + Part 2) VoiceCLAP Large Embeddings… See the full description on the dataset page: https://huggingface.co/datasets/laion/voice-acting-data-annotated.audioaudio-classification0 likes382 downloads13d agoHugging Face26laion /laions_got_talent_previewaudio1K<n<10K1 likes380 downloads2y agoHugging Face27KBlueLeaf /laion-coco-13m-taraudio10M<n<100M1 likes365 downloads1y agoHugging Face28laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes318 downloads1y agoHugging Face29laion /Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated audio1M<n<10M3 likes293 downloads1y agoHugging Face30laion /talent_plus_rl_groups_of_50_with_audiobox_scoresaudio1M<n<10M0 likes293 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.