CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /LAION-Audio-300Maudio100M<n<1B74 likes18k downloads2y agoHugging Face02laion /soundscapesaudio10M<n<100M7 likes15k downloads1y agoHugging Face03laion /laions_got_talent LAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a generated dataset comprising voice acting samples that exhibit a wide range of emotions, vocal bursts, topics, and content. This dataset is a component of the BUD-E project, spearheaded by LAION with support from Intel. Dataset Composition The dataset includes: Emotional Diversity: Samples portraying various emotions to facilitate research in emotional recognition and… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent.audio100K<n<1M41 likes9.7k downloads2y agoHugging Face04laion /laion-audio-previewaudio1M<n<10M11 likes4.5k downloads2y agoHugging Face05lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.6k downloads8mo agoHugging Face06laion /laions_got_talent_rawaudio10K<n<100K7 likes3.3k downloads2y agoHugging Face07zh-liu799 /0813audio100K<n<1M0 likes2.6k downloads1y agoHugging Face08laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.1k downloads10mo agoHugging Face09laion /captioned-ai-music-snippets Dataset Overview A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models. Source Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository. Captioning All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions. License Apache 2.0 audio1M<n<10M15 likes1.9k downloads11mo agoHugging Face10laion /synthetic_vocal_burstsThis repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. https://huggingface.co/datasets/sleeping-ai/Vocal-burst We captioned them using Gemini Flash Audio 2.0. This dataset contains, this dataset contains ~ 365,000 vocal bursts from all kinds of categories. It might be helpful for pre-training audio text foundation models to generate and understand all kinds of nuances in vocal bursts. audio100K<n<1M6 likes1.6k downloads2y agoHugging Face11laion /majestrino-dataaudio1M<n<10M1 likes1.5k downloads7mo agoHugging Face12mitermix /audiosnippets_long_2_5Maudio1M<n<10M3 likes1.4k downloads2y agoHugging Face13laion /laions_got_talent_enhanced_no_metadataaudio10K<n<100K0 likes978 downloads2y agoHugging Face14litagin /Galgame_Speech_SER_16kHz Dataset Card for Galgame_Speech_SER_16kHz [!IMPORTANT]The following rules (in the original repository) must be followed: 必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关! 训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。 English: You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_SER_16kHz.audioautomatic-speech-recognition1M<n<10M17 likes972 downloads2y agoHugging Face15laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes894 downloads1y agoHugging Face16laion /laions_got_talent_german_bicodecaudio100K<n<1M0 likes868 downloads2y agoHugging Face17mitermix /audiosnippets_long_1Maudio100K<n<1M0 likes823 downloads2y agoHugging Face18litagin /reazon-speech-v2-clonegated Reazon Speech v2 dataset mirror Original Dataset Source Hugging Face Dataset Page: reazon-research/reazonspeech Project Page: Reazon Research License This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction: TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/litagin/reazon-speech-v2-clone.audioautomatic-speech-recognition10K<n<100K12 likes747 downloads2y agoHugging Face19laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes738 downloads1y agoHugging Face20ClareNie /Light-Omni-Training Light-Omni Training Dataset This repository contains the training data used by Light-Omni, a multimodal agent framework for reflexive video understanding with long-term memory. Light-Omni uses memory-augmented multimodal streams to train adapters for memory construction, response generation, and reaction/action control. Links Project page: https://clare-nie.github.io/Light-Omni/ Code: https://github.com/Clare-Nie/Light-Omni Dataset:… See the full description on the dataset page: https://huggingface.co/datasets/ClareNie/Light-Omni-Training.audiovisual-question-answering100K<n<1M3 likes704 downloads3mo agoHugging Face21laion /unsupervised_peoples_speech_raw_voice_activity_detection_snippets_part_1audio100M<n<1B4 likes671 downloads1y agoHugging Face22leungtianle /AgentChataudio100K<n<1M2 likes646 downloads5mo agoHugging Face23zh-liu799 /89823872audio100K<n<1M0 likes620 downloads1y agoHugging Face24laion /voiceclap-data VoiceCLAP Data The audio + dense-caption mixture used to train laion/voiceclap-small and laion/voiceclap-large. Each tar shard is a WebDataset of paired <key>.flac (48 kHz mono audio) + <key>.json (caption + metadata) samples. Captions and structured attribute annotations are produced automatically by a pipeline of audio-aware LLMs — Qwen-Audio, Gemini Flash 2.5, and a thinking-mode reasoning model that scores emotion under the EmoNet taxonomy plus per-clip vocal-burst, timbre… See the full description on the dataset page: https://huggingface.co/datasets/laion/voiceclap-data.audioaudio-classification1M<n<10M0 likes562 downloads5mo agoHugging Face25laion /common-voice-subset-for-clapaudion<1K1 likes521 downloads9mo agoHugging Face26litagin /Galgame_Speech_ASR_16kHz Dataset Card for Galgame_Speech_ASR_16kHz [!IMPORTANT]The following rules (in the original repository) must be followed: 必须遵守GNU General Public License v3.0内的所有协议!附加:禁止商用,本数据集以及使用本数据集训练出来的任何模型都不得用于任何商业行为,如要用于商业用途,请找数据列表内的所有厂商授权(笑),因违反开源协议而出现的任何问题都与本人无关! 训练出来的模型必须开源,是否在README内引用本数据集由训练者自主决定,不做强制要求。 English: You must comply with all the terms of the GNU General Public License v3.0!Additional note: Commercial use is prohibited. This dataset and any model trained using this dataset… See the full description on the dataset page: https://huggingface.co/datasets/litagin/Galgame_Speech_ASR_16kHz.audioautomatic-speech-recognition1M<n<10M48 likes488 downloads2y agoHugging Face27labhamlet /NatHEARPlease see LICENSE.txt for each dataset licenses audio10K<n<100K0 likes488 downloads8mo agoHugging Face28FYQ12138 /log_prompt_dataaudio100K<n<1M0 likes486 downloads3mo agoHugging Face29CASIA-LM /OpenS2S_Datasets How to Use? Download, merge the files, and extract You can run the following command to merge the compressed file parts after downloading. cat en_response_wav.tar.gz.* > en_response_wav.tar.gz cat zh_response_wav.tar.gz.* > zh_response_wav.tar.gz audio100K<n<1M8 likes471 downloads1y agoHugging Face30laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes382 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.