CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes46k downloads2y agoHugging Face02Vyvo-Research /Emilia-YODAS-ENaudio10M<n<100M4 likes3.8k downloads11mo agoHugging Face03lighthouse-emnlp2024 /Clotho-Moment Clotho-Moment This repository provides wav files used in Language-based Audio Moment Retrieval. Each sample includes long audio containing some audio events with the temporal and textual annotation. Project page: https://h-munakata.github.io/Language-based-Audio-Moment-Retrieval/ Code: https://github.com/line/lighthouse Split Train train/train-{000..715}.tar 37930 audio samples Valid valid/valid-{000..108}.tar 5741 audio samples Test test/test-{000..142}.tar 7569… See the full description on the dataset page: https://huggingface.co/datasets/lighthouse-emnlp2024/Clotho-Moment.audioaudio-text-to-text10K<n<100K2 likes3.6k downloads8mo agoHugging Face04TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.1k downloads2y agoHugging Face05laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.2k downloads10mo agoHugging Face06Vyvo-Research /Emilia-ENaudio10M<n<100M2 likes1.6k downloads11mo agoHugging Face07krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face08krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face09krishnakalyan3 /emo_webdsaudio10K<n<100K5 likes1.3k downloads2y agoHugging Face10laion /laions_got_talent_enhanced_no_metadataaudio10K<n<100K0 likes1.1k downloads2y agoHugging Face11laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes894 downloads1y agoHugging Face12echodict /NeMo NVIDIA NeMo Speech Checkout our HuggingFace🤗 collection for the latest open weight checkpoints and demos! Updates 2026-03: Nemotron 3 VoiceChatis now released in Early Access. Built on the Nemotron Nano v2 LLM backbone with Nemotron speech and TTS decoder, VoiceChat delivers full-duplex, natural, interruptible conversations with low latency. Try out the demo and apply for early access. 2026-03: Nemotron-Speech-Streaming v2603 has been updated. It has been… See the full description on the dataset page: https://huggingface.co/datasets/echodict/NeMo.audio10K<n<100K0 likes875 downloads5mo agoHugging Face13laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes743 downloads1y agoHugging Face14akuzdeuov /qwen3-tts-multilingual-emotional-speechaudio1M<n<10M0 likes677 downloads12d agoHugging Face15NandemoGHS /Japanese-Eroge-Voice Japanese-Eroge-Voice Description This dataset contains pairs of audio data and corresponding transcriptions extracted from Japanese eroge (adult games) that I have personally purchased. The transcriptions are generated using the litagin/anime-whisper model. Preprocessing Steps The raw audio data has undergone the following preprocessing steps: Loudness Normalization: Audio loudness is normalized using ffmpeg's 2-pass loudnorm filter to target parameters of… See the full description on the dataset page: https://huggingface.co/datasets/NandemoGHS/Japanese-Eroge-Voice.audiotext-to-speech100K<n<1M37 likes596 downloads1y agoHugging Face16VoiceNet /emolia emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired) This is the emolia-balanced-5M-subset corpus repackaged for high-quality audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz (PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json samples. The JSON sidecar carries the full annotation stack: Original metadata (id, text, duration, speaker, language, dnsmos). A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.audioaudio-classification1M<n<10M1 likes385 downloads5mo agoHugging Face17laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes383 downloads1y agoHugging Face18ESpeech /ESpeech-webinars2 Webinar Audio Dataset Dataset Description This dataset contains 850 hours processed webinar audio segments with corresponding metadata. Each audio file represents a segment extracted from webinar recordings, processed at 44.1kHz sample rate. Dataset Summary Language: Russian Task: TTS, ASR, Quality Asessment Audio format: MP3, 44.1kHz sample rate Structure: Segmented audio files with JSON metadata Dataset Structure Data Fields… See the full description on the dataset page: https://huggingface.co/datasets/ESpeech/ESpeech-webinars2.audiotext-to-speech100K<n<1M8 likes378 downloads1y agoHugging Face19laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes317 downloads1y agoHugging Face20EaseZh /musanMUSAN Identifier: SLR17 Summary: A corpus of music, speech, and noise Category: Audio License: Attribution 4.0 International (CC BY 4.0) Downloads (use a mirror closer to you): musan.tar.gz [11G] ( The corpus ) Mirrors: [EU] [EU] [CN] About this resource: MUSAN is a corpus of music, speech, and noise recordings. This work was supported by the National Science Foundation Graduate Research Fellowship under Grant No. 1232825 and by Spoken Communications. You can cite the data using the… See the full description on the dataset page: https://huggingface.co/datasets/EaseZh/musan.audio1K<n<10K0 likes297 downloads7mo agoHugging Face21amphion /Emilia-NVgated NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.audiotext-to-speech100K<n<1M52 likes280 downloads1y agoHugging Face22TTS-AGI /mls-enhanced-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.audioautomatic-speech-recognition100K<n<1M0 likes223 downloads6mo agoHugging Face23krishnakalyan3 /emo_speech_filtered_v12 second filtered emotional speech in webdataset format https://huggingface.co/datasets/EQ4You/Emotional_Speech audio10K<n<100K0 likes181 downloads2y agoHugging Face24TTS-AGI /emolia-hq Emolia-HQ Emolia-HQ is a high-quality, speaker-paired subset of the LAION Emolia dataset. Each sample includes a target utterance and a reference utterance from the same speaker, enabling speaker-conditioned tasks such as voice conversion, expressive TTS, and speaker-aware emotion recognition. Source Derived from laion/Emolia by: Quality filtering: Only samples with dnsmos >= 3.0 are retained. Speaker pairing: Each target sample is matched with a reference audio from the… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-hq.audioaudio-classification10M<n<100M3 likes156 downloads7mo agoHugging Face25JusperLee /EchoSetaudio10K<n<100K10 likes155 downloads2y agoHugging Face26Vyvo-Research /Emilia-YODAS-DEaudio1M<n<10M0 likes151 downloads11mo agoHugging Face27AffectDF /AffectDF_EmotionSDD AffectDF: Emotionally Expressive Speech Deepfake Benchmark Overview AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks. AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.audioaudio-classification100K<n<1M0 likes146 downloads4mo agoHugging Face28zszhong /Lyra-Evalaudio10K<n<100K1 likes110 downloads2y agoHugging Face29novateur /cosyvoice2_enaudio100K<n<1M0 likes105 downloads2y agoHugging Face30humanify /common_voice_englishaudio1M<n<10M1 likes105 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.