CoolFace
26 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01amphion /Emilia-Datasetgated Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline. News 🔥 2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.audiotext-to-speech10M<n<100M489 likes46k downloads2y agoHugging Face02Vyvo-Research /Emilia-YODAS-ENaudio10M<n<100M4 likes4.3k downloads11mo agoHugging Face03TTS-AGI /emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset. https://huggingface.co/datasets/amphion/Emilia-Dataset audiotext-to-speech10M<n<100M5 likes3.2k downloads2y agoHugging Face04Vyvo-Research /Emilia-ENaudio10M<n<100M2 likes1.6k downloads11mo agoHugging Face05laion /Emilia-with-Emotion-Annotations4audio10M<n<100M1 likes894 downloads1y agoHugging Face06laion /Emilia-with-Emotion-Annotations5audio10M<n<100M3 likes738 downloads1y agoHugging Face07laion /Emilia-with-Emotion-Annotations3audio10M<n<100M1 likes382 downloads1y agoHugging Face08laion /Emilia-with-Emotion-Annotations2audio10M<n<100M1 likes316 downloads1y agoHugging Face09amphion /Emilia-NVgated NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.audiotext-to-speech100K<n<1M52 likes269 downloads1y agoHugging Face10Vyvo-Research /Emilia-YODAS-DEaudio1M<n<10M0 likes150 downloads11mo agoHugging Face11seastar105 /Emilia-YODAS-KO-filteredaudio100K<n<1M0 likes64 downloads2y agoHugging Face12jspaulsen /emilia-yodas-alignedaudio100K<n<1M0 likes64 downloads6mo agoHugging Face13Vyvo-Research /Emilia-ZHaudio10M<n<100M0 likes49 downloads11mo agoHugging Face14lonesamurai /emilia_clean_10k EMILIA Clean 10k A filtered subset of the amphion/Emilia-Dataset (English split), designed for single-speaker TTS training. Dataset Statistics Total clips: 10,000 Speakers: 200 (single-speaker English) Train / Val split: 8,000 / 2,000 Duration per clip: 3–10 seconds Sample rate: 24 kHz (mono) Language: English (EN) Filtering Pipeline Candidate selection — Filtered EMILIA EN clips for duration (3–10s) and DNSMOS quality (≥3.2). Selected top 400 speakers with… See the full description on the dataset page: https://huggingface.co/datasets/lonesamurai/emilia_clean_10k.audio10K<n<100K1 likes40 downloads5mo agoHugging Face15Vyvo-Research /Emilia-JAaudio1M<n<10M0 likes28 downloads11mo agoHugging Face16Vyvo-Research /Emilia-YODAS-KOaudio1M<n<10M0 likes21 downloads11mo agoHugging Face17Vyvo-Research /Emilia-YODAS-FRaudio1M<n<10M0 likes20 downloads11mo agoHugging Face18Vyvo-Research /Emilia-DEaudio100K<n<1M0 likes19 downloads11mo agoHugging Face19Vyvo-Research /Emilia-YODAS-JAaudio100K<n<1M0 likes17 downloads11mo agoHugging Face20Vyvo-Research /Emilia-FRaudio100K<n<1M0 likes15 downloads11mo agoHugging Face21Vyvo-Research /Emilia-KOaudio10K<n<100K0 likes14 downloads11mo agoHugging Face22Vyvo-Research /Emilia-EN-Betaaudio10M<n<100M0 likes12 downloads11mo agoHugging Face23OpenSound /CapSpeech_Emiliagated CapSpeech-Emilia Audio DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details. Overview 🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech_Emilia.audio100K<n<1M3 likes5 downloads1y agoHugging Face24OpenBot /Emilia-ZHaudio1M<n<10M0 likes4 downloads11mo agoHugging Face25Vyvo-Research /Emilia-YODAS-ZHaudio100K<n<1M0 likes4 downloads11mo agoHugging Face26MrDragonFox /DE_Emilia_Yodas_680h_raw_timestampsgatedadditional files for https://huggingface.co/datasets/MrDragonFox/DE_Emilia_Yodas_680h word timestamps with events raw as from elevenlabs scribe v1 used as companion to the main dataset NC licensed text100K<n<1M3 likes3 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.