datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.Emilia-YODAS-ENemilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
Emilia-ENEmilia-with-Emotion-Annotations4Emilia-with-Emotion-Annotations5Emilia-with-Emotion-Annotations3Emilia-with-Emotion-Annotations2Emilia-NV
NVSpeech Dataset
Overview
The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.Emilia-YODAS-DEEmilia-YODAS-KO-filteredemilia-yodas-alignedEmilia-ZHemilia_clean_10k
EMILIA Clean 10k
A filtered subset of the amphion/Emilia-Dataset (English split), designed for single-speaker TTS training.
Dataset Statistics
Total clips: 10,000
Speakers: 200 (single-speaker English)
Train / Val split: 8,000 / 2,000
Duration per clip: 3–10 seconds
Sample rate: 24 kHz (mono)
Language: English (EN)
Filtering Pipeline
Candidate selection — Filtered EMILIA EN clips for duration (3–10s) and DNSMOS quality (≥3.2). Selected top 400 speakers with… See the full description on the dataset page: https://huggingface.co/datasets/lonesamurai/emilia_clean_10k.Emilia-JAEmilia-YODAS-KOEmilia-YODAS-FREmilia-DEEmilia-YODAS-JAEmilia-FREmilia-KOEmilia-EN-BetaCapSpeech_Emilia
CapSpeech-Emilia Audio
DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech
Please refer to 🤗CapSpeech for the whole dataset and 🚀CapSpeech repo for more details.
Overview
🔥 CapSpeech is a new benchmark designed for style-captioned TTS (CapTTS) tasks, including style-captioned text-to-speech synthesis with sound effects (CapTTS-SE), accent-captioned TTS (AccCapTTS), emotion-captioned TTS (EmoCapTTS) and… See the full description on the dataset page: https://huggingface.co/datasets/OpenSound/CapSpeech_Emilia.Emilia-ZHEmilia-YODAS-ZHDE_Emilia_Yodas_680h_raw_timestampsadditional files for https://huggingface.co/datasets/MrDragonFox/DE_Emilia_Yodas_680h
word timestamps with events raw as from elevenlabs scribe v1
used as companion to the main dataset
NC licensed
