datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Emilia-Dataset.Reazon-Speech
Reazon Speech v2 dataset mirror
Original Dataset Source
Hugging Face Dataset Page: reazon-research/reazonspeech
Project Page: Reazon Research
License
This dataset is a mirror of the original Reazon Speech v2 dataset, but on 🤗 server (so may be faster). This dataset is licensed under the CDLA-Sharing-1.0. The original dataset comes with the following restriction:
TO USE THIS DATASET, YOU MUST AGREE THAT YOU WILL USE THE DATASET SOLELY FOR THE PURPOSE OF… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Reazon-Speech.
