datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.Malaysian-Emilia
Malaysian Emilia
An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Malaysian and Singaporean Speech Generation, where originally from Emilia.
We are improving Malaysian Emilia due to https://github.com/open-mmlab/Amphion/issues/436, check out mesolitica/Malaysian-Emilia-v2
Dataset
Clone and Extract
We upload as split zip files so you can clone and extract distributedly,
huggingface-cli download --repo-type dataset \
--include… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia.Emilia-YODAS-ENEmilia-Dataset-tokenisedEmilia-Dataset-JA-Plus
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2024/08/28: Welcome to join Amphion's Discord channel to stay connected and engage with our community!
2024/08/27: The Emilia dataset is now publicly available! Discover the most extensive and diverse speech generation dataset with… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/Emilia-Dataset-JA-Plus.emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
Emilia-with-Emotion-Annotations
Dataset Card for Emilia with Emotion Annotations
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emilia-with-Emotion-Annotations.Malaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.EmiliaEmilia-ENraw_tts_emilia_ESPnet_espnet_mls-multi_soundstream_16kEmilia-with-Emotion-Annotations4emilia-Dataset
Emilia (Re:Zero) — Illustrious SDXL Character LoRA Training Dataset
中文说明 | English (current)
This is the actual training subset used to train abbuibuibui/emilia-Lora, an unofficial Emilia character LoRA on waiIllustriousSDXL v17. The files here are a copy of the directory named in dataset.toml (image_dir = .../03_captioned/main). Nothing was added from unused candidates, and captions were not rewritten for this release.
This is a fan-made derivative, not an official product.… See the full description on the dataset page: https://huggingface.co/datasets/abbuibuibui/emilia-Dataset.emilia-yodas-en-speaker-embeddings
Emilia-YODAS English Qwen3-TTS Speaker Embeddings
This dataset contains precomputed speaker embeddings for the English subset of
Emilia-YODAS. Each row maps an Emilia-YODAS sample ID to one speaker embedding
extracted from the corresponding audio.
Dataset Details
Source dataset: amphion/Emilia-Dataset
Source subset: Emilia-YODAS English
Embedding model: Qwen/Qwen3-TTS-12Hz-1.7B-Base
Embedding shape: (2048,)
Embedding dtype: float16
Rows: 4,516,833
Split: train
Additional… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-speaker-embeddings.YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.Emilia-with-Emotion-Annotations5Emilia-ZH-BetaMalaysian-Emilia
Malaysian Emilia
Gather Malaysian Emilia from,
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2
https://huggingface.co/datasets/Scicom-intl/Malaysian-Chinese-Emilia
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia#malaysian-dialect
And do,
Trim silent.
Permutation for Voice Conversion include post-filtering during permutation.
Convert to Neucodec speech tokens.
Emilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.Malaysian-Chinese-Emilia
Malaysian-Chinese-Emilia
Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Chinese audio.
Total rows: 605169
Total hours: 1857.611445057867 hours
Permutation for Voice Conversion
Also we already calculated speaker permutation to prepare for voice conversion.
emilia-expressive-zh
Emilia Expressive — Phase-1 Filtered Subset
Auto-generated by emilia_pipeline.scoring.phase1_hf. This is the Phase-1
filtered view: every clip that survived the S0+S1 acoustic funnel, physically
partitioned into quality tiers so you can download exactly the strictness
level you want -- before Phase-2 emotion labeling.
Derived from amphion/Emilia-Dataset
(CC-BY-NC-4.0); the same license and usage restrictions apply.
Pipeline version: voxsift-emilia-v1.3-full (schema 1.3)
Clips:… See the full description on the dataset page: https://huggingface.co/datasets/leeoxiang/emilia-expressive-zh.emilia-en-snac
Stats (EN)
Emilia: 46,349 hours
Emilia-YODAS: 87,258 hours
Total: 133,607 hours
License
The Emilia subset is licensed under CC BY-NC 4.0.
The Emilia-YODAS subset is licensed under CC BY 4.0.
Reference
@inproceedings{emilialarge,
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/emilia-en-snac.Emilia-DE
Emilia - DE
Clean version with only text and audio from the Emilia Dataset.
Samples: 653,109
Language: DE
Usage
from datasets import load_dataset
dataset = load_dataset("Vyvo/Emilia-DE")
sample = dataset['train'][0]
text = sample['text']
audio = sample['audio']['array']
sampling_rate = sample['audio']['sampling_rate']
emilia-yodas-fr-parquetEmilia-dataset-french-splitEmilia-YODAS-Voice-Conversion
Emilia-YODAS-Voice-Conversion
We sample https://huggingface.co/datasets/amphion/Emilia-Dataset YODAS set for voice conversion.
Filter transcriptions based on character repetitiveness and word ngrams.
Filter speaker similarity using https://huggingface.co/nvidia/speakerverification_en_titanet_large during speaker permutation.
Convert audio to speech tokens using https://huggingface.co/neuphonic/neucodec
We also upload the full permutation as zip files.
Speech Tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Emilia-YODAS-Voice-Conversion.emilia_hifitts_fullEmilia-with-Emotion-Annotations3EmiliaEmilia-Annotated-WIPStill a WIP, full dataset is still being annotated
