datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.Emilia-YODAS-ENemilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
Malaysian-Emilia-annotated
Malaysian Emilia Annotated
Annotate Malaysian-Emilia using Data-Speech pipeline.
Malaysian Youtube
Originally from malaysia-ai/crawl-youtube
Total 3168.8 hours.
Gender prediction, filtered-24k_processed_24k_gender.zip
Language prediction, filtered-24k_processed_language.zip
Force alignment.
Post cleaned to 24k and 44k sampling rates,
24k, filtered-24k_processed_24k.zip
44k, filtered-24k_processed_44k.zip
Synthetic description… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-annotated.Emilia-ENEmilia-with-Emotion-Annotations4emilia-Dataset
Emilia (Re:Zero) — Illustrious SDXL Character LoRA Training Dataset
中文说明 | English (current)
This is the actual training subset used to train abbuibuibui/emilia-Lora, an unofficial Emilia character LoRA on waiIllustriousSDXL v17. The files here are a copy of the directory named in dataset.toml (image_dir = .../03_captioned/main). Nothing was added from unused candidates, and captions were not rewritten for this release.
This is a fan-made derivative, not an official product.… See the full description on the dataset page: https://huggingface.co/datasets/abbuibuibui/emilia-Dataset.emilia-yodas-en-speaker-embeddings
Emilia-YODAS English Qwen3-TTS Speaker Embeddings
This dataset contains precomputed speaker embeddings for the English subset of
Emilia-YODAS. Each row maps an Emilia-YODAS sample ID to one speaker embedding
extracted from the corresponding audio.
Dataset Details
Source dataset: amphion/Emilia-Dataset
Source subset: Emilia-YODAS English
Embedding model: Qwen/Qwen3-TTS-12Hz-1.7B-Base
Embedding shape: (2048,)
Embedding dtype: float16
Rows: 4,516,833
Split: train
Additional… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-speaker-embeddings.Emilia-with-Emotion-Annotations5YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.Emilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.Malaysian-Emilia
Malaysian Emilia
Gather Malaysian Emilia from,
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2
https://huggingface.co/datasets/Scicom-intl/Malaysian-Chinese-Emilia
https://huggingface.co/datasets/mesolitica/Malaysian-Emilia#malaysian-dialect
And do,
Trim silent.
Permutation for Voice Conversion include post-filtering during permutation.
Convert to Neucodec speech tokens.
Malaysian-Chinese-Emilia
Malaysian-Chinese-Emilia
Use https://github.com/mesolitica/Emilia to pseudo-label Malaysian Chinese audio.
Total rows: 605169
Total hours: 1857.611445057867 hours
Permutation for Voice Conversion
Also we already calculated speaker permutation to prepare for voice conversion.
emilia-en-snac
Stats (EN)
Emilia: 46,349 hours
Emilia-YODAS: 87,258 hours
Total: 133,607 hours
License
The Emilia subset is licensed under CC BY-NC 4.0.
The Emilia-YODAS subset is licensed under CC BY 4.0.
Reference
@inproceedings{emilialarge,
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/emilia-en-snac.Emilia-DE
Emilia - DE
Clean version with only text and audio from the Emilia Dataset.
Samples: 653,109
Language: DE
Usage
from datasets import load_dataset
dataset = load_dataset("Vyvo/Emilia-DE")
sample = dataset['train'][0]
text = sample['text']
audio = sample['audio']['array']
sampling_rate = sample['audio']['sampling_rate']
emilia-yodas-fr-parquetEmilia-dataset-french-splitEmilia-YODAS-Voice-Conversion
Emilia-YODAS-Voice-Conversion
We sample https://huggingface.co/datasets/amphion/Emilia-Dataset YODAS set for voice conversion.
Filter transcriptions based on character repetitiveness and word ngrams.
Filter speaker similarity using https://huggingface.co/nvidia/speakerverification_en_titanet_large during speaker permutation.
Convert audio to speech tokens using https://huggingface.co/neuphonic/neucodec
We also upload the full permutation as zip files.
Speech Tokenizer… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/Emilia-YODAS-Voice-Conversion.emilia_hifitts_fullEmilia-with-Emotion-Annotations3Emilia-Annotated-WIPStill a WIP, full dataset is still being annotated
emilia-yodas-en-mimiEmilia-with-Emotion-Annotations2Malaysian-Emilia-v2
Malaysian Emilia v2
This version 2 should fixed https://github.com/open-mmlab/Amphion/issues/436, an Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Malaysian and Singaporean Speech Generation. Replicating Emilia on,
Dataset
Clone and Extract
We upload as split zip files so you can clone and extract distributedly,
huggingface-cli download --repo-type dataset \
--include '*.zip' \
--local-dir './' \
--max-workers 20 \… See the full description on the dataset page: https://huggingface.co/datasets/mesolitica/Malaysian-Emilia-v2.Emilia-YODAS-DE
Emilia-YODAS - DE
Clean version with only text and audio from the Emilia Dataset.
Samples: 2,005,364
Language: DE
Usage
from datasets import load_dataset
dataset = load_dataset("Vyvo/Emilia-YODAS-DE")
sample = dataset['train'][0]
text = sample['text']
audio = sample['audio']['array']
sampling_rate = sample['audio']['sampling_rate']
emilia_mfa_correctExtra-Emilia
Extra Emilia
Extra dataset to extend Tamil and Mandarin capability for Malaysian-Emilia.
Tamil
Total length is 891 hours.
Mandarin
Total length is 301 hours.
Emilia-R-POSTPROCESS-44d9db35Emilia-NV
NVSpeech Dataset
Overview
The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.emilia-captions-v3
