datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour Emilia-YODAS… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-Dataset.Emilia-Dataset-JA-Plus
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2024/08/28: Welcome to join Amphion's Discord channel to stay connected and engage with our community!
2024/08/27: The Emilia dataset is now publicly available! Discover the most extensive and diverse speech generation dataset with… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/Emilia-Dataset-JA-Plus.emilia-yodasA mirror of the Emilia-YODAS dataset. Only includes the YODAS subset from the original dataset.
https://huggingface.co/datasets/amphion/Emilia-Dataset
YouTube-Cantonese-Emilia
YouTube Cantonese — Emilia
2,064,679 speaker-homogeneous Cantonese speech segments — 5,312.6 hours — produced by
running alvanlii/cantonese-youtube
through the Emilia
speech-data pipeline (source separation → diarization → VAD segmentation → ASR → MOS filtering).
Each row is one clean, single-speaker segment of 3–30 s with a transcript, a speaker turn
label and a DNSMOS quality score. Audio is shipped separately as MP3s inside zip parts, in
both an original and a silence-trimmed… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/YouTube-Cantonese-Emilia.Emilia-dataset-french-with-gender
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/AdrienB134/Emilia-dataset-french-with-gender.emilia-en-snac
Stats (EN)
Emilia: 46,349 hours
Emilia-YODAS: 87,258 hours
Total: 133,607 hours
License
The Emilia subset is licensed under CC BY-NC 4.0.
The Emilia-YODAS subset is licensed under CC BY 4.0.
Reference
@inproceedings{emilialarge,
author={He, Haorui and Shang, Zengqiang and Wang, Chaoren and Li, Xuyuan and Gu, Yicheng and Hua, Hua and Liu, Liwei and Yang, Chen and Li, Jiaqi and Shi, Peiyang and Wang, Yuancheng and Chen, Kai and Zhang, Pengyuan and Wu… See the full description on the dataset page: https://huggingface.co/datasets/nytopop/emilia-en-snac.Emilia-NV
NVSpeech Dataset
Overview
The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at enhancing the capabilities of automatic speech recognition (ASR) and text-to-speech (TTS) systems. The dataset features explicit word-level annotations for 18 categories of paralinguistic vocalizations, including non-verbal sounds like laughter and breathing, as well as lexicalized interjections like "uhm" and "oh."… See the full description on the dataset page: https://huggingface.co/datasets/amphion/Emilia-NV.emilia-yodas-en-aligned
Emilia-YODAS EN Word-Aligned
Word-level forced-alignment timestamps for the English subset of
amphion/Emilia-Dataset
(Emilia-YODAS split), produced with
Qwen/Qwen3-ForcedAligner-0.6B.
No audio is redistributed — this dataset contains only metadata (IDs,
transcripts already present in Emilia-YODAS, and per-word [start, end]
timestamps). To use it, join on id with the original Emilia-YODAS audio.
Stats
Metric
Value
Utterances
4,516,833
Total audio
11,572.7… See the full description on the dataset page: https://huggingface.co/datasets/duplexio/emilia-yodas-en-aligned.emilia-subset
JA_Emilia_Yodas_ScribeEvents
JA Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/JA_Emilia_Yodas_266h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (4433 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/JA_Emilia_Yodas_ScribeEvents.EN_Emilia_Yodas_ScribeEvents
EN Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/EN_Emilia_Yodas_616h containing only samples with ElevenLabs Scribe v1 audio events (vocal bursts, background sounds, etc.).
Changes from source
Filtered to only include rows where events_scribe is non-empty (16017 rows out of 228,265 original)
Bracket format unified: Round brackets (laughs) in text_scribe replaced with square brackets [laughs] for consistency with vocal burst annotation format… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/EN_Emilia_Yodas_ScribeEvents.DE_Emilia_Yodas_ScribeEvents
DE Emilia Yodas - Scribe Events
Filtered subset of MrDragonFox/DE_Emilia_Yodas_680h containing only samples with ElevenLabs Scribe v1 audio events.
Changes from source
Filtered to rows where events_scribe is non-empty (12173 rows kept)
Bracket format unified: (event) in text_scribe replaced with [event]
Event types include
Vocal bursts: <laughs>, <sighs>, <clears throat>, etc.
Background: <background noise>, <music>, etc.
Other: <pause>, <unintelligible>… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/DE_Emilia_Yodas_ScribeEvents.emilia-token
Emilia EN Pocket Mimi continuous latents
This gated repository contains the English Emilia training data representation
used by the LatentTTS experiments in this project. Audio was encoded offline
with the continuous Gaussian Mimi speech VAE used by Pocket TTS. The files are
intended to let an authorized researcher reproduce latent-domain training
without encoding the source audio again.
Access and licensing
This is a derived representation of the
official Emilia… See the full description on the dataset page: https://huggingface.co/datasets/Gong1212/emilia-token.MakeSense-Emilia-DatasetThis is a dataset of simultaneous interpretation policy / trajectory, includeing EN, ZH, JA, KO translation.
Each language has 2,000 records of asr data and 6,000 records of simultaneous interpretation trajectory data.
Source: amphion/Emilia-Dataset
Please refer to github to see the usage
Emilia-Dataset
Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
This is the official repository 👑 for the Emilia dataset and the source code for the Emilia-Pipe speech data preprocessing pipeline.
News 🔥
2025/02/26: The Emilia-Large dataset, featuring over 200,000 hours of data, is now available!!! Emilia-Large combines the original 101k-hour Emilia dataset (licensed under CC BY-NC 4.0) with the brand-new 114k-hour… See the full description on the dataset page: https://huggingface.co/datasets/samson-ailabs/Emilia-Dataset.emilia-ja-plus-metadata
Emilia Dataset JA Plus — normalized metadata
This is an audit-backed, metadata-only derivative of
ayousanz/Emilia-Dataset-JA-Plus. It does not bundle audio payloads.
Verified snapshot statistics
Metric
Value
Metadata rows
78,748
Unique IDs
78,748
Duplicate IDs
0
Unique speakers
10,046
Duration represented by metadata
145.06 hours
Language labels
ja: 78,748
Unique transcripts
64,034
Duplicate transcript rows
14,714
Japanese-labelled rows… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/emilia-ja-plus-metadata.
