datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emilia-expressive-zh
Emilia Expressive — Phase-1 Filtered Subset
Auto-generated by emilia_pipeline.scoring.phase1_hf. This is the Phase-1
filtered view: every clip that survived the S0+S1 acoustic funnel, physically
partitioned into quality tiers so you can download exactly the strictness
level you want -- before Phase-2 emotion labeling.
Derived from amphion/Emilia-Dataset
(CC-BY-NC-4.0); the same license and usage restrictions apply.
Pipeline version: voxsift-emilia-v1.3-full (schema 1.3)
Clips:… See the full description on the dataset page: https://huggingface.co/datasets/leeoxiang/emilia-expressive-zh.ExpressiveSpeech
ExpressiveSpeech Dataset
Project Webpage
中文版 (Chinese Version)
About The Dataset
ExpressiveSpeech is a high-quality, expressive, and bilingual (Chinese-English) speech dataset created to address the common lack of consistent vocal expressiveness in existing dialogue datasets.
This dataset is meticulously curated from five renowned open-source emotional dialogue datasets: Expresso, NCSSD, M3ED, MultiDialog, and IEMOCAP. Through a rigorous processing and selection pipeline… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ExpressiveSpeech.ExpressiveSpeech
ExpressiveSpeech
Expressive Speech dataset,
Default, we build our own by combining multiple classifier models and use LLM to generate synthetic description, https://github.com/Scicom-AI-Enterprise-Organization/Multilingual-TTS/issues/2
gigaspeech, from https://speechcraft2024.github.io/speechcraft2024/
libritts_r, from https://speechcraft2024.github.io/speechcraft2024/
Data source for Default
You can follow… See the full description on the dataset page: https://huggingface.co/datasets/Scicom-intl/ExpressiveSpeech.SALMon_Spirit-LM-Expressive-depPrahaTTS-ML-Expressive-DatasetSALMon_Spirit-LM-Expressive
SALMon Normalized Dataset
This repo preserves the SALMon per-config folder layout while normalizing
mismatched schema details across model families.
Expressiveness
Expressiveness — UltraVoice 指令语音合成数据集
基于 UltraVoice 的 instruction_text 字段,
使用 IndexTTS-2.5 零样本音色克隆合成的语音数据集。
55,352 条音频,共 166.71 小时,22050 Hz 单声道 WAV (PCM16)
平均时长 10.84 秒
参考音色随机取自 865 个说话人片段,每个音色被使用 41–93 次
仅合成 instruction(提问)部分,response 未合成
数据构成
大类
条数
时长(小时)
小类
accent
6,000
22.64
AU, CA, GB, IN, SG, ZA
composite
4,143
10.23
en
emotion
21,209
64.81
angry, disgusted, fearful, happy, neutral, sad, surprised
generalqa
6,000
9.89
en
language
6,000… See the full description on the dataset page: https://huggingface.co/datasets/leungtianle/Expressiveness.expressive_speech_24khzseamless-align-expressive
Dataset Card for Seamless-Align-Expressive (WIP). Inspired by https://huggingface.co/datasets/allenai/nllb
Dataset Summary
This dataset was created based on metadata for mined expressive Speech-to-Speech(S2S) released by Meta AI. The S2S contains data for 5 language pairs. The S2S dataset is ~228GB compressed.
How to use the data
There are two ways to access the data:
Via the Hugging Face Python datasets library
Scripts coming soon
Clone the git repo
git… See the full description on the dataset page: https://huggingface.co/datasets/jhu-clsp/seamless-align-expressive.tts-en-zonos2-expressive
ZONOS2 Expressive — English voice cloning
English Mandarin-pipeline counterpart: a diverse reference voice is generated with
Qwen3-TTS VoiceDesign, then cloned with Zyphra/ZONOS2
in expressive mode (accurate_mode=false). Conversational, human-sounding texts
(some emotional, mixed lengths); deliberately not anime/cartoon-style voices.
Content is verified with Qwen/Qwen3-ASR-1.7B:
every clone is transcribed and compared to its target text (asr_wer).
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-en-zonos2-expressive.Expressive_CodecFake
Expressive CodecFake
Expressive CodecFake is a codec-fake expressive speech dataset for audio deepfake detection research. The dataset contains codec-generated expressive speech and nonverbal vocalization samples organized into verified TAR shards.
Current Dataset Structure
Expressive_CodecFake/
├── Verbal speech CF/
│ ├── emodb_2.0_CF/
│ │ ├── emodb_2.0_CF-0000.tar
│ │ └── ...
│ ├── EMOVO_CF/
│ │ ├── EMOVO_CF-0000.tar
│ │ └── ...
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/ggirishg/Expressive_CodecFake.sonora-expressive-registers
Sonora Expressive Registers
Labeled expressive speech clips for training and evaluating directable TTS —
the companion dataset of the Sonora
model line (Project Prosodia, Artificial Humanity).
Openly-licensed expressive speech is scarce: most emotion-labeled corpora are
non-commercial-encumbered. This dataset is CC-BY-4.0 by construction: every clip is
synthesized by Apache/MIT-licensed teacher engines from in-repo authored or public-domain
text, with exact intended emotion… See the full description on the dataset page: https://huggingface.co/datasets/artificial-humanity/sonora-expressive-registers.nuzzle_hand_v1_expressiveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/TECHIdesu/nuzzle_hand_v1_expressive.tts-zh-zonos2-expressive
ZONOS2 — Accurate vs Expressive (Mandarin voice cloning)
Side-by-side A/B comparison of Zyphra/ZONOS2
accurate mode (accurate_mode=true) vs expressive mode (accurate_mode=false).
Same reference voice and same target text per row, cloned twice — one per mode —
so each can be heard back to back. Reference voices are clean Qwen3 generations.
Columns
column
meaning
index
row id
ref_text
text of the reference voice
ref_audio
reference voice (cloning… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-zonos2-expressive.expressive_s2s_zhexpressive_s2s_enukrainian-expressive-single-speaker-tts
Ukrainian Expressive Single-Speaker TTS Dataset
Description
This dataset contains Ukrainian expressive single-speaker speech samples prepared for non-commercial research in speech synthesis, speech processing, and related machine learning tasks.
The dataset was prepared as part of research and development work on Ukrainian text-to-speech and digital avatar generation systems. It is intended to support experiments with expressive Ukrainian speech, TTS model… See the full description on the dataset page: https://huggingface.co/datasets/Roman33111/ukrainian-expressive-single-speaker-tts.expressive-eng-tts
r-labs/expressive-eng-tts
Expressive synthetic Ugandan English speech dataset for conversational Text-to-Speech (TTS) fine-tuning.
Dataset Summary
r-labs/expressive-eng-tts is a fully synthetic expressive Ugandan English TTS dataset designed for fine-tuning conversational speech models with authentic Ugandan English accent, prosody, and expressive speaking behaviors.
The dataset contains speech generated from 3 synthetic speakers:
2 Female speakers
1 Male… See the full description on the dataset page: https://huggingface.co/datasets/r-labs/expressive-eng-tts.speechEvaluation_expressiveYoutube_Expressive_Artistsexpressivenessexpressive-voice-dataset-with-scene
