datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ExpressiveSpeech
ExpressiveSpeech Dataset
Project Webpage
中文版 (Chinese Version)
About The Dataset
ExpressiveSpeech is a high-quality, expressive, and bilingual (Chinese-English) speech dataset created to address the common lack of consistent vocal expressiveness in existing dialogue datasets.
This dataset is meticulously curated from five renowned open-source emotional dialogue datasets: Expresso, NCSSD, M3ED, MultiDialog, and IEMOCAP. Through a rigorous processing and selection pipeline… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ExpressiveSpeech.SALMon_Spirit-LM-Expressive
SALMon Normalized Dataset
This repo preserves the SALMon per-config folder layout while normalizing
mismatched schema details across model families.
tts-en-zonos2-expressive
ZONOS2 Expressive — English voice cloning
English Mandarin-pipeline counterpart: a diverse reference voice is generated with
Qwen3-TTS VoiceDesign, then cloned with Zyphra/ZONOS2
in expressive mode (accurate_mode=false). Conversational, human-sounding texts
(some emotional, mixed lengths); deliberately not anime/cartoon-style voices.
Content is verified with Qwen/Qwen3-ASR-1.7B:
every clone is transcribed and compared to its target text (asr_wer).
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-en-zonos2-expressive.Expressive_CodecFake
Expressive CodecFake
Expressive CodecFake is a codec-fake expressive speech dataset for audio deepfake detection research. The dataset contains codec-generated expressive speech and nonverbal vocalization samples organized into verified TAR shards.
Current Dataset Structure
Expressive_CodecFake/
├── Verbal speech CF/
│ ├── emodb_2.0_CF/
│ │ ├── emodb_2.0_CF-0000.tar
│ │ └── ...
│ ├── EMOVO_CF/
│ │ ├── EMOVO_CF-0000.tar
│ │ └── ...
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/ggirishg/Expressive_CodecFake.tts-zh-zonos2-expressive
ZONOS2 — Accurate vs Expressive (Mandarin voice cloning)
Side-by-side A/B comparison of Zyphra/ZONOS2
accurate mode (accurate_mode=true) vs expressive mode (accurate_mode=false).
Same reference voice and same target text per row, cloned twice — one per mode —
so each can be heard back to back. Reference voices are clean Qwen3 generations.
Columns
column
meaning
index
row id
ref_text
text of the reference voice
ref_audio
reference voice (cloning… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-zonos2-expressive.expressive-eng-tts
r-labs/expressive-eng-tts
Expressive synthetic Ugandan English speech dataset for conversational Text-to-Speech (TTS) fine-tuning.
Dataset Summary
r-labs/expressive-eng-tts is a fully synthetic expressive Ugandan English TTS dataset designed for fine-tuning conversational speech models with authentic Ugandan English accent, prosody, and expressive speaking behaviors.
The dataset contains speech generated from 3 synthetic speakers:
2 Female speakers
1 Male… See the full description on the dataset page: https://huggingface.co/datasets/r-labs/expressive-eng-tts.expressiveness
