expressive-tts
tts-en-zonos2-expressive
ZONOS2 Expressive — English voice cloning
English Mandarin-pipeline counterpart: a diverse reference voice is generated with
Qwen3-TTS VoiceDesign, then cloned with Zyphra/ZONOS2
in expressive mode (accurate_mode=false). Conversational, human-sounding texts
(some emotional, mixed lengths); deliberately not anime/cartoon-style voices.
Content is verified with Qwen/Qwen3-ASR-1.7B:
every clone is transcribed and compared to its target text (asr_wer).
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-en-zonos2-expressive.tts-zh-zonos2-expressive
ZONOS2 — Accurate vs Expressive (Mandarin voice cloning)
Side-by-side A/B comparison of Zyphra/ZONOS2
accurate mode (accurate_mode=true) vs expressive mode (accurate_mode=false).
Same reference voice and same target text per row, cloned twice — one per mode —
so each can be heard back to back. Reference voices are clean Qwen3 generations.
Columns
column
meaning
index
row id
ref_text
text of the reference voice
ref_audio
reference voice (cloning… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-zonos2-expressive.ukrainian-expressive-single-speaker-tts
Ukrainian Expressive Single-Speaker TTS Dataset
Description
This dataset contains Ukrainian expressive single-speaker speech samples prepared for non-commercial research in speech synthesis, speech processing, and related machine learning tasks.
The dataset was prepared as part of research and development work on Ukrainian text-to-speech and digital avatar generation systems. It is intended to support experiments with expressive Ukrainian speech, TTS model… See the full description on the dataset page: https://huggingface.co/datasets/Roman33111/ukrainian-expressive-single-speaker-tts.expressive-eng-tts
r-labs/expressive-eng-tts
Expressive synthetic Ugandan English speech dataset for conversational Text-to-Speech (TTS) fine-tuning.
Dataset Summary
r-labs/expressive-eng-tts is a fully synthetic expressive Ugandan English TTS dataset designed for fine-tuning conversational speech models with authentic Ugandan English accent, prosody, and expressive speaking behaviors.
The dataset contains speech generated from 3 synthetic speakers:
2 Female speakers
1 Male… See the full description on the dataset page: https://huggingface.co/datasets/r-labs/expressive-eng-tts.
