Aynursusuz/tts-zh-zonos2-expressive
ZONOS2 — Accurate vs Expressive (Mandarin voice cloning) Side-by-side A/B comparison of Zyphra/ZONOS2 accurate mode (accurate_mode=true) vs expressive mode (accurate_mode=false). Same reference voice and same target text per row, cloned twice — one per mode — so each can be heard back to back. Reference voices are clean Qwen3 generations. Columns column meaning index row id ref_text text of the reference voice ref_audio reference voice (cloning… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-zonos2-expressive.
ZONOS2 — Accurate vs Expressive (Mandarin voice cloning)
Side-by-side A/B comparison of Zyphra/ZONOS2 accurate mode (accurate_mode=true) vs expressive mode (accurate_mode=false).
Same reference voice and same target text per row, cloned twice — one per mode — so each can be heard back to back. Reference voices are clean Qwen3 generations.
Columns
DNSMOS measures signal quality only — not expressiveness/prosody, which is best judged by ear.
