zonos
Datasets
All datasets matching “zonos”zonos_finetune_datazonos2-vi
ZONOS2 VI — giọng Việt + Colab (L4)
Dataset phụ trợ cho Zyphra/ZONOS2:
6 giọng tham chiếu tiếng Việt (audio + transcript)
2 notebook Colab (clone giọng + đọc SRT), runtime GPU L4
Model weights ZONOS2 không nằm trong repo này — notebook tải từ Zyphra/ZONOS2.
Cấu trúc
voices/<slug>/
ref.wav | ref.mp3
ref_text.txt
colab/
ZONOS2_VI_Clone_Colab.ipynb
ZONOS2_VI_SRT_Colab.ipynb
Giọng
slug
file
ban_mai
Ban_Mai.mp3
lan_trinh… See the full description on the dataset page: https://huggingface.co/datasets/STBack23/zonos2-vi.tts-en-zonos2-expressive
ZONOS2 Expressive — English voice cloning
English Mandarin-pipeline counterpart: a diverse reference voice is generated with
Qwen3-TTS VoiceDesign, then cloned with Zyphra/ZONOS2
in expressive mode (accurate_mode=false). Conversational, human-sounding texts
(some emotional, mixed lengths); deliberately not anime/cartoon-style voices.
Content is verified with Qwen/Qwen3-ASR-1.7B:
every clone is transcribed and compared to its target text (asr_wer).
Columns… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-en-zonos2-expressive.tts-zh-zonos2-expressive
ZONOS2 — Accurate vs Expressive (Mandarin voice cloning)
Side-by-side A/B comparison of Zyphra/ZONOS2
accurate mode (accurate_mode=true) vs expressive mode (accurate_mode=false).
Same reference voice and same target text per row, cloned twice — one per mode —
so each can be heard back to back. Reference voices are clean Qwen3 generations.
Columns
column
meaning
index
row id
ref_text
text of the reference voice
ref_audio
reference voice (cloning… See the full description on the dataset page: https://huggingface.co/datasets/Aynursusuz/tts-zh-zonos2-expressive.tts-zh-zonos2-trialSimpson_Emotion_Labeling_for_ZONOS
