CoolFace
Datasetpublic

VoiceHub/dacvae-tts-tr-trc-w640

dacvae-tts Turkish 115M: trc-w640 (tr-combined + tr-dataset-12) Checkpoints of the training run trc-w640, each with 5 outputs: Turkish zero-shot voice-cloning TTS (dacvae-tts, branch turkish-tts; flow-matching DiT on frozen Meta DACVAE latents, 48 kHz). Training in progress (update 61200 of 100000): every 10000 updates a new checkpoint and its 5 outputs are added here. 5 outputs per checkpoint The same 5 unseen sentences, each spoken in the voice of a different… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-trc-w640.

sourceHugging Facecc-by-nc-4.0updated 1d agoView on Hugging Face
1likes297downloads
Dataset Card

dacvae-tts Turkish 115M: trc-w640 (tr-combined + tr-dataset-12)

Checkpoints of the training run trc-w640, each with 5 outputs: Turkish zero-shot voice-cloning TTS (dacvae-tts, branch turkish-tts; flow-matching DiT on frozen Meta DACVAE latents, 48 kHz). Training in progress (update 61200 of 100000): every 10000 updates a new checkpoint and its 5 outputs are added here.

5 outputs per checkpoint

The same 5 unseen sentences, each spoken in the voice of a different held-out speaker (samples/prompts/, real recordings never seen in training, decoded through the codec), with the demo settings: guidance 5, 32 Euler steps, duration auto, one sample (no best-of-N). The dataset viewer above plays them (pick the checkpoint in the menu). Raw model output: at guidance 5 it is loud (about -12 LUFS); the demo normalizes to -16 LUFS.

checkpoint**Freya WER****Freya CER**Freya SIMFreya DNSMOS5 outputs: WERCERSIMDNSMOSlisten
step-0010000----17.2%8.3%0.9573.11samples/step-0010000
step-00200006.93%4.05%0.9572.875.1%4.0%0.9613.21samples/step-0020000
step-0030000----8.1%2.7%0.9573.07samples/step-0030000
step-00400004.70%2.48%0.9552.8811.1%5.4%0.9523.31samples/step-0040000
step-0050000running7.1%3.7%0.9493.17samples/step-0050000
step-0060000----8.1%4.1%0.9503.22samples/step-0060000

Freya = Freya-TR-Eval, all 495 sentences (3,911 words, 24 held-out voices), the published protocol (guidance 5, 32 steps, prompt-rate duration, one sample, faster-whisper large-v3 beam 5). It runs on the CPU next to the training (about 2 hours per checkpoint), so a checkpoint gets it when the evaluator is free, newest first; the audio is not uploaded, per-sentence transcripts are in freya/<step>/results.jsonl. WER is measured only on Freya-TR-Eval. The published model VoiceHub/dacvae-tts-tr-w512 (66.5M) scores WER 4.32% / CER 2.50%, SIM 0.946, DNSMOS 2.89 on it; FreyaTTS-183M 8.0% / 3.0%, XTTS-v2 11.1%. The 5 outputs are for listening (their 99 words make one wrong word worth about 1 WER point).

step-0060000

#textoutputprompt voiceWhisper transcript of the outputWERSIMDNSMOS
01Bu sabah erkenden kalkıp sahilde uzun bir yürüyüş yaptım; deniz o kadar sakindi ki martıların sesi bile net duyuluyordu.01.wavprompt-00.wavBu sabah erkenden kalkıp sahilde uzun bir yürüyüş yaptım. Deniz Odar sakindi ki martıların sesi bile net duyuluyordu.0.1050.9853.07
02Yapay zeka modelleri, büyük miktarda veriden örüntüler öğrenerek konuşma sentezi gibi karmaşık görevleri artık şaşırtıcı bir doğallıkla yerine getirebiliyor.02.wavprompt-01.wavYapay zeka modelleri büyük miktarda veriden örüntüler öğrenerek konuşma sentezi gibi karmaşık görevleri artık şaşırtıcı bir doğallıkla yerine getirebiliyor.0.0000.9393.46
03Anneannem her bayram sabahı mutfakta baklava açar, evin içi tereyağı ve fıstık kokusuyla dolar, biz çocuklar da sofranın kurulmasını sabırsızlıkla beklerdik.03.wavprompt-02.wavAnneannem her bayram sabahı mutfakta baklava açar, evin içi tereyağı ve fıstık kokusuyla dolar. Biz çocuklar da sofranın kurulmasını sabırsızlıkla beklerdik.0.0000.8673.16
04Toplantı yarın saat 14'te başlayacak; lütfen sunum dosyalarınızı öğleden önce paylaşın ve %20'lik bütçe artışı önerisini gündeme eklemeyi unutmayın.04.wavprompt-03.wavToplantı yarın saat 14'te başlayacak. Lütfen sunum dosyalarınızı öğleden önce paylaşın ve %20'lik bütçe artışı önerisini gündeme eklemeyi unutmayın. Altyazı M.K.0.1430.9843.17
05İstanbul'da akşam trafiği başlamadan Boğaz Köprüsü'nden geçmek istiyorsanız en geç dörtte yola çıkmalısınız, yoksa bir saatlik yol üçe katlanır.05.wavprompt-04.wavİstanbul'da akşam trafiği başlamadan, Boğaz Köprüsü'nün başlamadan geçmek istiyorsanız en geç dörtte yola çıkmalısınız. Yoksa yoksa bir saatlik yol üçe katlanır.0.1580.9753.25

Model and data

Model: 640x14 generator (10 heads, low-rank AdaLN, LARoPE cross-attention, CTC at block 9), 114.8M parameters; byte-level text with the turkish-v2 normalization (the serving frontend: numbers, dates, currencies, abbreviations and acronyms spelled out). Prompts: 70% another recording of the same speaker, otherwise cut from the clip itself. Muon, WSD learning rate (0.0007, cooldown over the last 20%), 2x RTX 4090.

Data: Codyfederer/tr-combined (221,531 clips, 277 h) filtered by Whisper-turbo transcript agreement (CER <= 0.15 after removing fillers such as "eee"), DNSMOS OVRL >= 2.5, speaker-embedding consistency within each speaker label and speaking rate: 155,812 clips, 210.7 h, 1,904 speakers; plus the clean subset of Vyvo/tr-dataset-12 (32,226 clips, 72 h). Together 188,038 clips, 283 h, 2,561 speakers; validation/test speakers (tr-combined: whole videos) are never trained on.

Checkpoints

  • —checkpoints/step-0010000.pt
  • —checkpoints/step-0020000.pt
  • —checkpoints/step-0030000.pt
  • —checkpoints/step-0040000.pt
  • —checkpoints/step-0050000.pt
  • —checkpoints/step-0060000.pt

Each file holds the EMA weights used for synthesis (plus the raw weights). dacvae_tts/ is the package version that reads them:

python
import sys
from huggingface_hub import hf_hub_download, snapshot_download
sys.path.insert(0, snapshot_download("VoiceHub/dacvae-tts-tr-trc-w640", repo_type="dataset", allow_patterns=["dacvae_tts/*"]))
from dacvae_tts.inference import Synthesizer

tts = Synthesizer(hf_hub_download("VoiceHub/dacvae-tts-tr-trc-w640", "checkpoints/step-0060000.pt", repo_type="dataset"), device="cuda")
tts.synthesize("Merhaba, bugün nasılsın?", "prompt.wav", reference_text="Prompt kaydının metni.",
               output="out.wav", guidance=5.0, steps=32, duration_mode="auto")

train.jsonl holds the training and validation curves, config.json the full configuration.