dacvae
Datasets
All datasets matching “dacvae”commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.dacvae-tts-tr-w512-clean
DACVAE-TTS Turkish run C (width 512, clean data)
Generated audio of every evaluated checkpoint of the training run tr-w512-clean (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
Run C: configs/nano_tr_w512.yaml (66.5M parameters: width 512, 8 heads, batch expansion 4, frame budget 6000), trained from scratch on… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-w512-clean.dacvae-tts-tr-nano-b-ke4
DACVAE-TTS Turkish run B (batch expansion 4)
Generated audio of every evaluated checkpoint of the training run tr-nano-b-ke4 (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
Run B: configs/nano_tr_ke4.yaml (same 51.4M model as run A, context-sharing batch expansion 4, frame budget 7000), same data (~77 h), one RTX… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-nano-b-ke4.dacvae-tts-tr-nano-a
DACVAE-TTS Turkish run A (nano recipe, full data)
Generated audio of every evaluated checkpoint of the training run tr-nano-a (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
Run A: configs/nano_tr.yaml (51.4M parameters), 17 shards of Vyvo/tr-dataset-12 with quality >= 55 (~77 h train), one RTX 4090, frame budget… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-nano-a.mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.dacvae-tts-tr-stage2-b
DACVAE-TTS Turkish stage 2 (from run B 40k, clean data)
Generated audio of every evaluated checkpoint of the training run tr-stage2-b (Turkish zero-shot voice-cloning TTS,
dacvae-tts, frozen Meta DACVAE latents, 48 kHz). This repository holds
model outputs and metrics, not training data. Training data: Vyvo/tr-dataset-12 (Turkish podcast segments).
Stage 2: warm start (--init-from) from run B's 40k checkpoint (VoiceHub/dacvae-tts-tr-nano-b-ke4), trained 20k more updates on the… See the full description on the dataset page: https://huggingface.co/datasets/VoiceHub/dacvae-tts-tr-stage2-b.
