datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
commonvoice22-sidon-dacvae
CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents
Source
sarulab-speech/commonvoice22_sidon
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.mls-enhanced-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.maestrino-data-DACVAEbalanced-audio-snippets-40x3k-DACVAEenhanced-audiosnippets-DACVAEvocal-bursts-taxonomy-DACVAE
Vocal Bursts Taxonomy — DACVAE + MaestroClap Embeddings & Scores
Processed version of with DACVAE latents, MaestroClap embeddings, derived attribute/quality/speaker scores, and Gemini-verified labels.
Overview
Metric
Value
Total samples
16,175
Categories
82
Genders
male, female
Female samples
8,097
Male samples
8,078
Gemini Label Verification
Every sample was sent to Gemini 3.1 Flash Lite for two independent tasks:
Match scoring:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/vocal-bursts-taxonomy-DACVAE.emolia-3k-speaker-clusters-DACVAE
Emolia 3K Speaker Clusters
A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster.
Overview
The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure:
Outlier… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-3k-speaker-clusters-DACVAE.Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.emotion-conditioning-test-dacvae
Emotion Conditioning Test - DAC-VAE
This dataset contains emotion-ranked subsets of speech samples encoded as DAC-VAE latents.
Contents
For each of the 40 emotion categories, the top 1000 samples (ranked by annotation score)
are packaged into a separate tar file. Each tar file contains triplets of files per sample:
{sample_key}.npy — audio latent (DAC-VAE encoded)
{sample_key}.ref.npy — speaker reference latent
{sample_key}.json — metadata including text, caption… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emotion-conditioning-test-dacvae.podcast-dramabox-dacvae-pairs
podcast-dramabox-dacvae-pairs
Paired audio-codec latents for training a DramaBox → DACVAE latent translator.
Both codecs share an identical grid: 25 Hz, 128-dim, frame-aligned (same length).
Derived from TTS-AGI/podcast-tokenized-bg3.5-enj5.
How it was built (per sample)
DACVAE latent (from source dataset, = target) → DACVAE.decode → 48 kHz mono wav
→ duplicate to stereo → DramaBox/LTX-2.3 audio VAE encode → patchify → DramaBox latent (= input).
Both latents… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/podcast-dramabox-dacvae-pairs.enhanced-emo-snippets-balanced-DACVAE
Enhanced Emotion Snippets — Balanced DACVAE
A balanced, emotion-bucketed subset of TTS-AGI/enhanced-audiosnippets-DACVAE,
organized by Empathic Insight Voice+ emotion and voice attribute categories.
Overview
This dataset provides up to 100 samples per magnitude bucket for each of the
40 emotion categories and 15 voice attribute dimensions scored by
Empathic Insight Voice+.
Selection Criteria
Emotion Categories (40 dimensions)
For each emotion (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/enhanced-emo-snippets-balanced-DACVAE.balanced-audio-score-datasets-DACVAEurdu-dacvae---
language:
- ur
- en
license: cc-by-4.0
task_categories:
- automatic-speech-recognition
tags:
- urdu
- speech
- audio
- asr
- tts
size_categories:
- 10K<n<100K
---
# urdu-dacvae
A cleaned, normalised Urdu (+ limited English) speech dataset derived from
multiple sources, intended for ASR / TTS / VAE latent modelling.
## Dataset Summary
| Field | Value |
|---|---|
| **Samples** | 182,472 |
| **Total audio** | 445.1 hours |
| **Duration filter** | 2.0s < duration < 20.0s… See the full description on the dataset page: https://huggingface.co/datasets/zuhri025/urdu-dacvae.mls-dacvae
Multilingual LibriSpeech converted to DAC VAE latents
Source
facebook/multilingual_librispeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-dacvae.DACVAE-latentsemotion-attribute-conditioning-dacvae
Echo TTS - Emotion & Attribute Conditioning Dataset (DAC-VAE Latents)
Pre-bucketed speech dataset with DAC-VAE latent representations organized by 40 emotion categories and 13 vocal/audio attributes. Built for conditioning fine-tuning of Echo TTS and similar DiT-based TTS models.
Overview
Total emotion samples: 163,271 (across 40 emotions, 10K cap per emotion)
Total attribute samples: ~785K (across 13 attributes x 7 buckets, 10K cap per bucket)
Format: WebDataset .tar… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emotion-attribute-conditioning-dacvae.Emotion-Voice-Attribute-Reference-Snippets-DACVAE
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: ~98 GB (latents-only, no WAV)
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.eurospeech-enhanced-dacvae
EuroSpeech parliamentary speech converted to DAC VAE latents
Source
disco-eth/EuroSpeech
Format
Each tar shard (~2GB) contains samples with three files per sample:
{sample_key}.audio.flac # Original audio (FLAC, original sample rate)
{sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32
{sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second
DAC VAE Latent Format
Model:… See the full description on the dataset page: https://huggingface.co/datasets/laion/eurospeech-enhanced-dacvae.emilia-yodas-dacvae_tokenized
