CoolFace
14 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01TTS-AGI /commonvoice22-sidon-dacvae CommonVoice 22 (Sidon-enhanced) converted to DAC VAE latents Source sarulab-speech/commonvoice22_sidon Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/commonvoice22-sidon-dacvae.audioautomatic-speech-recognition1M<n<10M1 likes1.3k downloads6mo agoHugging Face02TTS-AGI /mls-enhanced-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-enhanced-dacvae.audioautomatic-speech-recognition100K<n<1M0 likes225 downloads6mo agoHugging Face03TTS-AGI /maestrino-data-DACVAEtext1M<n<10M0 likes207 downloads6mo agoHugging Face04TTS-AGI /balanced-audio-snippets-40x3k-DACVAEtext100K<n<1M0 likes179 downloads6mo agoHugging Face05TTS-AGI /enhanced-audiosnippets-DACVAEtext1M<n<10M1 likes116 downloads6mo agoHugging Face06TTS-AGI /emolia-3k-speaker-clusters-DACVAE Emolia 3K Speaker Clusters A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster. Overview The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure: Outlier… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-3k-speaker-clusters-DACVAE.textaudio-classification10K<n<100K0 likes90 downloads6mo agoHugging Face07TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: 1.54 TB Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.audiotext-to-speech100K<n<1M0 likes87 downloads6mo agoHugging Face08TTS-AGI /enhanced-emo-snippets-balanced-DACVAE Enhanced Emotion Snippets — Balanced DACVAE A balanced, emotion-bucketed subset of TTS-AGI/enhanced-audiosnippets-DACVAE, organized by Empathic Insight Voice+ emotion and voice attribute categories. Overview This dataset provides up to 100 samples per magnitude bucket for each of the 40 emotion categories and 15 voice attribute dimensions scored by Empathic Insight Voice+. Selection Criteria Emotion Categories (40 dimensions) For each emotion (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/enhanced-emo-snippets-balanced-DACVAE.textaudio-classification10K<n<100K0 likes40 downloads6mo agoHugging Face09TTS-AGI /mls-dacvae Multilingual LibriSpeech converted to DAC VAE latents Source facebook/multilingual_librispeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/mls-dacvae.audioautomatic-speech-recognition10K<n<100K0 likes28 downloads6mo agoHugging Face10TTS-AGI /balanced-audio-score-datasets-DACVAEtext100K<n<1M0 likes25 downloads6mo agoHugging Face11TTS-AGI /DACVAE-latentstext100K<n<1M0 likes20 downloads7mo agoHugging Face12TTS-AGI /emotion-attribute-conditioning-dacvae Echo TTS - Emotion & Attribute Conditioning Dataset (DAC-VAE Latents) Pre-bucketed speech dataset with DAC-VAE latent representations organized by 40 emotion categories and 13 vocal/audio attributes. Built for conditioning fine-tuning of Echo TTS and similar DiT-based TTS models. Overview Total emotion samples: 163,271 (across 40 emotions, 10K cap per emotion) Total attribute samples: ~785K (across 13 attributes x 7 buckets, 10K cap per bucket) Format: WebDataset .tar… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emotion-attribute-conditioning-dacvae.text1M<n<10M0 likes15 downloads7mo agoHugging Face13TTS-AGI /Emotion-Voice-Attribute-Reference-Snippets-DACVAE Emotion and Voice Attribute Reference Snippets - DACVAE and Wave Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio. Overview Total samples: 606,178 Filtered out: 363,331 (samples with speech_quality < 1.8) Total tar files: 328 Total size: ~98 GB (latents-only, no WAV) Audio format: WAV, 48kHz, PCM 16-bit mono Latents: DAC-VAE float16 [T, 128] at 25 frames/sec Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.texttext-to-speech100K<n<1M0 likes9 downloads6mo agoHugging Face14laion /eurospeech-enhanced-dacvae EuroSpeech parliamentary speech converted to DAC VAE latents Source disco-eth/EuroSpeech Format Each tar shard (~2GB) contains samples with three files per sample: {sample_key}.audio.flac # Original audio (FLAC, original sample rate) {sample_key}.dacvae.npy # DAC VAE latent [T_latent, 128] numpy float32 {sample_key}.metadata.json # All metadata + duration_seconds + chars_per_second DAC VAE Latent Format Model:… See the full description on the dataset page: https://huggingface.co/datasets/laion/eurospeech-enhanced-dacvae.audioautomatic-speech-recognition1M<n<10M0 likes5 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.