datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.reference-voices-enhanced
Reference Voices Enhanced
2,004 AI voice samples enhanced with ClearerVoice-Studio MossFormer2_SE_48K speech enhancement, annotated with Empathic Insight Voice Plus (59 quality + emotion scores).
Dataset Summary
Source: laion/ai-voices-deduplicated (2,004 speaker-deduplicated, quality-filtered AI voice samples)
Speech Enhancement: ClearerVoice MossFormer2_SE_48K — background noise removal and speech clarity improvement
Output Format: Enhanced WAV files at 48kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference-voices-enhanced.clustered-reference-voices
Clustered Reference Voices (EMOLIA 3K)
3,000 enhanced reference voice MP3s — one high-quality representative sample per speaker cluster, selected and scored by a multi-expert neural quality model.
Overview
Property
Value
Total clips
3,000
Total duration
11.3 hours
Mean duration
13.5 s (range: 3.5 – 29.9 s)
Format
MP3, 192 kbps, 48 kHz
Language
English
Naming
{cluster_id}.mp3 (0 – 2999)
Source
The source data is… See the full description on the dataset page: https://huggingface.co/datasets/laion/clustered-reference-voices.en_and_de_reference_voice_files_for_emotion_cloningEmotion-Voice-Attribute-Reference-Snippets-DACVAE
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: ~98 GB (latents-only, no WAV)
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.
