datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
moss-character-reference-voices
MOSS character reference voices (1336 voices)
1336 distinct synthetic character voices, each mined from a cluster of generated MOSS-VA-v2 character
audio and auto-annotated by Gemini-3-Flash. For every cluster the model was shown the 3 cluster samples
their automatic voice scores, chose the single most representative sample, and wrote a full
casting-style profile.
Contents
dataset.jsonl — one row per voice: cid, name, tagline, description, age, gender, register… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-character-reference-voices.6k-diverse-reference-voices
6k Diverse Reference Voices
6,064 permissively licensed reference voices for casting expressive voice-acting generations.
All voices in this collection are permissively usable: they were either synthetically created or
extracted from the CC-BY part of Emilia. Licensed under CC-BY-4.0.
Source / attribution: derived from TTS-AGI/moss-reference-voices-consolidated (CC-BY-4.0),
re-published under LAION with clarified metadata documentation. If you use this dataset,
please attribute… See the full description on the dataset page: https://huggingface.co/datasets/laion/6k-diverse-reference-voices.Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: 1.54 TB
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40 emotions + 15 voice… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE-Wave.clustered-reference-voices
Clustered Reference Voices (EMOLIA 3K)
3,000 enhanced reference voice MP3s — one high-quality representative sample per speaker cluster, selected and scored by a multi-expert neural quality model.
Overview
Property
Value
Total clips
3,000
Total duration
11.3 hours
Mean duration
13.5 s (range: 3.5 – 29.9 s)
Format
MP3, 192 kbps, 48 kHz
Language
English
Naming
{cluster_id}.mp3 (0 – 2999)
Source
The source data is… See the full description on the dataset page: https://huggingface.co/datasets/laion/clustered-reference-voices.reference_ai_voices_with_timbre_annotations
Overview
reference_voice_dataset__mp3 is a collection of around 32,000 AI-generated voice samples.
The clips were created with different neural TTS / voice-conversion tools and are designed to
cover a broad emotional spectrum and a wide variety of vocal timbres and character types.
Voices span:
Young adults to elderly speakers
Masculine, feminine, and androgynous presentations
Dark vs. bright, soft vs. harsh, warm vs. cool timbres
Neutral, everyday voices and highly stylised… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference_ai_voices_with_timbre_annotations.reference-voices-enhanced
Reference Voices Enhanced
2,004 AI voice samples enhanced with ClearerVoice-Studio MossFormer2_SE_48K speech enhancement, annotated with Empathic Insight Voice Plus (59 quality + emotion scores).
Dataset Summary
Source: laion/ai-voices-deduplicated (2,004 speaker-deduplicated, quality-filtered AI voice samples)
Speech Enhancement: ClearerVoice MossFormer2_SE_48K — background noise removal and speech clarity improvement
Output Format: Enhanced WAV files at 48kHz… See the full description on the dataset page: https://huggingface.co/datasets/laion/reference-voices-enhanced.moss-reference-voices-consolidated
MOSS reference voices — consolidated (6,064 voices)
6,064 reference voices for casting MOSS-VA-v2 voice-acting generations. Every voice was auto-annotated
by Gemini (name, tagline, language, accent, age/gender read, register, timbre, distinctive features,
emotional range, casting suggestions for 4 genres, free-text tags, search text) and scored on 99 measured
dimensions: 57 VoiceNet voice-quality axes (timbre/prosody/register/speaking-style, e.g. brightness,
roughness, warmth… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/moss-reference-voices-consolidated.voice-referenceen_and_de_reference_voice_files_for_emotion_cloningemotion_dataset_for_tts_with_transcriptions_and_reference_voice_v1Emotion-Voice-Attribute-Reference-Snippets-DACVAE
Emotion and Voice Attribute Reference Snippets - DACVAE and Wave
Merged dataset combining TTS-AGI/enhanced-emo-snippets-balanced-DACVAE and
TTS-AGI/emotion-attribute-conditioning-dacvae with decoded WAV audio.
Overview
Total samples: 606,178
Filtered out: 363,331 (samples with speech_quality < 1.8)
Total tar files: 328
Total size: ~98 GB (latents-only, no WAV)
Audio format: WAV, 48kHz, PCM 16-bit mono
Latents: DAC-VAE float16 [T, 128] at 25 frames/sec
Dimensions: 57 (40… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/Emotion-Voice-Attribute-Reference-Snippets-DACVAE.guzman-clone-voice-reference
