laion/clustered-reference-voices
Clustered Reference Voices (EMOLIA 3K) 3,000 enhanced reference voice MP3s — one high-quality representative sample per speaker cluster, selected and scored by a multi-expert neural quality model. Overview Property Value Total clips 3,000 Total duration 11.3 hours Mean duration 13.5 s (range: 3.5 – 29.9 s) Format MP3, 192 kbps, 48 kHz Language English Naming {cluster_id}.mp3 (0 – 2999) Source The source data is… See the full description on the dataset page: https://huggingface.co/datasets/laion/clustered-reference-voices.
Clustered Reference Voices (EMOLIA 3K)
3,000 enhanced reference voice MP3s — one high-quality representative sample per speaker cluster, selected and scored by a multi-expert neural quality model.
Overview
Source
The source data is laion/emolia-3k-speaker-clusters, which contains 3,000 speaker clusters with approximately 20 samples each (59,977 total utterances). Clusters were produced by grouping speaker embeddings from a diverse collection of English speech.
Processing Pipeline
Each of the 59,977 source utterances was processed through a two-stage pipeline:
1. Speech Enhancement — ClearerVoice MossFormer2SE48K
All audio was enhanced at 48 kHz using the MossFormer2_SE_48K speech enhancement model. This removes background noise, music, reverb, and other non-speech artifacts while preserving the natural characteristics of the speaker's voice.
2. Quality Scoring — Empathic Insight Voice Plus
Enhanced audio was scored by the Empathic Insight Voice Plus model, which employs 59 MLP expert heads on top of Whisper encoder embeddings. The model produces multiple quality dimensions:
3. Selection — Top Sample per Cluster
For each of the 3,000 speaker clusters, the single sample with the highest `overall_quality` score was selected as the cluster's representative reference voice.
Quality Statistics
Dataset Files
Metadata Schema (parquet)
Intended Uses
- TTS reference voices: High-quality, diverse speaker references for text-to-speech systems
- Voice cloning: Clean, enhanced single-speaker clips suitable as cloning targets
- Speaker verification benchmarks: One representative per cluster for speaker ID tasks
- Quality filtering research: Studying the relationship between quality scores and perceptual quality
Interactive Gallery
The included gallery.html file provides a self-contained, browser-based interface to explore all 3,000 samples. Features:
- Embedded base64 audio playback (no server required)
- Sortable columns (click any header)
- Full-text search across cluster IDs and transcripts
- Quality score display for all dimensions
Citation
@dataset{clustered_reference_voices_2026,
title={Clustered Reference Voices (EMOLIA 3K)},
author={LAION},
year={2026},
url={https://huggingface.co/datasets/laion/clustered-reference-voices}
}License
This dataset is released under the CC-BY-4.0 license.
