datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Emolia
Dataset Card for Emolia
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.emolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.emolia-hq
Emolia-HQ
Emolia-HQ is a high-quality, speaker-paired subset of the LAION Emolia dataset. Each sample includes a target utterance and a reference utterance from the same speaker, enabling speaker-conditioned tasks such as voice conversion, expressive TTS, and speaker-aware emotion recognition.
Source
Derived from laion/Emolia by:
Quality filtering: Only samples with dnsmos >= 3.0 are retained.
Speaker pairing: Each target sample is matched with a reference audio from the… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-hq.emolia-3k-speaker-clusters-DACVAE
Emolia 3K Speaker Clusters
A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster.
Overview
The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure:
Outlier… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-3k-speaker-clusters-DACVAE.emolia-3k-speaker-clusters
Emolia 3K Speaker Clusters
A curated set of 3,000 diverse speaker clusters derived from the TTS-AGI/emolia-hq dataset, with up to 20 representative audio samples per cluster.
Overview
The original emolia-hq dataset contains hundreds of thousands of speech samples with 128-dimensional WavLM speaker timbre embeddings. These were first clustered into 10,000 centroids, then intelligently pruned to 3,000 using density-aware farthest-point sampling to ensure:
Outlier… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-3k-speaker-clusters.emolia-balanced-5M-subset
emolia-balanced-5M-subset
A balanced ~5.26M-sample subset of laion/Emolia (80.5M speech samples), packaged as WebDataset-compatible tar shards for direct use in training pipelines.
How this subset was filtered
Samples were selected if they met either of two criteria:
1. Emotion thresholds
Each sample carries 40 emotion annotation scores (from the Emonet taxonomy) in its metadata. A sample qualifies for an emotion bucket if its score for that emotion meets or… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-balanced-5M-subset.
