emolia
Datasets
All datasets matching “emolia”emolia-thinking
Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia
Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.Emolia
Dataset Card for Emolia
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.emolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.emolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.emolia_filtered_nano_codec_21_dataset
Emolia · Filtered · NanoCodec (FSQ) Tokens
A cleaned, pre-tokenized version of laion/Emolia
prepared for text-to-speech (TTS) training.
The pipeline is two steps:
Quality filtering with the open-source
audio_filter tool — this
removes the dirtiest recordings (noise, clipping, band-limiting, robotic artifacts,
overlapping speakers), which matters a lot for TTS quality.
Discrete audio tokenization with NVIDIA
nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps
(an FSQ neural audio… See the full description on the dataset page: https://huggingface.co/datasets/nineninesix/emolia_filtered_nano_codec_21_dataset.emolia-hq
Emolia-HQ
Emolia-HQ is a high-quality, speaker-paired subset of the LAION Emolia dataset. Each sample includes a target utterance and a reference utterance from the same speaker, enabling speaker-conditioned tasks such as voice conversion, expressive TTS, and speaker-aware emotion recognition.
Source
Derived from laion/Emolia by:
Quality filtering: Only samples with dnsmos >= 3.0 are retained.
Speaker pairing: Each target sample is matched with a reference audio from the… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/emolia-hq.
