datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emolia-thinking
Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia
Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.emolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.
