datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.Emotion_new_collected_datasetemolia-thinking
Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia
Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.Emolia
Dataset Card for Emolia
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.emo_webds_2emo_parleremo_webdsemolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset
A balanced, per-dimension bucket subset of
VoiceNet/emolia-thinking,
derived from that dataset's zero-shot VoiceNet-dimension labels.
For every VoiceNet voice/prosody/timbre/style dimension, this subset draws a
roughly equal number of clips from each ordinal bucket (0–6), so that
downstream training / probing sees a balanced distribution along each axis
instead of the strongly skewed natural distribution.
How… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-thinking-balanced-buckets.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.Emilia-with-Emotion-Annotations4EmoFake_test
EmoFake Test
Benchmark-ready packaging of the EmoFake test set for speech anti-spoofing.
Overview
Emotional speech deepfake detection test set. Contains bonafide emotional utterances and spoofed samples with emotion conversion.
License
CC BY 4.0. See LICENSE.txt.
Schema
Column
Type
Description
path
string
Audio filename
audio
Audio(16000)
Audio waveform, 16 kHz mono
label
ClassLabel
bonafide (index 0) or spoof (index 1)… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/EmoFake_test.EmoSpoofTTS
EmoSpoofTTS
A spoof-only attack corpus of emotional text-to-speech (TTS) synthesis:
36,000 clips spanning 3 modern TTS systems, 10 speakers, and 4 emotions, all
synthesized from transcripts of the Emotional Speech Dataset (ESD).
Overview
EmoSpoof-TTS (Mahapatra et al., "Can Emotion Fool Anti-spoofing?", Interspeech
2025, arXiv:2505.23962) was built to study whether emotionally expressive TTS
is harder for anti-spoofing systems to detect than neutral TTS. For 10… See the full description on the dataset page: https://huggingface.co/datasets/SpeechAntiSpoofingBenchmarks/EmoSpoofTTS.qwen3-tts-multilingual-emotional-speechEmilia-with-Emotion-Annotations5emo_speech_filtered_v12 second filtered emotional speech in webdataset format
https://huggingface.co/datasets/EQ4You/Emotional_Speech
irodori-clones-3m-v2-no-emoji
Irodori TTS Clones v2 (3.29M)
3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2,
using the 10,000 reference voices from SynData-2/irodori-refs-10k-v2.
329 clones per ref voice, each with a unique Japanese conversational text.
Companion refs: SynData-2/irodori-refs-10k-v2.
Note: Bu dataset irodori-clones-3m-v2'nin emoji-temizlenmis kopyasidir. Audio bytes binary-identical; yalnizca text kolonundaki emojiler kaldirilmistir (emoji kutuphanesi, Japonca/CJK… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA/irodori-clones-3m-v2-no-emoji.CASIA_speech_emotion_recognitionEmoV_DBQWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-EmotionITA-Corpus Emotion Dataset (100 Japanese Female Voices)
彼のあだ名は言い得て妙だよね
11:A lower-pitched female voice with a strong core
ヒューズが飛んだ
100:A slightly quirky female voice that leaves a strong impression
Overview
This dataset contains 100 female voices generated with Qwen3-TTS.
Format: 24kHz mono WAV
Source: Link to designed voices
About ITA-Corpus Emotion
The text is based on the ITA-Corpus Emotion, a public domain dataset containing 100… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-Emotion.emolia-hq
Emolia-HQ
Emolia-HQ is a high-quality, speaker-paired subset of the LAION Emolia dataset. Each sample includes a target utterance and a reference utterance from the same speaker, enabling speaker-conditioned tasks such as voice conversion, expressive TTS, and speaker-aware emotion recognition.
Source
Derived from laion/Emolia by:
Quality filtering: Only samples with dnsmos >= 3.0 are retained.
Speaker pairing: Each target sample is matched with a reference audio… See the full description on the dataset page: https://huggingface.co/datasets/laion/emolia-hq.EMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.EMO-transcribed-1lineNemoMusic_by_Emotion🎵 Music by Emotion Dataset
Dataset Summary
The Music by Emotion dataset is a custom audio dataset designed for supervised music emotion recognition tasks. It consists of 1,000 music samples, each a 30-second audio clip, sourced from publicly available SoundCloud content.
Each clip is labeled according to its emotional content, derived from musical characteristics such as tempo and musical key.
🎯 Task
Task type: Audio Classification
Domain: Music Emotion Recognition
Input: 30-second music… See the full description on the dataset page: https://huggingface.co/datasets/lossminimilization/Music_by_Emotion.kz_emo_speech_finalyue_emo_speech
Cantonese Emotional Speech
Crawled from YouTube and RTHK, this dataset contains 1,000 hours of Cantonese speech, each labeled with one of the following emotions: angry, disgusted, fearful, happy, neutral, other, sad, or surprised. The dataset also includes the confidence of the emotion label. The audio files are denoised with resemble-enhance. The transcriptions are generated by SenseVoiceSmall, and deduplicated using MinHash.
Emilia-with-Emotion-Annotations3emolia
emolia-balanced-5M-subset · flac 48 kHz · WebDataset (paired)
This is the emolia-balanced-5M-subset corpus repackaged for high-quality
audio–text contrastive training. Audio is re-encoded as mono FLAC at 48 kHz
(PCM 16-bit) and stored as a WebDataset of paired <key>.flac + <key>.json
samples.
The JSON sidecar carries the full annotation stack:
Original metadata (id, text, duration, speaker, language, dnsmos).
A free-text emotion_caption derived from the emotion-annotation scalars.… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia.Emilia-with-Emotion-Annotations2
