datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.Emotion_new_collected_datasetlaions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.Emilia-with-Emotion-Annotations4microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.Emilia-with-Emotion-Annotations5CASIA_speech_emotion_recognitionqwen3-tts-multilingual-emotional-speechQWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-EmotionITA-Corpus Emotion Dataset (100 Japanese Female Voices)
彼のあだ名は言い得て妙だよね
11:A lower-pitched female voice with a strong core
ヒューズが飛んだ
100:A slightly quirky female voice that leaves a strong impression
Overview
This dataset contains 100 female voices generated with Qwen3-TTS.
Format: 24kHz mono WAV
Source: Link to designed voices
About ITA-Corpus Emotion
The text is based on the ITA-Corpus Emotion, a public domain dataset containing 100… See the full description on the dataset page: https://huggingface.co/datasets/Akjava/QWEN3-TTS-Voice-Clone-100-Japanese-Female-ITA-Corpus-Emotion.EMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.Emilia-with-Emotion-Annotations3Music_by_Emotion🎵 Music by Emotion Dataset
Dataset Summary
The Music by Emotion dataset is a custom audio dataset designed for supervised music emotion recognition tasks. It consists of 1,000 music samples, each a 30-second audio clip, sourced from publicly available SoundCloud content.
Each clip is labeled according to its emotional content, derived from musical characteristics such as tempo and musical key.
🎯 Task
Task type: Audio Classification
Domain: Music Emotion Recognition
Input: 30-second music… See the full description on the dataset page: https://huggingface.co/datasets/lossminimilization/Music_by_Emotion.Emilia-with-Emotion-Annotations2speech-emotion-dataset-consolidatedmeld_emotion_test@article{poria2018meld,
title={Meld: A multimodal multi-party dataset for emotion recognition in conversations},
author={Poria, Soujanya and Hazarika, Devamanyu and Majumder, Navonil and Naik, Gautam and Cambria, Erik and Mihalcea, Rada},
journal={arXiv preprint arXiv:1810.02508},
year={2018}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large Language Models},
author={Wang, Bin and Zou, Xunlong and Lin, Geyu and Sun, Shuo and Liu, Zhuohan and… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/meld_emotion_test.emotional-speech-audio-dataset-3eng-4noneng-updatedspeech-emotion-recognition
Speech Emotion Recognition
Dataset comprises 30,000+ audio recordings featuring 4 distinct emotions: euphoria, joy, sadness, and surprise. This extensive collection is designed for research in emotion recognition, focusing on the nuances of emotional speech and the subtleties of speech signals as individuals vocally express their feelings.
By utilizing this dataset, researchers and developers can enhance their understanding of sentiment analysis and improve automatic speech… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/speech-emotion-recognition.AffectDF_EmotionSDD
AffectDF: Emotionally Expressive Speech Deepfake Benchmark
Overview
AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks.
AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.TTS_emotional
Dataset Summary
TTS_emotional is a speech dataset built for training and evaluating expressive / emotional
text-to-speech (TTS) systems. Each example pairs a short spoken-word audio clip with its
transcript, a natural-language description of how the line is delivered, and structured
metadata about the voice, speaking style/emotion, and speaker gender. Many of the transcripts
are short educational explanations (e.g. "why do we forget things", "how does a compass work"),
each read… See the full description on the dataset page: https://huggingface.co/datasets/SeifElden2342532/TTS_emotional.CAMEO-emotion-classification
CAMEO multilingual speech emotion classification (MTEB)
Speech emotion recognition across five languages, drawn from the CAMEO collection.
Labels index this list:
anger
fear
happiness
neutral
sadness
surprise
Source: amu-cai/CAMEO at revision 38e9e96, cc-by-nc-sa-4.0. Split by
speaker so no speaker appears in both train and test. Only the six emotions common
to every included language are kept. Audio is 16 kHz Opus.
Built by scripts/data/cameo_emotion/create_data.py in the… See the full description on the dataset page: https://huggingface.co/datasets/vnahata/CAMEO-emotion-classification.jvnv-emotional-speech-corpusemotional-roleplay-finetuning-dataset
Artificial Voice Roleplay Dataset
67,491 fully-synthetic speech clips (~184 hours) pairing expressive role-play / character
voice-direction captions with generated audio, across German, English, Spanish, and French
(German-dominant). Rich in exaggerated fantasy/creature voices (orc, goblin, troll, ogre,
zombie, dragon, demon, witch, banshee, imp, fairy, gnome, robot, murloc, harpy, skeleton, ghost,
vampire …) and high-arousal emotional delivery (rage, fear, grief, menace).
Every… See the full description on the dataset page: https://huggingface.co/datasets/laion/emotional-roleplay-finetuning-dataset.bangla-emotion-maleiemocap_emotion_recognition@article{busso2008iemocap,
title={IEMOCAP: Interactive emotional dyadic motion capture database},
author={Busso, Carlos and Bulut, Murtaza and Lee, Chi-Chun and Kazemzadeh, Abe and Mower, Emily and Kim, Samuel and Chang, Jeannette N and Lee, Sungbok and Narayanan, Shrikanth S},
journal={Language resources and evaluation},
volume={42},
pages={335--359},
year={2008},
publisher={Springer}
}
@article{wang2024audiobench,
title={AudioBench: A Universal Benchmark for Audio Large… See the full description on the dataset page: https://huggingface.co/datasets/AudioLLMs/iemocap_emotion_recognition.ssi-speech-emotion-recognition
Dataset Card for SSI: Speech Emotion Recognition - Stapes AI
Dataset Details
Dataset Format for Audio Files
This is the format for the audio files in the dataset. We'll open-source the dataset soon.
Gender
M - Male
F - Female
Age Group
CH - Child (0-12)
TE - Teenager (13-19)
AD - Adult (20-60)
SE - Senior (60+)
UNK - Unknown
Utterance Type
SEN: Sentence
WOR: Word
PHR: Phrase
Sentence
DFA: "Don't Forget A… See the full description on the dataset page: https://huggingface.co/datasets/stapesai/ssi-speech-emotion-recognition.arabic-multidialect-emotional-speech-demo
DataHive AI — Demo: Arabic Multi-Dialect Emotional Speech
A DataHive AI dataset: a stratified 1-hour demo sample from a full corpus of 50+ hours. We can also create larger audio datasets upon client request.
Most public Arabic speech corpora flatten dialect into a single label and ignore emotion entirely. This corpus does the opposite: every recording is tagged with one of four regional Arabic dialects (Najdi, Hejazi, Jordanian, Moroccan) and one of four target emotions (Sad, Happy… See the full description on the dataset page: https://huggingface.co/datasets/datahiveai/arabic-multidialect-emotional-speech-demo.egyption-with-emotion-dataset
Egption Text-Audio Dataset With Emotions and Diarization
Creating datasets for TTS and ASR models with emotions and Diarization
In case you want to focus only one speaker , you can fiter based on speaker_role
Source Code
if you want to collect more data from youtube, you can check this link
🙏 Acknowledgements
This project makes use of the forced alignment model and Cohere ASR model provided by:
MahmoudAshraf/mms-300m-1130-forced-aligner
Cohere ASR
Hubert… See the full description on the dataset page: https://huggingface.co/datasets/OmarAhmedSobhy/egyption-with-emotion-dataset.common_voice_17_0_emotion_25k_Whisper_Compatibleurdu-emotions
URDU-Dataset
1. General information
URDU dataset contains emotional utterances of Urdu speech gathered from Urdu talk shows. It contains 300 utterances of four basic emotions: Angry, Happy, and Neutral. There are 38 speakers (27 male and 11 female). This data is created from YouTube. Speakers are selected randomly.
For more details about dataset, please refer the complete paper "Cross Lingual Speech Emotion Recognition: Urdu vs. Western Languages".… See the full description on the dataset page: https://huggingface.co/datasets/2DamnWav/urdu-emotions.
