datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emotion
Dataset Card for "emotion"
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
An example looks as follows.
{
"text": "im feeling quite sad and sorry for myself but… See the full description on the dataset page: https://huggingface.co/datasets/dair-ai/emotion.emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)**
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.emotion
EmotionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise.
Task category
t2c
Domains
Social, Written
Reference
https://www.aclweb.org/anthology/D18-1404
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["EmotionClassification"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/emotion.emotionsRobot-EQ
RobotEQ-Data
Official dataset release for RobotEQ.
Evaluation & Scripts
For inference scripts, evaluation scripts, and data production tooling, see the RobotEQ code repository.
Dataset Statistics
Item
Count
Behavior judgment scenarios (synthetic)
1,812
Behavior judgment scenarios (real POV)
223
Behavior judgment scenarios (total)
2,035
Behavior judgment behavior annotations
3,171
Spatial grounding questions
825… See the full description on the dataset page: https://huggingface.co/datasets/Tongji-Emotion/Robot-EQ.laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.BRIGHTER-emotion-categories
BRIGHTER Emotion Categories Dataset
This dataset contains the emotion categories data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages.
Dataset Description
The BRIGHTER Emotion Categories dataset is a comprehensive multi-language, multi-label emotion classification dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.Emilia-with-Emotion-Annotations4microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.Emilia-with-Emotion-Annotations5laion-emotional-trajectory-t80
LAION Emotional-Trajectory Speech — tier T≥0.80
319,765 crossfaded speech trajectories · 4,482 audio-hours · 1,598,825 source clips
A trajectory is a short sequence of 5 consecutive utterances by one
speaker whose measured emotion or voice character moves monotonically from one end of
the corpus distribution to the other. The clips are joined into one continuous audio file
with equal-power crossfades, the joined audio is re-tokenized with MOSS-Audio-
Tokenizer-v2, and every… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-emotional-trajectory-t80.go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
eMotions
eMotions Dataset
The proposed eMotions dataset in our paper entitled Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline (ACM ICMR'25).
If you find our dataset useful, please cite our paper:
@inproceedings{wu2025towards,
title={Towards emotion analysis in short-form videos: A large-scale dataset and baseline},
author={Wu, Xuecheng and Sun, Heli and Xue, Junxiao and Nie, Jiayu and Kong, Xiangyan and Zhai, Ruofan and Huang, Danlei and He, Liang}… See the full description on the dataset page: https://huggingface.co/datasets/Conna/eMotions.qwen3-tts-multilingual-emotional-speechsynthetic-emotions
Synthetic Emotions Dataset
Overview
Synthetic Emotions is a video dataset of AI-generated human emotions created using OpenAI Sora. It features short (5-sec, 480p, 9:16) videos depicting diverse individuals expressing emotions like happiness, sadness, anger, fear, surprise, and more.
This dataset is ideal for emotion recognition, facial expression analysis, affective computing, and AI-human interaction research.
Dataset Details
Total Videos: 100
Video Format:… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/synthetic-emotions.EMID-Emotion-Matching
EMID-Emotion-Matching
orrzohar/EMID-Emotion-Matching is a derived dataset built on top of
the Emotionally paired Music and Image Dataset (EMID) from ECNU (ecnu-aigc/EMID).
It is designed for music ↔ image emotion matching with Qwen-Omni–style models.
Each example contains:
audio: mono waveform stored as datasets.Audio (HF Hub preview can play it)
sampling_rate: sampling rate used when decoding (typically 16 kHz)
image: a single image (datasets.Image)
same: bool, whether the audio… See the full description on the dataset page: https://huggingface.co/datasets/orrzohar/EMID-Emotion-Matching.Emotiontalk
EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
Introduction
EmotionTalk is an interactive Chinese multimodal emotion dataset with rich annotations. This dataset provides multimodal information from 19 actors participating in dyadic conversation settings, incorporating acoustic, visual, and textual modalities. It includes 23.6 hours of speech (19,250 utterances), annotations for 7 utterance-level emotion categories (happy… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/Emotiontalk.emotion-clever-datasetEmilia-with-Emotion-Annotations3short-text-labeled-emotion-classificationEmilia-with-Emotion-Annotations2ru_emotion_dvach
Русский датасет эмоций (E-Kulture)
Датасет содержит тексты с разметкой по пяти эмоциональным категориям:
Агрессия (aggression)
Тревожность (anxiety)
Сарказм (sarcasm)
Позитив (positive)
Нейтральное состояние (neutral)
Структура датасета
Датасет разделен на два сплита:
train - обучающая выборка
valid - валидационная выборка
Каждый пример содержит следующие поля:
text - текст сообщения
label - метка эмоции (одна из пяти категорий)
Создание
Исходные данные… See the full description on the dataset page: https://huggingface.co/datasets/Kostya165/ru_emotion_dvach.Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.many_emotions
Many Emotions
Many Emotions is a 2.7-million-row multilingual text-classification dataset for recognizing seven emotion categories in English, French, Italian, Spanish, and German. It combines examples from Emotion, DailyDialog, and GoEmotions.
The default unsplit corpus contains 2,710,740 non-empty rows derived from 550,123 source IDs. Every row records its source dataset and source-specific license.
The current release is 2.0.0. See CHANGELOG.md for changes from the original… See the full description on the dataset page: https://huggingface.co/datasets/ma2za/many_emotions.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.super-emotion
Super Emotion Dataset
Associated Paper
We provide full documentation of the dataset construction process in the accompanying paper: 📘 The Super Emotion Dataset (PDF)
Dataset Summary
The Super Emotion dataset is a large-scale, multilabel dataset for emotion classification, aggregated from six prominent emotion datasets:
MELD
GoEmotions
TwitterEmotion
ISEAR
SemEval
Crowdflower
It contains 552,821 unique text samples and 570,457 total emotion label assignments… See the full description on the dataset page: https://huggingface.co/datasets/cirimus/super-emotion.speech-emotion-dataset-consolidated
