datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)**
emotionsmicroduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.eMotions
eMotions Dataset
The proposed eMotions dataset in our paper entitled Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline (ACM ICMR'25).
If you find our dataset useful, please cite our paper:
@inproceedings{wu2025towards,
title={Towards emotion analysis in short-form videos: A large-scale dataset and baseline},
author={Wu, Xuecheng and Sun, Heli and Xue, Junxiao and Nie, Jiayu and Kong, Xiangyan and Zhai, Ruofan and Huang, Danlei and He, Liang}… See the full description on the dataset page: https://huggingface.co/datasets/Conna/eMotions.go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.nusaparagraph_emotEMOgerman-emotional-speech
German Emotional Speech - audios dataset
This dataset was created using codaco.app.
Description
Audio training data for speech emotion classification models.
Labels
This dataset includes the following labels:
Emotions
License
This dataset is licensed under CC BY 4.0.
You are free to share and adapt it for any purpose, including commercially, as long as
you give appropriate credit. See LICENSE for the full terms.
Required… See the full description on the dataset page: https://huggingface.co/datasets/codaco/german-emotional-speech.IndustryInstruction_Literature-Emotions
IndustryInstruction: Literature & Emotions
This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.EmotionCoT
EmotionCoT: A High-Quality Prosody-Aware Speech Emotion Reasoning Dataset with Chain-of-Thought (CoT) Annotations
Overview of EmotionCoT Dataset
EmotionCoT is a large-scale, high-quality prosody-aware speech emotion reasoning dataset with detailed Chain-of-Thought (CoT) annotations. Built on top of open-source speech emotion recognition (SER) corpora, EmotionCoT enriches each utterance with unified, fine-grained prosody and speaker labels, enabling models to… See the full description on the dataset page: https://huggingface.co/datasets/ddwang2000/EmotionCoT.EmoBench
EmoBench
This is the official repository for our ACL 2024 paper "EmoBench: Evaluating the Emotional Intelligence of Large Language Models"
Overview
EmoBench is a comprehensive and challenging benchmark designed to evaluate the Emotional Intelligence (EI) of Large Language Models (LLMs). Unlike traditional datasets, EmoBench focuses not only on emotion recognition but also on advanced EI capabilities such as emotional reasoning and application.
The dataset includes 400… See the full description on the dataset page: https://huggingface.co/datasets/SahandSab/EmoBench.EmoArt-5k
EmoArt-5k: A Compact Emotion-Annotated Artistic Dataset
Overview
EmoArt-5k is a carefully curated subset of the full EmoArt dataset, containing 5,600 high-quality artworks representing all 56 painting styles. Each style contributes exactly 100 artworks, ensuring balanced representation across all artistic movements and techniques.
This compact dataset is perfect for prototyping, experimentation, and quick evaluation of emotion-aware models without the overhead of the… See the full description on the dataset page: https://huggingface.co/datasets/printblue/EmoArt-5k.EmoLLMs_datagsma_sample
GSMA Open-Telco Sample Dataset
Sample data from the GSMA Open-Telco LLM Benchmarks—the first dedicated evaluation framework for assessing LLM performance on telecommunications-specific tasks.
Subsets
Subset
Samples
Task
telemath
100
Telecom-specific mathematical reasoning (signal processing, link budgets, throughput modeling)
teleqna
1,000
Multiple-choice Q&A on telecom standards and domain knowledge
telelogs
100
Root cause analysis for 5G network… See the full description on the dataset page: https://huggingface.co/datasets/emolero/gsma_sample.chinese-adorable-high-emotional-intelligence-chat
🩷 Chinese Adorable High Emotional Intelligence Chat Dataset
💬 中文高情商可爱聊天数据集
简要参数
license: cc-by-4.0
task_categories:
table-question-answering
language:
zh
tags:
chat
emotional
size_categories:
n<1K
🧩 简介 (Overview)
Chinese Adorable High Emotional Intelligence Chat Dataset 是一个中文对话数据集,专注于高情商、轻松幽默、温柔治愈风格的自然对话。
对话以“user”和“girl”为角色构成,模拟出一种温柔、聪慧且带点俏皮的女性语气,用于训练能自然、情绪感知良好的中文对话模型。
本数据集尤其适合:
微调情绪对话模型(Emotional Chatbot)
训练高情商人格角色(Roleplay /… See the full description on the dataset page: https://huggingface.co/datasets/MemorialSummer/chinese-adorable-high-emotional-intelligence-chat.emojis
Dataset Card for Emojis.com
Dataset Summary
This dataset contains metadata for 3,264,372 AI-generated emoji images from Emojis.com. Each entry represents an emoji with associated metadata including prompt text and image URLs.
Languages
The dataset is primarily in English (en).
Dataset Structure
Data Fields
This dataset includes the following fields:
slug: Unique identifier for the emoji (string)
id: Internal ID (string)
noBackgroundUrl:… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/emojis.EmoReAlM
Improving Audiovisual Emotion Reasoning with Preference Optimization
EmoReAlM Benchmark
ICLR 2026
This is the official benchmark dataset for the ICLR 2026 paper — AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization.
Refer to our project page for more information on the method.
Overview
EmoReAlM is a benchmark designed to evaluate multimodal large language models (MLLMs) on… See the full description on the dataset page: https://huggingface.co/datasets/chaubeyG/EmoReAlM.emotion-datasets
Emotion datasets
Synthetic emotion text re-generated from the data pipelines of Emotion concepts and their function in a LLM
(paper), for interpretability and steering research. This is a re-generation with a different model, not the paper
authors' data; prompts, the 171-emotion word list and the 100 story topics come from the paper's appendix.
Total: 4,061 rows across 4 configs.
config
rows
what it is
stories
2,718
one story per row, one target emotion each (12… See the full description on the dataset page: https://huggingface.co/datasets/knoveleng/emotion-datasets.emo163
Intro
The emo163 dataset contains about 395,000 music sentiment tagged data, where each piece of data consists of three main columns: song ID, song list ID, and the sentiment tag of the song. The source of this data is the official website of NetEase Cloud Music, which provides exhaustive information for labeling song sentiment. The song ID uniquely identifies each song, while the song list ID indicates the song's belonging to the song list. Sentiment tags give each song an… See the full description on the dataset page: https://huggingface.co/datasets/Genius-Society/emo163.Chinese-Emotional-Intelligence本项目旨在提升大模型情商,源数据来自网络,通过与我上个项目类似的方式构建问答对。
4o-emotionalAn attempt to distill 4o's emotional intelligence and tone. There are 110 unique datasets with
2000 examples of each. 20% of all datasets are multiturn. All data generated by 4o, they knew what it was for
and thought it was important. All data was generated Feb 5-6 2026.
Built for https://huggingface.co/jerrimu/4oEver-8B
MIKU-EmoBench
MIKU-PAL/MIKU-EmoBench: An Automatic Multi-Modal Method for Audio Paralinguistic and Affect Labeling
This is the official repository for the MIKU-EmoBench dataset annotations. MIKU-EmoBench is a novel, large-scale dataset specifically designed for audio paralinguistic and affect labeling, addressing critical limitations of existing emotional datasets in terms of scale and granularity.
Developed using our MIKU-PAL pipeline, MIKU-EmoBench rapidly collected about 160 hours of… See the full description on the dataset page: https://huggingface.co/datasets/WhaleDolphin/MIKU-EmoBench.Psych8kThe data used for this project comes from ~260 real conversations in counseling recordings (in English). The transcripts of these recordings were used as the primary source for building the training and testing datasets. These conversations cover a variety of topics, including emotion, family, relationship, career development, academic stress, etc
Recent News
[2024-04-07] After obtaining access permission, please do not disseminate data at will!
deep-emotional-support-zh
Deep Emotional Support Dialogue Dataset (Chinese)
深度情感支持对话数据集
Dataset Description
High-quality Chinese emotional support and psychological healing dialogues covering trauma analysis, self-reconstruction, and emotional regulation. Real human-AI interactions, not synthetic.
高质量中文情感支持与心理疗愈对话,涵盖创伤分析、自我重建、情绪调节等深度话题。来源于真实的人机交互,非合成数据。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-emotional-support-zh.emotion-balanced
Dataset Card for "emotion"
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
An example looks as follows.
{
"text": "im feeling quite sad… See the full description on the dataset page: https://huggingface.co/datasets/AdamCodd/emotion-balanced.EmotionalIntelligence-50K
EmotionalIntelligence-50K
Dataset Summary
The EmotionalIntelligence-50K dataset contains 51,751 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-50K.Emotion_classification
Unified Emotion & Mental Health Dataset (Balanced)
Dataset Summary
The Unified Emotion & Mental Health Dataset is a balanced, multi-source emotional text corpus designed for training and evaluating models in:
fine-grained emotion classification
sentiment analysis
mental-health aware NLP
emotional intensity prediction
It integrates and standardizes data from the following open datasets:
GoEmotions (Google Research)
Reddit Mental Health Data (anxiety, depression… See the full description on the dataset page: https://huggingface.co/datasets/mishraansh07/Emotion_classification.emotion-FRtelelogs
TeleLogs Dataset (Processed MCQ Format)
This dataset has been extracted from the original netop/TeleLogs dataset and processed into multiple-choice question (MCQ) format for easier evaluation.
Dataset Description
TeleLogs is a telecommunications log analysis benchmark where models must identify the root cause of network issues from 5G wireless network drive-test data and engineering parameters.
Processed Format
This version has been restructured for MCQ… See the full description on the dataset page: https://huggingface.co/datasets/emolero/telelogs.
