datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.ru_emotion_dvach
Русский датасет эмоций (E-Kulture)
Датасет содержит тексты с разметкой по пяти эмоциональным категориям:
Агрессия (aggression)
Тревожность (anxiety)
Сарказм (sarcasm)
Позитив (positive)
Нейтральное состояние (neutral)
Структура датасета
Датасет разделен на два сплита:
train - обучающая выборка
valid - валидационная выборка
Каждый пример содержит следующие поля:
text - текст сообщения
label - метка эмоции (одна из пяти категорий)
Создание
Исходные данные… See the full description on the dataset page: https://huggingface.co/datasets/Kostya165/ru_emotion_dvach.cleansed_emocontext
Dataset Card for "cleansed_emocontext"
cleansed_emocontext is a cleansed and normalized version of emo.
For cleansing and normalization, data_cleansing.py was used, modifying the code provided on the official EmoContext GitHub.
Dataset Summary
In this dataset, given a textual dialogue i.e. an utterance along with two previous turns of context, the goal was to infer the underlying emotion of the utterance by choosing from four emotion classes - Happy, Sad, Angry and… See the full description on the dataset page: https://huggingface.co/datasets/oneonlee/cleansed_emocontext.tweet_emotion_intensity
Tweet Emotion Intensity Dataset
Papers:
Emotion Intensities in Tweets. Saif M. Mohammad and Felipe Bravo-Marquez. In Proceedings of the sixth joint conference on lexical and computational semantics (*Sem), August 2017, Vancouver, Canada.
WASSA-2017 Shared Task on Emotion Intensity. Saif M. Mohammad and Felipe Bravo-Marquez. In Proceedings of the EMNLP 2017 Workshop on Computational Approaches to Subjectivity, Sentiment, and Social Media (WASSA), September 2017… See the full description on the dataset page: https://huggingface.co/datasets/stepp1/tweet_emotion_intensity.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.emojisA collection of 38,176 emoji images from Facebook, Google, Apple, WhatsApp, Samsung, JoyPixels, Twitter, emojidex, LG, OpenMoji, and Microsoft. It includes all the emojis for these apps/platforms as of early 2022.
Counts: Facebook=3664, Google=3664, Apple=3961, WhatsApp=3519, Samsung=3752, JoyPixels=3538, Twitter=3544, emojidex=2040, LG=3051, OpenMoji=3512, Microsoft=3931.
Sizes: Facebook=144x144, Google=144x144, Apple=144x144, WhatsApp=144x144, Samsung=108x108, JoyPixels=144x144… See the full description on the dataset page: https://huggingface.co/datasets/rocca/emojis.COVID-19_weibo_emotionCOVID-19 Epidemic Weibo Emotional Dataset, the content of Weibo in this dataset is the epidemic Weibo obtained by using relevant keywords to filter during the epidemic, and its content is related to COVID-19.
Each tweet is labeled as one of the following six categories: neutral (no emotion), happy (positive), angry (angry), sad (sad), fear (fear), surprise (surprise)
The COVID-19 Weibo training dataset includes 8,606 Weibos, the validation set contains 2,000 Weibos, and the test dataset… See the full description on the dataset page: https://huggingface.co/datasets/souljoy/COVID-19_weibo_emotion.emotion-dataset-20-emotions
20-Emotion Text Classification Dataset
A comprehensive dataset for fine-grained emotion classification containing 79,595 sentences labeled with 20 distinct emotions.
Dataset Description
This dataset is designed for training emotion classification models that can detect nuanced emotional states in text. Unlike basic sentiment analysis (positive/negative/neutral), this dataset provides fine-grained emotion labels that better capture the complexity of human emotions.… See the full description on the dataset page: https://huggingface.co/datasets/shreyaspullehf/emotion-dataset-20-emotions.EMOPIACheck https://github.com/annahung31/EMOPIA for more info.
social-behavior-emotionsukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.Moroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.Text-Emotion.
text_emotionmultilingual_go_emotions
Overview:
This dataset is updated from on the go_emotions dataset
With the same labels, but add 5 new languages: Arabic, French, Spanish , Dutch, and Turkish.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English Arabic, French, Spanish , Dutch, and Turkish
go_emotionsEmoEvent
EmoEvent: A Multilingual Emotion Corpus based on different Events
In recent years emotion detection in text has become more popular due to its potential applications in fields such as psychology,
marketing, political science, and artificial intelligence, among others. While opinion mining is a well-established task with many
standard datasets and well-defined methodologies, emotion mining has received less attention due to its complexity. In particular,
the annotated gold standard… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/EmoEvent.Emo3D
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
Citation
@inproceedings{dehghani-etal-2025-emo3d,
title = "{E}mo3{D}: Metric and Benchmarking Dataset for 3{D} Facial Expression Generation from Emotion Description",
author = "Dehghani, Mahshid and
Shafiee, Amirahmad and
Shafiei, Ali and
Fallah, Neda and
Alizadeh, Farahmand and
Gholinejad, Mohammad Mehdi and
Behroozi, Hamid… See the full description on the dataset page: https://huggingface.co/datasets/llm-lab/Emo3D.Emotion_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this CSV file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Emotion_Video_Facial_Landmarks.ukr-emotions-per-annotator
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Per annotator: specifically this dataset contains concatenated… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-per-annotator.Telugu_EmotionDo cite the below reference for using the dataset:
@article{marreddy2022resource, title={Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP tasks in Telugu Language},
author={Marreddy, Mounika and Oota, Subba Reddy and Vakada, Lakshmi Sireesha and Chinni, Venkata Charan and Mamidi, Radhika},
journal={Transactions on Asian and Low-Resource Language Information Processing}, publisher={ACM New York, NY} }
If you want to use the four classes (angry… See the full description on the dataset page: https://huggingface.co/datasets/mounikaiiith/Telugu_Emotion.emotion-rawFrench-emotion
French Synthetic Emotion Dataset (7 Emotions)
To read this documentation in French, please scroll down.
Dataset Description
This dataset was created for French text classification and contains 14,000 sentences annotated with one of the seven basic emotions: colère (anger), dégoût (disgust), joie (joy), neutre (neutral), peur (fear), surprise (surprise), and tristesse (sadness).
The main feature of this dataset is that it was entirely synthetically generated through a… See the full description on the dataset page: https://huggingface.co/datasets/JusteLeo/French-emotion.Cultural-Emo
CulEmo
Cultural Lenses on Emotion (CuLEmo) is the first benchmark to evaluate culture-aware emotion prediction across six languages: Amharic, Arabic, English, German, Hindi, and Spanish.
It comprises 400 crafted questions per language, each requiring nuanced cultural reasoning and understanding. It is designed for evaluating LLMs in Sentiment analysis and emotion prediction.
If you use this dataset, cite the paper below.
BibTeX entry and citation info.… See the full description on the dataset page: https://huggingface.co/datasets/llm-for-emotion/Cultural-Emo.Simplified_Chinese_Multi-Emotion_Dialogue_Dataset
Simplified_Chinese_Multi-Emotion_Dialogue_Dataset
数据说明
本数据是简体中文口语情感分类数据集
翻译自于:Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset
使用Qwen2.5-32B-instruce模型将其从繁体中文翻译为简体中文
共4159条
情感类别
原始类别
翻译后类别
条数
悲傷語調
伤心
486
憤怒語調
生气
527
關切語調
关心
560
驚奇語調
惊讶
499
開心語調
开心
592
平淡語氣
平静
705
厭惡語調
厌恶
404
adcumen-viewer-emotions
AdCumen Viewer Emotions Dataset
Dataset for the paper "Decoding Viewer Emotions in Video Ads" by Alexey Antonov, Shravan Sampath Kumar, Jiefei Wei, William Headley, Orlando Wood, and Giovanni Montana, published in Nature Scientific Reports.
Code: github.com/gmontana/DecodingViewerEmotions
Model weights: dnamodel/tsam-viewer-emotions
Dataset Description
The dataset consists of 26,637 five-second video clips extracted from video advertisements, annotated for seven… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/adcumen-viewer-emotions.silly-emoji-qaThe silly dataset to take text questions and return emoji-only answers.
Powered by ChatGPT
Examples
Why do we have different seasons? 🌍☀️🔄
How do fish breathe underwater? 🐟💦💨
Import
# !pip install -q datasets
from datasets import load_dataset
train_ds, test_ds = load_dataset("hululuzhu/silly-emoji-qa", split=["train", "test"])
# import pandas as pd
# train_df = pd.DataFrame(train)
emotion-prediction-comet-atomic-2020
emotion-prediction-comet-atomic-2020
This dataset extends the COMET-Atomic-2020 commonsense reasoning dataset by focusing on the xReact (subject’s emotional reaction) and oReact (other person’s emotional reaction) relations.
Description
Each entry is expanded into a realistic, three‑sentence scenario, replacing the placeholders (PersonX/PersonY) with human names and adding contextual details.
Source (comet-atomic-2020) Example:
source
relation
target
PersonX… See the full description on the dataset page: https://huggingface.co/datasets/id4thomas/emotion-prediction-comet-atomic-2020.Arabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.
