datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ru_emotion_dvach
Русский датасет эмоций (E-Kulture)
Датасет содержит тексты с разметкой по пяти эмоциональным категориям:
Агрессия (aggression)
Тревожность (anxiety)
Сарказм (sarcasm)
Позитив (positive)
Нейтральное состояние (neutral)
Структура датасета
Датасет разделен на два сплита:
train - обучающая выборка
valid - валидационная выборка
Каждый пример содержит следующие поля:
text - текст сообщения
label - метка эмоции (одна из пяти категорий)
Создание
Исходные данные… See the full description on the dataset page: https://huggingface.co/datasets/Kostya165/ru_emotion_dvach.Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.tweet_emotion_intensity
Tweet Emotion Intensity Dataset
Papers:
Emotion Intensities in Tweets. Saif M. Mohammad and Felipe Bravo-Marquez. In Proceedings of the sixth joint conference on lexical and computational semantics (*Sem), August 2017, Vancouver, Canada.
WASSA-2017 Shared Task on Emotion Intensity. Saif M. Mohammad and Felipe Bravo-Marquez. In Proceedings of the EMNLP 2017 Workshop on Computational Approaches to Subjectivity, Sentiment, and Social Media (WASSA), September 2017… See the full description on the dataset page: https://huggingface.co/datasets/stepp1/tweet_emotion_intensity.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.emotion-dataset-20-emotions
20-Emotion Text Classification Dataset
A comprehensive dataset for fine-grained emotion classification containing 79,595 sentences labeled with 20 distinct emotions.
Dataset Description
This dataset is designed for training emotion classification models that can detect nuanced emotional states in text. Unlike basic sentiment analysis (positive/negative/neutral), this dataset provides fine-grained emotion labels that better capture the complexity of human emotions.… See the full description on the dataset page: https://huggingface.co/datasets/shreyaspullehf/emotion-dataset-20-emotions.COVID-19_weibo_emotionCOVID-19 Epidemic Weibo Emotional Dataset, the content of Weibo in this dataset is the epidemic Weibo obtained by using relevant keywords to filter during the epidemic, and its content is related to COVID-19.
Each tweet is labeled as one of the following six categories: neutral (no emotion), happy (positive), angry (angry), sad (sad), fear (fear), surprise (surprise)
The COVID-19 Weibo training dataset includes 8,606 Weibos, the validation set contains 2,000 Weibos, and the test dataset… See the full description on the dataset page: https://huggingface.co/datasets/souljoy/COVID-19_weibo_emotion.text_emotionsocial-behavior-emotionsukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.Text-Emotion.
emotion-rawMoroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.Telugu_EmotionDo cite the below reference for using the dataset:
@article{marreddy2022resource, title={Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP tasks in Telugu Language},
author={Marreddy, Mounika and Oota, Subba Reddy and Vakada, Lakshmi Sireesha and Chinni, Venkata Charan and Mamidi, Radhika},
journal={Transactions on Asian and Low-Resource Language Information Processing}, publisher={ACM New York, NY} }
If you want to use the four classes (angry… See the full description on the dataset page: https://huggingface.co/datasets/mounikaiiith/Telugu_Emotion.go_emotionsEmotion_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this CSV file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Emotion_Video_Facial_Landmarks.ukr-emotions-per-annotator
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Per annotator: specifically this dataset contains concatenated… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-per-annotator.express-emotion-recognition
EXPRESS Dataset
Overview
This is the EXPRESS dataset from the paper:
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
EXPRESS (EXperiences and PRocessed Emotions in Self-disclosure Stories) is a benchmark dataset for evaluating fine-grained emotion recognition in language models. It contains 33,679 naturally occurring Reddit-based human experiences paired with self-disclosed emotion labels.
EXPRESS uses emotions explicitly disclosed by the… See the full description on the dataset page: https://huggingface.co/datasets/bangzhao/express-emotion-recognition.multilingual_go_emotions
Overview:
This dataset is updated from on the go_emotions dataset
With the same labels, but add 5 new languages: Arabic, French, Spanish , Dutch, and Turkish.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English Arabic, French, Spanish , Dutch, and Turkish
Simplified_Chinese_Multi-Emotion_Dialogue_Dataset
Simplified_Chinese_Multi-Emotion_Dialogue_Dataset
数据说明
本数据是简体中文口语情感分类数据集
翻译自于:Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset
使用Qwen2.5-32B-instruce模型将其从繁体中文翻译为简体中文
共4159条
情感类别
原始类别
翻译后类别
条数
悲傷語調
伤心
486
憤怒語調
生气
527
關切語調
关心
560
驚奇語調
惊讶
499
開心語調
开心
592
平淡語氣
平静
705
厭惡語調
厌恶
404
Cultural-Emo
CulEmo
Cultural Lenses on Emotion (CuLEmo) is the first benchmark to evaluate culture-aware emotion prediction across six languages: Amharic, Arabic, English, German, Hindi, and Spanish.
It comprises 400 crafted questions per language, each requiring nuanced cultural reasoning and understanding. It is designed for evaluating LLMs in Sentiment analysis and emotion prediction.
If you use this dataset, cite the paper below.
BibTeX entry and citation info.… See the full description on the dataset page: https://huggingface.co/datasets/llm-for-emotion/Cultural-Emo.adcumen-viewer-emotions
AdCumen Viewer Emotions Dataset
Dataset for the paper "Decoding Viewer Emotions in Video Ads" by Alexey Antonov, Shravan Sampath Kumar, Jiefei Wei, William Headley, Orlando Wood, and Giovanni Montana, published in Nature Scientific Reports.
Code: github.com/gmontana/DecodingViewerEmotions
Model weights: dnamodel/tsam-viewer-emotions
Dataset Description
The dataset consists of 26,637 five-second video clips extracted from video advertisements, annotated for seven… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/adcumen-viewer-emotions.Yemeni-Speech-Emotion-Dataset
YSED — Yemeni Speech Emotion Dataset (audio-classification repackaging)
A clean repackaging of YSED with a metadata.csv and stratified train/validation/test splits, for emotion classification on Yemeni Arabic.
Original dataset: Derhem, S., AL-Mekhlafi, E., AL-Majmar, N. A., & AL-Makhlafi, M. (2025). YSED: Yemeni Speech Emotion Dataset. Data in Brief. DOI: 10.1016/j.dib.2025.112233. Zenodo: https://zenodo.org/records/15227219.
What's in here
1432 audio clips across… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Yemeni-Speech-Emotion-Dataset.Arabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.emotion一个包含六种基本情绪(愤怒、恐惧、喜悦、爱、悲伤和惊讶)的英文Twitter消息数据集
Github 链接 https://github.com/dair-ai/emotion_dataset
emotion-prediction-comet-atomic-2020
emotion-prediction-comet-atomic-2020
This dataset extends the COMET-Atomic-2020 commonsense reasoning dataset by focusing on the xReact (subject’s emotional reaction) and oReact (other person’s emotional reaction) relations.
Description
Each entry is expanded into a realistic, three‑sentence scenario, replacing the placeholders (PersonX/PersonY) with human names and adding contextual details.
Source (comet-atomic-2020) Example:
source
relation
target
PersonX… See the full description on the dataset page: https://huggingface.co/datasets/id4thomas/emotion-prediction-comet-atomic-2020.korean-emotion-lexicon
Korean Emotion Lexicon
This repository contains a comprehensive dataset of Korean emotion lexicons developed through psychological research conducted by In-jo Park and Kyung-Hwan Min from Seoul National University. The dataset includes several key measures for each emotion lexicon:
lexicon: The lexicon that represents a specific emotion in the Korean language.
representative: The degree to which the lexicon is a representative example of the emotion.
prototypicality: A rating of… See the full description on the dataset page: https://huggingface.co/datasets/jonghwanhyeon/korean-emotion-lexicon.short-text-multi-labeled-emotion-classificationjournal-entries-emotion-detection-vadReddit Diary of a Redditor VAD Dataset
Dataset Creation Process
Scraping Reddit Posts
Posts were scraped from the r/diaryofaredditor subreddit using the Reddit API.
The script used for scraping is shown below:import requests
import csv
import time
access_token = ""
headers = {
"Authorization": f"bearer {access_token}",
"User-Agent": "ChangeMeClient/0.1"
}
url = "https://oauth.reddit.com/r/diaryofaredditor/new"
params = {"limit": 100}
after = None
csv_path =… See the full description on the dataset page: https://huggingface.co/datasets/mmarkusmalone/journal-entries-emotion-detection-vad.synthetic-emotion-detection-dataset-v1
Tanaos Emotion Detection Training Dataset
This dataset was created synthetically by Tanaos with the Artifex Python library.
The dataset is designed to train and evaluate emotion detection systems — models that classify the main emotion expressed in text as one of eight possible categories: joy, anger, fear, sadness, surprise, disgust, excitement, or neutral. It can be used to build emotion detection models for various applications, such as customer feedback analysis, social… See the full description on the dataset page: https://huggingface.co/datasets/tanaos/synthetic-emotion-detection-dataset-v1.
