datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.emotionsBRIGHTER-emotion-categories
BRIGHTER Emotion Categories Dataset
This dataset contains the emotion categories data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages.
Dataset Description
The BRIGHTER Emotion Categories dataset is a comprehensive multi-language, multi-label emotion classification dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
laion-emotional-trajectory-t80
LAION Emotional-Trajectory Speech — tier T≥0.80
319,765 crossfaded speech trajectories · 4,482 audio-hours · 1,598,825 source clips
A trajectory is a short sequence of 5 consecutive utterances by one
speaker whose measured emotion or voice character moves monotonically from one end of
the corpus distribution to the other. The clips are joined into one continuous audio file
with equal-power crossfades, the joined audio is re-tokenized with MOSS-Audio-
Tokenizer-v2, and every… See the full description on the dataset page: https://huggingface.co/datasets/laion/laion-emotional-trajectory-t80.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.ru_go_emotions
Description
This dataset is a translation of the Google GoEmotions emotion classification dataset.
All features remain unchanged, except for the addition of a new ru_text column containing the translated text in Russian.
For the translation process, I used the Deep translator with the Google engine.
You can find all the details about translation, raw .csv files and other stuff in this Github repository.
For more information also check the official original dataset card.… See the full description on the dataset page: https://huggingface.co/datasets/seara/ru_go_emotions.BRIGHTER-emotion-intensities
BRIGHTER Emotion Intensities Dataset
This dataset contains the emotion intensities data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages.
Dataset Description
The BRIGHTER Emotion Intensities dataset is a comprehensive multi-language emotion intensity dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple languages, providing… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-intensities.EmotionAnalysisFinal
Dataset Card for EmotionAnalysisFinal
EmotionAnalysisFinal is the official dataset for SemEval-2025 Task 11, Track C: Cross-lingual Emotion Detection in Social Media Text.
This dataset comprises multilingual social media posts annotated for six basic emotions: anger, disgust, fear, joy, sadness, and surprise.
The annotation schema is multi-label.
Each language-specific configuration (subset) contains validation and test splits.
Split
Original Source
Notes
validation
dev… See the full description on the dataset page: https://huggingface.co/datasets/llama-lang-adapt/EmotionAnalysisFinal.IndustryInstruction_Literature-Emotions
IndustryInstruction: Literature & Emotions
This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.tweet_emotion_intensity
Tweet Emotion Intensity Dataset
Papers:
Emotion Intensities in Tweets. Saif M. Mohammad and Felipe Bravo-Marquez. In Proceedings of the sixth joint conference on lexical and computational semantics (*Sem), August 2017, Vancouver, Canada.
WASSA-2017 Shared Task on Emotion Intensity. Saif M. Mohammad and Felipe Bravo-Marquez. In Proceedings of the EMNLP 2017 Workshop on Computational Approaches to Subjectivity, Sentiment, and Social Media (WASSA), September 2017… See the full description on the dataset page: https://huggingface.co/datasets/stepp1/tweet_emotion_intensity.team9-lepuppet_emotionsemotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.emotional_application
Data explorer and full leaderboard
https://huggingface.co/spaces/llm-council/emotional-intelligence-arena
The LMC-EA dataset
This dataset was developed to demonstrate how to benchmark foundation models on highly subjective tasks such as those in the domain of emotional intelligence by the collective consensus of a council of LLMs.
There are 4 subsets of the LMC-EA dataset:
test_set_formulation: Synthetic expansions of the EmoBench EA dataset, generated by 20… See the full description on the dataset page: https://huggingface.co/datasets/llm-council/emotional_application.audio2face-emotion-arkit-teacher
audio2face-emotion-arkit-teacher
Nyx avatar (Gaussian-splat head, ARKit-52 blendshape rig) driven by a surprise clip's blendshape labels derived from this dataset.
14,082 emotional-speech clips, each annotated with two parallel 52-channel ARKit blendshape sequences (NVIDIA Audio2Face-3D-v2.3.1-James and LAM_Audio2Expression) plus a 26-dimensional NVIDIA Audio2Emotion conditioning vector.
Reference-only dataset — the original audio is not shipped. Each row contains a… See the full description on the dataset page: https://huggingface.co/datasets/myned-ai/audio2face-emotion-arkit-teacher.ukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.emotion-analysis-8-categories-datasetMoroccan-Arabic-Multimodal-Emotion-Recognition
MDER-MA — Moroccan Arabic Multimodal Emotion Recognition (TTS-aligned repackaging)
A repackaging of the MDER-MA dataset that pairs every audio clip with its Arabic (Moroccan dialect / Darija) transcript and ships speaker-disjoint train/validation/test splits.
Original dataset: Ouali, S. & El Garouani, S. (2025). MDER-MA: A multimodal dataset for emotion recognition in low-resource Moroccan Arabic language. Data in Brief. DOI: 10.1016/j.dib.2025.112005. Mendeley:… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Moroccan-Arabic-Multimodal-Emotion-Recognition.Emotion_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this CSV file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Emotion_Video_Facial_Landmarks.emotion-datasets
Emotion datasets
Synthetic emotion text re-generated from the data pipelines of Emotion concepts and their function in a LLM
(paper), for interpretability and steering research. This is a re-generation with a different model, not the paper
authors' data; prompts, the 171-emotion word list and the 100 story topics come from the paper's appendix.
Total: 4,061 rows across 4 configs.
config
rows
what it is
stories
2,718
one story per row, one target emotion each (12… See the full description on the dataset page: https://huggingface.co/datasets/knoveleng/emotion-datasets.short-text-multi-labeled-emotion-classificationUnified_Dataset_with_Emotionsexpress-emotion-recognition
EXPRESS Dataset
Overview
This is the EXPRESS dataset from the paper:
Fluent but Unfeeling: The Emotional Blind Spots of Language Models
EXPRESS (EXperiences and PRocessed Emotions in Self-disclosure Stories) is a benchmark dataset for evaluating fine-grained emotion recognition in language models. It contains 33,679 naturally occurring Reddit-based human experiences paired with self-disclosed emotion labels.
EXPRESS uses emotions explicitly disclosed by the… See the full description on the dataset page: https://huggingface.co/datasets/bangzhao/express-emotion-recognition.ukr-emotions-per-annotator
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Per annotator: specifically this dataset contains concatenated… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-per-annotator.MELD-processed-v5-emotion2vec
MELD Processed Multi-Modal Emotion Recognition Dataset
Processed dataset containing Prosody, Whisper acoustic encodings, DistilBERT text hidden states, and Ekman emotion labels.
adcumen-viewer-emotions
AdCumen Viewer Emotions Dataset
Dataset for the paper "Decoding Viewer Emotions in Video Ads" by Alexey Antonov, Shravan Sampath Kumar, Jiefei Wei, William Headley, Orlando Wood, and Giovanni Montana, published in Nature Scientific Reports.
Code: github.com/gmontana/DecodingViewerEmotions
Model weights: dnamodel/tsam-viewer-emotions
Dataset Description
The dataset consists of 26,637 five-second video clips extracted from video advertisements, annotated for seven… See the full description on the dataset page: https://huggingface.co/datasets/dnamodel/adcumen-viewer-emotions.eeg-bme-emotion-parquet
BME EEG Emotion Parquet
赛题四脑电情绪识别数据的 Parquet 转换版本。数据按 HuggingFace Datasets 原生 split
组织:
train: 384 条样本
validation: 96 条样本
test: 80 条样本
每条样本对应一段视频的 EEG 数据。eeg 字段为扁平数组,可通过 eeg_shape 和
eeg_dtype 还原为二维数组。
from datasets import load_dataset
import numpy as np
ds = load_dataset("MigoXV/eeg-bme-emotion-parquet")
row = ds["train"][0]
eeg = np.asarray(row["eeg"], dtype=row["eeg_dtype"]).reshape(row["eeg_shape"])
训练集和验证集来自原训练数据,按被试级别拆分,避免同一被试同时出现在训练集和验
证集。测试集来自公开测试集,真实情绪标签未公开。tsterbak-lyrics-dataset-with-emotionsArabic-Emotional-Audio-Dataset-Baved
BAVED — Basic Arabic Vocal Emotions Dataset (TTS-ready repackaging)
A re-packaged, transcript-aligned version of the Basic Arabic Vocal Emotions Dataset (BAVED) with explicit Arabic transcripts, English glosses, speaker metadata, and speaker-disjoint train/validation/test splits.
Original dataset: Aouf Yacine, Basic Arabic Vocal Emotions Dataset (BAVED), GitHub: https://github.com/40uf411/Basic-Arabic-Vocal-Emotions-Dataset. This repackaging adds metadata; all audio is unchanged.… See the full description on the dataset page: https://huggingface.co/datasets/FatimahEmadEldin/Arabic-Emotional-Audio-Dataset-Baved.
