datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.emotionsemotion-selfstory-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-selfstory-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it-postfix.microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.emotion-selfstory-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it probed on its OWN self-generated stories
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: abotresol/emotion-stories-gemma-4-31b-it — stories written by the probed model itself (generator = probed model, the reference's convention; 12 the twelve emotions, up to 256 stories each — the E6 scale corpus)
Per-story pooled residual-stream activations and per-emotion mean vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it.go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
eMotions
eMotions Dataset
The proposed eMotions dataset in our paper entitled Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline (ACM ICMR'25).
If you find our dataset useful, please cite our paper:
@inproceedings{wu2025towards,
title={Towards emotion analysis in short-form videos: A large-scale dataset and baseline},
author={Wu, Xuecheng and Sun, Heli and Xue, Junxiao and Nie, Jiayu and Kong, Xiangyan and Zhai, Ruofan and Huang, Danlei and He, Liang}… See the full description on the dataset page: https://huggingface.co/datasets/Conna/eMotions.synthetic-emotions
Synthetic Emotions Dataset
Overview
Synthetic Emotions is a video dataset of AI-generated human emotions created using OpenAI Sora. It features short (5-sec, 480p, 9:16) videos depicting diverse individuals expressing emotions like happiness, sadness, anger, fear, surprise, and more.
This dataset is ideal for emotion recognition, facial expression analysis, affective computing, and AI-human interaction research.
Dataset Details
Total Videos: 100
Video Format:… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/synthetic-emotions.many_emotions
Many Emotions
Many Emotions is a 2.7-million-row multilingual text-classification dataset for recognizing seven emotion categories in English, French, Italian, Spanish, and German. It combines examples from Emotion, DailyDialog, and GoEmotions.
The default unsplit corpus contains 2,710,740 non-empty rows derived from 550,123 source IDs. Every row records its source dataset and source-specific license.
The current release is 2.0.0. See CHANGELOG.md for changes from the original… See the full description on the dataset page: https://huggingface.co/datasets/ma2za/many_emotions.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.ru_go_emotions
Description
This dataset is a translation of the Google GoEmotions emotion classification dataset.
All features remain unchanged, except for the addition of a new ru_text column containing the translated text in Russian.
For the translation process, I used the Deep translator with the Google engine.
You can find all the details about translation, raw .csv files and other stuff in this Github repository.
For more information also check the official original dataset card.… See the full description on the dataset page: https://huggingface.co/datasets/seara/ru_go_emotions.emotions-dataset
🌟 Emotions Dataset — Infuse Your AI with Human Feelings! 😊😢😡
Tap into the Soul of Human Emotions 💖The Emotions Dataset is your key to unlocking emotional intelligence in AI. With 131,306 text entries labeled across 13 vivid emotions 😊😢😡, this dataset empowers you to build empathetic chatbots 🤖, mental health tools 🩺, social media analyzers 📱, and more!
The Emotions Dataset is a carefully curated collection designed to elevate emotion classification, sentiment… See the full description on the dataset page: https://huggingface.co/datasets/boltuix/emotions-dataset.go_emotions-en
GoEmotions dataset
The original dataset: GoEmotions (paper).
The derived dataset contains an additional labels_ekman column with emotion labels as per Paul Ekman's theory.
The original 27 + neutral emotion labels (may contain more than one label per sample):
0: admiration
1: amusement
2: anger
3: annoyance
4: approval
5: caring
6: confusion
7: curiosity
8: desire
9: disappointment
10: disapproval
11: disgust
12: embarrassment
13: excitement
14: fear
15: gratitude
16: grief
17: joy… See the full description on the dataset page: https://huggingface.co/datasets/AiLab-IMCS-UL/go_emotions-en.IndustryInstruction_Literature-Emotions
IndustryInstruction: Literature & Emotions
This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.team9-lepuppet_emotionsafrikaans-english-emotions-corpus
Afrikaans-english Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Afrikaans-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-english-emotions-corpus.AffectDF_EmotionSDD
AffectDF: Emotionally Expressive Speech Deepfake Benchmark
Overview
AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks.
AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.SemEval_training_data_emotions
Dataset Card for "SemEval_traindata_emotions"
Как был получен
from datasets import load_dataset
import datasets
from torchvision.io import read_video
import json
import torch
import os
from torch.utils.data import Dataset, DataLoader
import tqdm
dataset_path = "./SemEval-2024_Task3/training_data/Subtask_2_train.json"
dataset = json.loads(open(dataset_path).read())
print(len(dataset))
all_conversations = []
for item in dataset:
all_conversations.extend(item["conversation"])… See the full description on the dataset page: https://huggingface.co/datasets/dim/SemEval_training_data_emotions.emotion-dataset-20-emotions
20-Emotion Text Classification Dataset
A comprehensive dataset for fine-grained emotion classification containing 79,595 sentences labeled with 20 distinct emotions.
Dataset Description
This dataset is designed for training emotion classification models that can detect nuanced emotional states in text. Unlike basic sentiment analysis (positive/negative/neutral), this dataset provides fine-grained emotion labels that better capture the complexity of human emotions.… See the full description on the dataset page: https://huggingface.co/datasets/shreyaspullehf/emotion-dataset-20-emotions.go_emotions_ptbr
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English and Brazilian Portuguese… See the full description on the dataset page: https://huggingface.co/datasets/antoniomenezes/go_emotions_ptbr.Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.social-behavior-emotionstask518_emo_different_dialogue_emotions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task518_emo_different_dialogue_emotions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task518_emo_different_dialogue_emotions.ukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.urdu-emotions
URDU-Dataset
1. General information
URDU dataset contains emotional utterances of Urdu speech gathered from Urdu talk shows. It contains 300 utterances of four basic emotions: Angry, Happy, and Neutral. There are 38 speakers (27 male and 11 female). This data is created from YouTube. Speakers are selected randomly.
For more details about dataset, please refer the complete paper "Cross Lingual Speech Emotion Recognition: Urdu vs. Western Languages".… See the full description on the dataset page: https://huggingface.co/datasets/2DamnWav/urdu-emotions.twitter_emotions-en
Twitter Emotions dataset
The original dataset: Emotions
The derived dataset contains an additional labels_ekman column that maps the original emotion classes to the Paul Ekman's classification.
The column labels contains the emotion classes of the original dataset:
0: sadness
1: joy
2: love - not distinguished in the Ekman'sclassification
3: anger
4: fear
5: surprise
The column labels_ekman contains the corresponding Ekman's emotion classes:
0: anger
1: disgust - omitted, since not… See the full description on the dataset page: https://huggingface.co/datasets/AiLab-IMCS-UL/twitter_emotions-en.emotions-test
Test Version of humair025/TTS-Dataset-Batched
TTS-Dataset-Batched
Dataset Overview
TTS-Dataset-Batched is a large-scale, multi-speaker English text-to-speech dataset optimized for efficient processing and training. The Original dataset contains 556,667 high-quality audio samples across 30 unique speakers, totaling over 1,024 hours of speech data.
This is a batched version of a larger consolidated dataset, split into manageable chunks for easier downloading… See the full description on the dataset page: https://huggingface.co/datasets/humair025/emotions-test.go_emotions_raw
Dataset Card for go_emotions_raw
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
It contains the raw version of go_emotions as a FeedbackDataset. Each of the original questions are defined a single
FeedbackRecord and contain the responses from each annotator. The final labels in the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/go_emotions_raw.reachy2_emotions_library
