datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.emotionsmicroduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.eMotions
eMotions Dataset
The proposed eMotions dataset in our paper entitled Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline (ACM ICMR'25).
If you find our dataset useful, please cite our paper:
@inproceedings{wu2025towards,
title={Towards emotion analysis in short-form videos: A large-scale dataset and baseline},
author={Wu, Xuecheng and Sun, Heli and Xue, Junxiao and Nie, Jiayu and Kong, Xiangyan and Zhai, Ruofan and Huang, Danlei and He, Liang}… See the full description on the dataset page: https://huggingface.co/datasets/Conna/eMotions.go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
synthetic-emotions
Synthetic Emotions Dataset
Overview
Synthetic Emotions is a video dataset of AI-generated human emotions created using OpenAI Sora. It features short (5-sec, 480p, 9:16) videos depicting diverse individuals expressing emotions like happiness, sadness, anger, fear, surprise, and more.
This dataset is ideal for emotion recognition, facial expression analysis, affective computing, and AI-human interaction research.
Dataset Details
Total Videos: 100
Video Format:… See the full description on the dataset page: https://huggingface.co/datasets/aadityaubhat/synthetic-emotions.many_emotions
Many Emotions
Many Emotions is a 2.7-million-row multilingual text-classification dataset for recognizing seven emotion categories in English, French, Italian, Spanish, and German. It combines examples from Emotion, DailyDialog, and GoEmotions.
The default unsplit corpus contains 2,710,740 non-empty rows derived from 550,123 source IDs. Every row records its source dataset and source-specific license.
The current release is 2.0.0. See CHANGELOG.md for changes from the original… See the full description on the dataset page: https://huggingface.co/datasets/ma2za/many_emotions.ukr-emotions-binary
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Binary: specifically this dataset contains binary labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-binary.ru_go_emotions
Description
This dataset is a translation of the Google GoEmotions emotion classification dataset.
All features remain unchanged, except for the addition of a new ru_text column containing the translated text in Russian.
For the translation process, I used the Deep translator with the Google engine.
You can find all the details about translation, raw .csv files and other stuff in this Github repository.
For more information also check the official original dataset card.… See the full description on the dataset page: https://huggingface.co/datasets/seara/ru_go_emotions.emotions-dataset
🌟 Emotions Dataset — Infuse Your AI with Human Feelings! 😊😢😡
Tap into the Soul of Human Emotions 💖The Emotions Dataset is your key to unlocking emotional intelligence in AI. With 131,306 text entries labeled across 13 vivid emotions 😊😢😡, this dataset empowers you to build empathetic chatbots 🤖, mental health tools 🩺, social media analyzers 📱, and more!
The Emotions Dataset is a carefully curated collection designed to elevate emotion classification, sentiment… See the full description on the dataset page: https://huggingface.co/datasets/boltuix/emotions-dataset.afrikaans-english-emotions-corpus
Afrikaans-english Emotion Analysis Corpus
Dataset Description
This dataset contains emotion-labeled text data in Afrikaans-english for emotion classification (joy, sadness, anger, fear, surprise, disgust, neutral). Emotions were extracted and processed from the English meanings of the sentences using the model j-hartmann/emotion-english-distilroberta-base. The dataset is part of a larger collection of African language emotion analysis resources.
Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-english-emotions-corpus.go_emotions-en
GoEmotions dataset
The original dataset: GoEmotions (paper).
The derived dataset contains an additional labels_ekman column with emotion labels as per Paul Ekman's theory.
The original 27 + neutral emotion labels (may contain more than one label per sample):
0: admiration
1: amusement
2: anger
3: annoyance
4: approval
5: caring
6: confusion
7: curiosity
8: desire
9: disappointment
10: disapproval
11: disgust
12: embarrassment
13: excitement
14: fear
15: gratitude
16: grief
17: joy… See the full description on the dataset page: https://huggingface.co/datasets/AiLab-IMCS-UL/go_emotions-en.AffectDF_EmotionSDD
AffectDF: Emotionally Expressive Speech Deepfake Benchmark
Overview
AffectDF is a large-scale benchmark for speech deepfake detection under emotionally expressive spoofing conditions. The dataset is designed to evaluate whether current speech deepfake detection (SDD) systems can generalize beyond conventional neutral-speech benchmarks to modern emotional and expressive speech attacks.
AffectDF contains approximately 260 hours of audio generated using 21 spoofing… See the full description on the dataset page: https://huggingface.co/datasets/AffectDF/AffectDF_EmotionSDD.IndustryInstruction_Literature-Emotions
IndustryInstruction: Literature & Emotions
This repository contains the IndustryInstruction: Literature & Emotions domain subset of BAAI/IndustryInstruction.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryInstruction:
@misc{shi2024industryinstruction,
title = {IndustryInstruction},
author = {Xiaofeng Shi and Lulu Zhao and Hua… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryInstruction_Literature-Emotions.emotion-dataset-20-emotions
20-Emotion Text Classification Dataset
A comprehensive dataset for fine-grained emotion classification containing 79,595 sentences labeled with 20 distinct emotions.
Dataset Description
This dataset is designed for training emotion classification models that can detect nuanced emotional states in text. Unlike basic sentiment analysis (positive/negative/neutral), this dataset provides fine-grained emotion labels that better capture the complexity of human emotions.… See the full description on the dataset page: https://huggingface.co/datasets/shreyaspullehf/emotion-dataset-20-emotions.task518_emo_different_dialogue_emotions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task518_emo_different_dialogue_emotions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task518_emo_different_dialogue_emotions.SemEval_training_data_emotions
Dataset Card for "SemEval_traindata_emotions"
Как был получен
from datasets import load_dataset
import datasets
from torchvision.io import read_video
import json
import torch
import os
from torch.utils.data import Dataset, DataLoader
import tqdm
dataset_path = "./SemEval-2024_Task3/training_data/Subtask_2_train.json"
dataset = json.loads(open(dataset_path).read())
print(len(dataset))
all_conversations = []
for item in dataset:
all_conversations.extend(item["conversation"])… See the full description on the dataset page: https://huggingface.co/datasets/dim/SemEval_training_data_emotions.social-behavior-emotionsukr-emotions-intensity
EmoBench-UA: Emotions Detection Dataset in Ukrainian Texts
EmoBench-UA: the first of its kind emotions detection dataset in Ukrainian texts. This dataset covers the detection of basic emotions: Joy, Anger, Fear, Disgust, Surprise, Sadness, or None.
Any text can contain any amount of emotion -- only one, several, or none at all. The texts with None emotions are the ones where the labels per emotions classes are 0.
Intensity: specifically this dataset contains intensity labels… See the full description on the dataset page: https://huggingface.co/datasets/ukr-detect/ukr-emotions-intensity.go_emotions-es-mt
GoEmotions Spanish
A Spanish translation (using EasyNMT) of the GoEmotions dataset.
For more information check the official Model Card
talromur3_without_emotions
Overview
This corpus is an emotion-less version of Talromur3_with_prompts. talromur_3_without_emotions is a prompt-labelled corpus that can be used for fine-tuning models, such as ParlerTTS.The corpus consists of approximately 15,000 utterances, spoken by 7 named speakers in 6 different emotions (see more info here).
The dataset is an expanded version of Talromur-3: an Icelandic emotional speech corpus.We have added natural-language descriptions of utterance-level pitch, speech… See the full description on the dataset page: https://huggingface.co/datasets/atlithor/talromur3_without_emotions.go_emotions_raw
Dataset Card for go_emotions_raw
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
This dataset contains:
A dataset configuration file conforming to the Argilla dataset format named argilla.yaml. This configuration file will be used to configure the dataset when using the… See the full description on the dataset page: https://huggingface.co/datasets/plaguss/go_emotions_raw.emotions-test
Test Version of humair025/TTS-Dataset-Batched
TTS-Dataset-Batched
Dataset Overview
TTS-Dataset-Batched is a large-scale, multi-speaker English text-to-speech dataset optimized for efficient processing and training. The Original dataset contains 556,667 high-quality audio samples across 30 unique speakers, totaling over 1,024 hours of speech data.
This is a batched version of a larger consolidated dataset, split into manageable chunks for easier downloading… See the full description on the dataset page: https://huggingface.co/datasets/humair025/emotions-test.go_emotions_raw
Dataset Card for go_emotions_raw
This dataset has been created with Argilla.
As shown in the sections below, this dataset can be loaded into Argilla as explained in Load with Argilla, or used directly with the datasets library in Load with datasets.
Dataset Summary
It contains the raw version of go_emotions as a FeedbackDataset. Each of the original questions are defined a single
FeedbackRecord and contain the responses from each annotator. The final labels in the… See the full description on the dataset page: https://huggingface.co/datasets/argilla/go_emotions_raw.go_emotionslv_go_emotionsOriginal dataset: GoEmotions dataset
The dataset was machine translated to Latvian using free Google Translate API.
Tool used for translation: deep-translator
Translation script:
from datasets import load_dataset
from deep_translator import GoogleTranslator
from deep_translator.exceptions import TranslationNotFound
original_dataset = load_dataset("go_emotions", name="simplified")
translator = GoogleTranslator(source="en", target="lv")
def translate_batch(batch):
original_text =… See the full description on the dataset page: https://huggingface.co/datasets/SkyWater21/lv_go_emotions.nepali-ethnic-groups-Face-emotionstiny-stories-emotionsmultilingual_go_emotions
Overview:
This dataset is updated from on the go_emotions dataset
With the same labels, but add 5 new languages: Arabic, French, Spanish , Dutch, and Turkish.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English Arabic, French, Spanish , Dutch, and Turkish
