datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emotion
Dataset Card for "emotion"
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
An example looks as follows.
{
"text": "im feeling quite sad and sorry for myself but… See the full description on the dataset page: https://huggingface.co/datasets/dair-ai/emotion.emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)**
go_emotions
Dataset Card for GoEmotions
Dataset Summary
The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral.
The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test
splits.
Supported Tasks and Leaderboards
This dataset is intended for multi-class, multi-label emotion classification.
Languages
The data is in English.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.reachy-mini-emotions-library
Reachy Mini Emotions Library
Curated emotion recordings for the Reachy Mini robot, maintained by
Pollen Robotics. Each move is a JSON trajectory (head pose, antennas,
body yaw, sampled over time) paired with an Opus audio track.
Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by
the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves
non-.wav audio sidecars).
File layout
Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.Emotional_SpeechThis dataset contains audio-text pairs in the webdataset format.
The audio files are short speech segments from publicly available videos & the texts are descriptions of emotions the speakers seems to be feeling. Some captions also describe the speakers gender and age.
All files with the substring "part1" in the name contain unique audio files with unique captions.
All files with the substring "part2" , "part3", ... in the name contain the same audio files as in "part1", but with different… See the full description on the dataset page: https://huggingface.co/datasets/EQ4You/Emotional_Speech.emotion-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.Emotion_new_collected_datasetemotion
EmotionClassification
An MTEB dataset
Massive Text Embedding Benchmark
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise.
Task category
t2c
Domains
Social, Written
Reference
https://www.aclweb.org/anthology/D18-1404
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_tasks(["EmotionClassification"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/emotion.emotionsRobot-EQ
RobotEQ-Data
Official dataset release for RobotEQ.
Evaluation & Scripts
For inference scripts, evaluation scripts, and data production tooling, see the RobotEQ code repository.
Dataset Statistics
Item
Count
Behavior judgment scenarios (synthetic)
1,812
Behavior judgment scenarios (real POV)
223
Behavior judgment scenarios (total)
2,035
Behavior judgment behavior annotations
3,171
Spatial grounding questions
825… See the full description on the dataset page: https://huggingface.co/datasets/Tongji-Emotion/Robot-EQ.laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset
Overview
"LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants.
Updated Composition
Voices and Languages
English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.Emilia-with-Emotion-Annotations
Dataset Card for Emilia with Emotion Annotations
Dataset Description
This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?".
The annotations include 54 scores for each sample, covering a wide range… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emilia-with-Emotion-Annotations.emotion-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.emotion-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus
Data provenance (what made these activations)
Probed model (whose activations these are): google/gemma-4-31b (base)
Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model)
Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.BRIGHTER-emotion-categories
BRIGHTER Emotion Categories Dataset
This dataset contains the emotion categories data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages.
Dataset Description
The BRIGHTER Emotion Categories dataset is a comprehensive multi-language, multi-label emotion classification dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories.emotion-vectors-gemma-4-31b-postfix
Emotion vectors, google/gemma-4-31b (corrected extraction)
Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.dialogs-ru-emotional-conversations
Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus
Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational
Russian speech, designed for dialog-oriented and emotional text-to-speech.
Unlike existing Russian corpora — mostly single-speaker read speech or large but
low-quality web-mined audio — Dialogs was recorded by professional theatre actors
performing scripted dialogs face-to-face, capturing natural turn-taking,
timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.emotion6-testemotion-selfstory-vectors-gemma-4-31b-it-postfix
Emotion vectors, google/gemma-4-31b-it (corrected extraction)
Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per
emotion. Each emotion ends up as one direction in the model's activation space.
Read LINEAGE.md before using this. This set supersedes
abotresol/emotion-selfstory-vectors-gemma-4-31b-it. The earlier extraction ran
while the tokenizer padded on the left, so the step that skips a story's first
50 tokens skipped padding instead. This… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it-postfix.emotion-combined-trajectories-gemma-4-31b-it-v2
Per-token emotion trajectories, instruction-tuned model (primary set)
google/gemma-4-31b-it read over stories written to move through three emotions in sequence, scored against three different emotion-vector sets (corpus-built, self-generated and DeepSeek-written). This is the primary trajectory set behind the project's story-following results.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b-it-v2.emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it
Per-token emotion trajectories, DeepSeek-written stories
google/gemma-4-31b-it read over the DeepSeek-written three-emotion stories. Pairs with the Gemma-written set to separate what the model does from what the story writer does.
Each story is stored as one .npz. The arrays are per token, so a trajectory
can be replayed word by word rather than only summarised.
Contents
Path
Contents
shards/<story_id>.npz
one story, arrays below
manifest.jsonl
one… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-deepseek-stories-gemma-4-31b-it.emotion-probes-raw-activationsEmilia-with-Emotion-Annotations4microduck-emotions
Microduck Emotions
A collection of emotions for the Microduck robot. Each one is a motion and a sound designed together, beat by
beat, with the beak opening on the sound, rendered in the physics simulation and validated on the real robot. Every
emotion is three files: the motion (emotions/<name>.json, keyframes at 30 fps: head and body offsets played on
top of whichever trained policy is active, plus the policy hand-overs, such as the sit that devastated and play dead
start)… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/microduck-emotions.emotion-selfstory-vectors-gemma-4-31b-it
Emotion vectors — gemma-4-31b-it probed on its OWN self-generated stories
Data provenance (what made these activations)
Probed model: google/gemma-4-31b-it (instruct)
Input corpus: abotresol/emotion-stories-gemma-4-31b-it — stories written by the probed model itself (generator = probed model, the reference's convention; 12 the twelve emotions, up to 256 stories each — the E6 scale corpus)
Per-story pooled residual-stream activations and per-emotion mean vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-selfstory-vectors-gemma-4-31b-it.IndustryCorpus2_literature_emotion
IndustryCorpus2: Literature & Emotions
This repository contains the IndustryCorpus2: Literature & Emotions domain subset of BAAI/IndustryCorpus2.
Refer to the parent dataset card for data construction, intended use, limitations,
and licensing details.
Citation
If you use this dataset in your work, please cite IndustryCorpus2:
@misc{shi2024industrycorpus2,
title = {IndustryCorpus2},
author = {Xiaofeng Shi and Lulu Zhao and Hua Zhou and Donglin Hao}… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus2_literature_emotion.emotion-dialogue-vectors-gemma-4-31b
Emotion vectors — gemma-4-31b (base) probed on base-generated dialogues
Data provenance (what made these activations)
Probed model: google/gemma-4-31b (base)
Input corpus: abotresol/emotion-dialogues-gemma-4-31b — two-person dialogues written by the base model (generator = probed model; 44% emotion-word leakage, documented)
Per-story pooled residual-stream activations and per-emotion mean vectors,
extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-dialogue-vectors-gemma-4-31b.Emilia-with-Emotion-Annotations5go_emotions
GoEmotions
This dataset is a port of the official go_emotions dataset on the Hub. It only contains the simplified subset as these are the only fields we need for text classification.
emotion-combined-trajectories-gemma-4-31b
emotion-combined-trajectories-gemma-4-31b
BASE-model condition of the per-token trajectory stories (companion to
emotion-combined-trajectories-gemma-4-31b-it, same corpus and layout). One
shard per story from a teacher-forced forward pass through google/gemma-4-31b
(base, bf16, right padding), residual captured at layers [6, 15, 24, 33, 42, 51].
Provenance caveats specific to this condition
The stories were generated by the INSTRUCT model (base cannot chat-write… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-combined-trajectories-gemma-4-31b.
