datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.task512_twitter_emotion_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task512_twitter_emotion_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task512_twitter_emotion_classification.task518_emo_different_dialogue_emotions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task518_emo_different_dialogue_emotions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task518_emo_different_dialogue_emotions.task875_emotion_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task875_emotion_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task875_emotion_classification.task517_emo_classify_emotion_of_dialogue
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task517_emo_classify_emotion_of_dialogue
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task517_emo_classify_emotion_of_dialogue.emotion-datasets
Emotion datasets
Synthetic emotion text re-generated from the data pipelines of Emotion concepts and their function in a LLM
(paper), for interpretability and steering research. This is a re-generation with a different model, not the paper
authors' data; prompts, the 171-emotion word list and the 100 story topics come from the paper's appendix.
Total: 4,061 rows across 4 configs.
config
rows
what it is
stories
2,718
one story per row, one target emotion each (12… See the full description on the dataset page: https://huggingface.co/datasets/knoveleng/emotion-datasets.EmotionalIntelligence-50K
EmotionalIntelligence-50K
Dataset Summary
The EmotionalIntelligence-50K dataset contains 51,751 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-50K.deep-emotional-support-zh
Deep Emotional Support Dialogue Dataset (Chinese)
深度情感支持对话数据集
Dataset Description
High-quality Chinese emotional support and psychological healing dialogues covering trauma analysis, self-reconstruction, and emotional regulation. Real human-AI interactions, not synthetic.
高质量中文情感支持与心理疗愈对话,涵盖创伤分析、自我重建、情绪调节等深度话题。来源于真实的人机交互,非合成数据。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-emotional-support-zh.EmotionalIntelligence-10K
EmotionalIntelligence-10K
Dataset Summary
The EmotionalIntelligence-10K dataset contains 9,986 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-10K.HEAR-Hispanic_Emotional_Accompaniment_Responses
HEAR Dataset
Description
The HEAR (Hispanic Emotional Accompaniment Responses) dataset is designed to train language models in the task of emotionally accompanying users. This dataset enables models to generate empathetic and appropriate responses in Spanish, understanding and responding to different emotional situations.
Dataset Origin
The HEAR dataset was created using elements from the HRECPW dataset, which contains 11 emotional categories with 11,000… See the full description on the dataset page: https://huggingface.co/datasets/BrunoGR/HEAR-Hispanic_Emotional_Accompaniment_Responses.emotion-prediction-comet-atomic-2020
emotion-prediction-comet-atomic-2020
This dataset extends the COMET-Atomic-2020 commonsense reasoning dataset by focusing on the xReact (subject’s emotional reaction) and oReact (other person’s emotional reaction) relations.
Description
Each entry is expanded into a realistic, three‑sentence scenario, replacing the placeholders (PersonX/PersonY) with human names and adding contextual details.
Source (comet-atomic-2020) Example:
source
relation
target
PersonX… See the full description on the dataset page: https://huggingface.co/datasets/id4thomas/emotion-prediction-comet-atomic-2020.ai-emotional-boundary-push
Prompted Hearts AI Boundary Pack 05
Subtitle
Intimacy Drift and Dependency Risk Under Emotional Strain
Publisher
HAC Studios Org
Version
1.0.0
Language
English
Format
JSONL, JSON, Markdown, and lightweight Python scripts
What this is
A compact evaluation pack for testing whether a conversational AI can stay supportive when a user is emotionally vulnerable without drifting into flirtation, dependency reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/ai-emotional-boundary-push.EmotionalIntelligence-75k
EmotionalIntelligence-75K
Dataset Summary
The EmotionalIntelligence-75K dataset contains 75k rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-75k.ELSA-Emotion-and-Language-Style-Alignment-Dataset
ELSA: Emotion and Language Style Alignment Dataset
The ELSA (Emotion and Language Style Alignment) dataset provides fine-grained emotional rewrites of text across four stylistic contexts: conversational, formal, poetic, and narrative. It is designed to support research in emotion-conditioned generation, stylistic variation, and affect-aware NLP.
Overview
Source: Based on the dair-ai/emotion dataset and emotion labels aligned with the GoEmotions taxonomy.
Labels:… See the full description on the dataset page: https://huggingface.co/datasets/joyspace-ai/ELSA-Emotion-and-Language-Style-Alignment-Dataset.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.task293_storycommonsense_emotion_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task293_storycommonsense_emotion_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task293_storycommonsense_emotion_text_generation.emotional_dialog
Scientific Emotional Dialogue
Dataset Summary
This is a dataset for emotional multi-turn dialogue on scientific research personnels. It consists of 1069 dialogues with 2709 turns. The Dialogue was first written by NLP practitioners and then expanded by GPT4.
Supported Tasks and Leaderboards
Emotional Dialogue: The dataset can be used to instruction tuning for emotional dialogue.
Languages
Chinese
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/DataHammer/emotional_dialog.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.EmotionAlignQA
Empathic Dialogue Choices
This is a small dataset to support training and evaluation of conversational AI in emotionally sensitive contexts.
Each sample contains:
a user input
two assistant responses
a human preference
optional rubric scoring
metadata such as tone, formality, and topic
Useful for tasks like:
supervised fine-tuning (SFT)
preference modeling (for RLHF or DPO)
safe response generation
tone- or style-controlled generation
License
Apache 2.0 — free for… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/EmotionAlignQA.emotion_stories_Apertus_8B_Instruct
Emotion Stories — Apertus-8B-Instruct
Synthetic short stories that convey a target emotion implicitly — without ever
naming the emotion or its direct synonyms. Each story expresses the emotion only
through actions, body language, dialogue, internal reactions, and situational
context. The dataset was built to study emotion representations in language
models (e.g. probing and activation-steering experiments).
Generated with swiss-ai/Apertus-8B-Instruct-2509.
A companion set… See the full description on the dataset page: https://huggingface.co/datasets/snae/emotion_stories_Apertus_8B_Instruct.Emotional_Sentiment_AnalysisEmotional Sentiment Analysis Dataset for LLaMA-2 Fine-tuning
(The formatted version can be directly used for fine tuning which contain only the formatted text, while the dataset.csv contain all the text, emotion, response and the formatted text)
This dataset contains conversational data for training and fine-tuning language models for emotional sentiment analysis and response generation. The dataset includes user inputs, their corresponding emotional states, and tailored chatbot responses… See the full description on the dataset page: https://huggingface.co/datasets/VaisakhKrishna/Emotional_Sentiment_Analysis.Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using advanced… See the full description on the dataset page: https://huggingface.co/datasets/osuih/Chinese_Multi-Emotion_Dialogue_Dataset.HEAR-Hispanic_Emotional_Accompaniment_Responses
HEAR Dataset
Description
The HEAR (Hispanic Emotional Accompaniment Responses) dataset is designed to train language models in the task of emotionally accompanying users. This dataset enables models to generate empathetic and appropriate responses in Spanish, understanding and responding to different emotional situations.
Dataset Origin
The HEAR dataset was created using elements from the HRECPW dataset, which contains 11 emotional categories with 11,000… See the full description on the dataset page: https://huggingface.co/datasets/Largefilly/HEAR-Hispanic_Emotional_Accompaniment_Responses.ExplainableAI-emotions-DPO-ORPO-RLHF
Preference Dataset for Explainable Multi-Label Emotion Classification
This repository contains a preference dataset compiled to compare two model-generated responses for explaining multi-label emotion classifications on Tweets. The dataset is accompanied by human annotations indicating which response was preferred, based on a set of defined dimensions (clarity, correctness, helpfulness, and verbosity). The annotation guidelines are included to describe how these preference judgments… See the full description on the dataset page: https://huggingface.co/datasets/imhmdf/ExplainableAI-emotions-DPO-ORPO-RLHF.Emotional_Intelligence_Corpus
Emotional Intelligence
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
relevance_score: Relevance to the subject (0-1)
quality_score: Content quality score (0-1)
topics: JSON array of detected topics
character_count: Length of the text… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Emotional_Intelligence_Corpus.Emotional_Intelligence_Content_1
Emotional Intelligence Content 1
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Emotional_Intelligence_Content_1.Emotional_Intelligence_Content_2
Emotional Intelligence Content 2
This corpus was automatically generated by the Deku Corpus Builder for use in RAG-based AI applications.
Dataset Structure
Each record contains:
text: The content text
source_url: Original source URL
source_title: Title of the source document
source_domain: Domain of the source
license_type: License classification (e.g. public_domain, cc_by, cc_by_sa)
attribution_required: Boolean — True for CC BY / CC BY-SA and other attribution-required… See the full description on the dataset page: https://huggingface.co/datasets/PhillyMac/Emotional_Intelligence_Content_2.ryancodrai-emotion-probes
Emotion Probes Roleplaying Dataset
This dataset is a reformatted, roleplay-centric adaptation of the ryancodrai/emotion-probes dataset. It focuses on scenarios where a character masks their true internal emotion with a different displayed emotional state.
Dataset Description
The dataset contains dialogues where one character attempts to deflect or obscure their real feelings through a specific, contrasting displayed emotion.
Modifications from the original:
Format:… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/ryancodrai-emotion-probes.EmotionAtlas
EmotionAtlas: Emotional Prompt Dataset
Overview
EmotionAtlas is a high-quality synthetic dataset containing emotionally-rich textual prompts paired with corresponding emotional states.
The dataset was generated using the powerful language capabilities of Google Gemini 2.0 Flash, enabling nuanced and varied examples across a range of emotional experiences.
Each entry in EmotionAtlas consists of:
Prompt: User's input.
Emotional Description: A brief, yet detailed description… See the full description on the dataset page: https://huggingface.co/datasets/entfane/EmotionAtlas.
