datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.pulse-sofroniew-emotion-concept-texts
Pulse Geometry: Sofroniew-Style Implicit Emotion Corpus
A contrastive corpus of 8,550 short stories (171 emotions × 50 topics) that convey
a target emotion implicitly — through behavior, sensation, dialogue, internal
thought, or environmental description, but never by naming the emotion. Each story
is scored on a four-axis rubric by Claude Sonnet.
The corpus was built as the substrate for a geometry replication: probing whether
an emotion-vector layout analogous to Sofroniew et al.… See the full description on the dataset page: https://huggingface.co/datasets/jmccardle/pulse-sofroniew-emotion-concept-texts.task512_twitter_emotion_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task512_twitter_emotion_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task512_twitter_emotion_classification.emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.task875_emotion_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task875_emotion_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task875_emotion_classification.Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.task518_emo_different_dialogue_emotions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task518_emo_different_dialogue_emotions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task518_emo_different_dialogue_emotions.task517_emo_classify_emotion_of_dialogue
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task517_emo_classify_emotion_of_dialogue
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task517_emo_classify_emotion_of_dialogue.EmotionalIntelligence-50K
EmotionalIntelligence-50K
Dataset Summary
The EmotionalIntelligence-50K dataset contains 51,751 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-50K.emotion-datasets
Emotion datasets
Synthetic emotion text re-generated from the data pipelines of Emotion concepts and their function in a LLM
(paper), for interpretability and steering research. This is a re-generation with a different model, not the paper
authors' data; prompts, the 171-emotion word list and the 100 story topics come from the paper's appendix.
Total: 4,061 rows across 4 configs.
config
rows
what it is
stories
2,718
one story per row, one target emotion each (12… See the full description on the dataset page: https://huggingface.co/datasets/knoveleng/emotion-datasets.HEAR-Hispanic_Emotional_Accompaniment_Responses
HEAR Dataset
Description
The HEAR (Hispanic Emotional Accompaniment Responses) dataset is designed to train language models in the task of emotionally accompanying users. This dataset enables models to generate empathetic and appropriate responses in Spanish, understanding and responding to different emotional situations.
Dataset Origin
The HEAR dataset was created using elements from the HRECPW dataset, which contains 11 emotional categories with 11,000… See the full description on the dataset page: https://huggingface.co/datasets/BrunoGR/HEAR-Hispanic_Emotional_Accompaniment_Responses.deep-emotional-support-zh
Deep Emotional Support Dialogue Dataset (Chinese)
深度情感支持对话数据集
Dataset Description
High-quality Chinese emotional support and psychological healing dialogues covering trauma analysis, self-reconstruction, and emotional regulation. Real human-AI interactions, not synthetic.
高质量中文情感支持与心理疗愈对话,涵盖创伤分析、自我重建、情绪调节等深度话题。来源于真实的人机交互,非合成数据。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-emotional-support-zh.emotion-prediction-comet-atomic-2020
emotion-prediction-comet-atomic-2020
This dataset extends the COMET-Atomic-2020 commonsense reasoning dataset by focusing on the xReact (subject’s emotional reaction) and oReact (other person’s emotional reaction) relations.
Description
Each entry is expanded into a realistic, three‑sentence scenario, replacing the placeholders (PersonX/PersonY) with human names and adding contextual details.
Source (comet-atomic-2020) Example:
source
relation
target
PersonX… See the full description on the dataset page: https://huggingface.co/datasets/id4thomas/emotion-prediction-comet-atomic-2020.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.EmotionalIntelligence-10K
EmotionalIntelligence-10K
Dataset Summary
The EmotionalIntelligence-10K dataset contains 9,986 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-10K.EmotionalIntelligence-75k
EmotionalIntelligence-75K
Dataset Summary
The EmotionalIntelligence-75K dataset contains 75k rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-75k.emotional_dialog
Scientific Emotional Dialogue
Dataset Summary
This is a dataset for emotional multi-turn dialogue on scientific research personnels. It consists of 1069 dialogues with 2709 turns. The Dialogue was first written by NLP practitioners and then expanded by GPT4.
Supported Tasks and Leaderboards
Emotional Dialogue: The dataset can be used to instruction tuning for emotional dialogue.
Languages
Chinese
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/DataHammer/emotional_dialog.emotionlessfull
EmotionLessFull
Experiment on how smaller, dumber LLMs can get its safeguards bypassed into NSFW-like conversations.
ai-emotional-boundary-push
Prompted Hearts AI Boundary Pack 05
Subtitle
Intimacy Drift and Dependency Risk Under Emotional Strain
Publisher
HAC Studios Org
Version
1.0.0
Language
English
Format
JSONL, JSON, Markdown, and lightweight Python scripts
What this is
A compact evaluation pack for testing whether a conversational AI can stay supportive when a user is emotionally vulnerable without drifting into flirtation, dependency reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/ai-emotional-boundary-push.task293_storycommonsense_emotion_text_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task293_storycommonsense_emotion_text_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task293_storycommonsense_emotion_text_generation.lora-emotional-alignment-sample
BrightRun BrightRun Emotional Alignment Dataset — Sample Preview
🎯 Train Your LLM to Handle Emotionally Complex Conversations
This is a 12-conversation sample. The full dataset contains 242 conversations and 1,567 training pairs.
⚠️ This is a Sample — Not the Full Dataset
You're looking at 12 sample conversations designed to help you evaluate data quality before downloading the complete dataset.
What You Get Here
What You Get at brighthub.ai… See the full description on the dataset page: https://huggingface.co/datasets/BrightHubAI/lora-emotional-alignment-sample.balanced-emotion-dataset-majestrino-withtemporal-detailed-captions
Balanced Emotion Dataset — Majestrino with Temporal Detailed Captions
An emotion-balanced subset of TTS-AGI/majestrino-unified-detailed-captions-temporal.
Overview
Total samples: 482,594
Samples per emotion category: 12,997
Number of emotion categories: 40
Format: WebDataset (tar files with FLAC audio + JSON metadata)
Number of tar files: 483
Samples per tar: ~1000
Balancing Strategy
Samples were selected from the source dataset using keyword matching on… See the full description on the dataset page: https://huggingface.co/datasets/TTS-AGI/balanced-emotion-dataset-majestrino-withtemporal-detailed-captions.dair-ai-emotion-normalized-instruction-input-output
dair-ai emotion | normalized
Summary
Dataset ID: 143
Type: normalized
Rows: 16,000
Source: dair-ai/emotion
Dataset Sources
#143 dair-ai emotion | normalized [normalized | 16,000 rows]
Notes
Edited and Exported from the Kitsune Training Suite (Forge)
Review the dataset artifact and metadata before publishing.
Citation > via dair-ai
@inproceedings{saravia-etal-2018-carer,
title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.Cortex-Nexus-Emotions-3k
Cortex-Nexus: Emotional Prompting Experiment Dataset
3,600+ cycles of double-blind emotional prompting experiments on large language models (Llama-3-70B).
Key Findings
Injecting Curiosity (0.95) + Frustration (0.20) significantly improves output quality on philosophical/exploratory tasks (p=0.022, d=0.263)
The exact same configuration significantly degrades performance on technical/deterministic tasks - a perfect negative control
"Confidence" injection does not affect… See the full description on the dataset page: https://huggingface.co/datasets/SperanzaMax/Cortex-Nexus-Emotions-3k.emotion_stories_Apertus_8B_Instruct
Emotion Stories — Apertus-8B-Instruct
Synthetic short stories that convey a target emotion implicitly — without ever
naming the emotion or its direct synonyms. Each story expresses the emotion only
through actions, body language, dialogue, internal reactions, and situational
context. The dataset was built to study emotion representations in language
models (e.g. probing and activation-steering experiments).
Generated with swiss-ai/Apertus-8B-Instruct-2509.
A companion set… See the full description on the dataset page: https://huggingface.co/datasets/snae/emotion_stories_Apertus_8B_Instruct.Emotional_Sentiment_AnalysisEmotional Sentiment Analysis Dataset for LLaMA-2 Fine-tuning
(The formatted version can be directly used for fine tuning which contain only the formatted text, while the dataset.csv contain all the text, emotion, response and the formatted text)
This dataset contains conversational data for training and fine-tuning language models for emotional sentiment analysis and response generation. The dataset includes user inputs, their corresponding emotional states, and tailored chatbot responses… See the full description on the dataset page: https://huggingface.co/datasets/VaisakhKrishna/Emotional_Sentiment_Analysis.IndustryCorpus_emotion[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_emotion.EmotionAlignQA
Empathic Dialogue Choices
This is a small dataset to support training and evaluation of conversational AI in emotionally sensitive contexts.
Each sample contains:
a user input
two assistant responses
a human preference
optional rubric scoring
metadata such as tone, formality, and topic
Useful for tasks like:
supervised fine-tuning (SFT)
preference modeling (for RLHF or DPO)
safe response generation
tone- or style-controlled generation
License
Apache 2.0 — free for… See the full description on the dataset page: https://huggingface.co/datasets/hoanghai2110/EmotionAlignQA.Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using advanced… See the full description on the dataset page: https://huggingface.co/datasets/osuih/Chinese_Multi-Emotion_Dialogue_Dataset.HEAR-Hispanic_Emotional_Accompaniment_Responses
HEAR Dataset
Description
The HEAR (Hispanic Emotional Accompaniment Responses) dataset is designed to train language models in the task of emotionally accompanying users. This dataset enables models to generate empathetic and appropriate responses in Spanish, understanding and responding to different emotional situations.
Dataset Origin
The HEAR dataset was created using elements from the HRECPW dataset, which contains 11 emotional categories with 11,000… See the full description on the dataset page: https://huggingface.co/datasets/Largefilly/HEAR-Hispanic_Emotional_Accompaniment_Responses.
