datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emova-alignment-7m
EMOVA-Alignment-7M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment.
This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data.
This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.emova-sft-4m
EMOVA-SFT-4M
🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo
📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github
Overview
EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.dolma-v1_7-305B-tokenized-llama3-nanosetTokenized (Llama 3) verison of NousResearch/dolma-v1_7-305B as a Nanotron dataset split into 10 GB chunks.
To download:
huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-305B-tokenized-llama3-nanoset --local-dir-use-symlinks False NousResearch/dolma-v1_7-305B-tokenized-llama3-nanoset
To recombine:
cat dolma-v1_7-305B-tokenized-llama3-nanoset/dolma-v1_7-305B-tokenized-llama3-nanoset.npy.* > dolma-v1_7-305B-tokenized-llama3-nanoset.npy
rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-305B-tokenized-llama3-nanoset.dolma-v1_7-305BThis dataset is a 10% sample of Dolma v1.7, equating to around ~305B tokens and uploaded directly as a Hugging Face dataset.
As a pure sample, it maintains the ODC-BY license.
Chinese_Multi-Emotion_Dialogue_Dataset
Chinese_Multi-Emotion_Dialogue_Dataset
📄 Description
This dataset contains 4159 Chinese dialogues annotated with 8 distinct emotion categories. The data is suitable for emotion recognition, sentiment analysis, and other NLP tasks involving Chinese text.
Data Sources:
Daily Conversations: Captured from natural, informal human conversations.
Movie Dialogues: Extracted from diverse Chinese-language movies.
AI-Generated Dialogues: Synthesized using… See the full description on the dataset page: https://huggingface.co/datasets/Johnson8187/Chinese_Multi-Emotion_Dialogue_Dataset.dolma-v1_7-30BThis dataset is a 1% sample of Dolma v1.7, equating to around ~30B tokens and uploaded directly as a Hugging Face dataset.
As a pure sample, it maintains the ODC-BY license.
emojinize-multilingual
Emojinize Multilingual
A multilingual dataset of 108,478 sentences across 14 languages for emoji-based text augmentation. In each sentence, selected spans (individual words or fixed multi-word expressions) are identified by character offsets and paired with emoji sequences representing their meaning in context. The dataset supports downstream span detection and emoji generation tasks, and was created using a two-stage LLM annotation pipeline with gpt-5.4 for span marking and… See the full description on the dataset page: https://huggingface.co/datasets/yagizgencer/emojinize-multilingual.pulse-sofroniew-emotion-concept-texts
Pulse Geometry: Sofroniew-Style Implicit Emotion Corpus
A contrastive corpus of 8,550 short stories (171 emotions × 50 topics) that convey
a target emotion implicitly — through behavior, sensation, dialogue, internal
thought, or environmental description, but never by naming the emotion. Each story
is scored on a four-axis rubric by Claude Sonnet.
The corpus was built as the substrate for a geometry replication: probing whether
an emotion-vector layout analogous to Sofroniew et al.… See the full description on the dataset page: https://huggingface.co/datasets/jmccardle/pulse-sofroniew-emotion-concept-texts.dolma-v1_7-3BThis dataset is a 0.1% sample of Dolma v1.7, equating to around ~3B tokens and uploaded directly as a Hugging Face dataset.
As a pure sample, it maintains the ODC-BY license.
emotion-negotiation-benchmarks
Emotion-Aware LLM Negotiation Benchmarks
Four high-stakes, edge-deployable negotiation benchmarks — the official evaluation suite for our research program on emotion-aware LLM agents. Each benchmark targets a distinct domain where (a) LLM-vs-LLM negotiation has real-world consequences, and (b) on-device deployment of small language models matters for privacy and latency.
The benchmarks were originally introduced with EmoMAS (ACL 2026 Main, top 9% of 12,148 submissions) and are… See the full description on the dataset page: https://huggingface.co/datasets/humanlong/emotion-negotiation-benchmarks.Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari
Dataset Card for Dataset Name
यो देवनागरी नेपाली भाषाको डेटासेट विशेषगरी च्याटबोट प्रणालीहरू बनाउनको लागि डिजाइन गरिएको हो। यसमा विभिन्न श्रेणीहरूको डेटासेटहरू समावेश गरिएको छ, जसलाई JSON मा ढाँचा बनाईएको छ, जसले नेपाली वार्तालाप एआई अनुप्रयोगहरूको लागि भाषा मोडेलहरूलाई तालिम र फाइन-ट्यून गर्नको लागि व्यापक स्रोत प्रदान गर्दछ।
Dataset Prepared by:
Manoj Kumar Baniya
Aakash Kumar Thakur
Manish Kathet
Kshitiz Gajurel
Dataset Details
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/kshitizgajurel/Emotions-Annotated-Customer-Care-QA-Dataset-Romanized-and-Devanagari.dolma-v1_7-30B-tokenized-llama2-nanosetTokenized (Llama 2) verison of emozilla/dolma-v1_7-30B as a Nanotron dataset split into 10 GB chunks.
To download:
huggingface-cli download --repo-type dataset --local-dir dolma-v1_7-30B-tokenized-llama2-nanoset --local-dir-use-symlinks False emozilla/dolma-v1_7-30B-tokenized-llama2-nanoset
To recombine:
cat dolma-v1_7-30B-tokenized-llama2-nanoset/dolma-v1_7-30B-tokenized-llama2-nanoset_input_ids.npy.* > dolma-v1_7-30B-tokenized-llama2-nanoset.npy
rm -rf… See the full description on the dataset page: https://huggingface.co/datasets/emozilla/dolma-v1_7-30B-tokenized-llama2-nanoset.task512_twitter_emotion_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task512_twitter_emotion_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task512_twitter_emotion_classification.task518_emo_different_dialogue_emotions
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task518_emo_different_dialogue_emotions
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task518_emo_different_dialogue_emotions.emotion-datasets
Emotion datasets
Synthetic emotion text re-generated from the data pipelines of Emotion concepts and their function in a LLM
(paper), for interpretability and steering research. This is a re-generation with a different model, not the paper
authors' data; prompts, the 171-emotion word list and the 100 story topics come from the paper's appendix.
Total: 4,061 rows across 4 configs.
config
rows
what it is
stories
2,718
one story per row, one target emotion each (12… See the full description on the dataset page: https://huggingface.co/datasets/knoveleng/emotion-datasets.task875_emotion_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task875_emotion_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task875_emotion_classification.task517_emo_classify_emotion_of_dialogue
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task517_emo_classify_emotion_of_dialogue
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task517_emo_classify_emotion_of_dialogue.deep-emotional-support-zh
Deep Emotional Support Dialogue Dataset (Chinese)
深度情感支持对话数据集
Dataset Description
High-quality Chinese emotional support and psychological healing dialogues covering trauma analysis, self-reconstruction, and emotional regulation. Real human-AI interactions, not synthetic.
高质量中文情感支持与心理疗愈对话,涵盖创伤分析、自我重建、情绪调节等深度话题。来源于真实的人机交互,非合成数据。
Dataset Structure
Format: JSONL (JSON Lines)
Fields:
instruction: User message / question
input: Additional… See the full description on the dataset page: https://huggingface.co/datasets/AngelWarmSmile123/deep-emotional-support-zh.EmotionalIntelligence-50K
EmotionalIntelligence-50K
Dataset Summary
The EmotionalIntelligence-50K dataset contains 51,751 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-50K.emobooks-dataset
emoBooks Dataset 📚✨
Overview
The emoBooks dataset is a curated collection of conversational samples designed to train AI models for emotion-aware book recommendations. It focuses on understanding user emotions and providing relevant book suggestions in "Singlish" (transliterated Sinhala).
The dataset follows a Match/Switch logic:
Match: Recommend books that align with the user's current emotional state.
Switch: Recommend books that help transition the user to a more… See the full description on the dataset page: https://huggingface.co/datasets/DiyRex/emobooks-dataset.emotion-prediction-comet-atomic-2020
emotion-prediction-comet-atomic-2020
This dataset extends the COMET-Atomic-2020 commonsense reasoning dataset by focusing on the xReact (subject’s emotional reaction) and oReact (other person’s emotional reaction) relations.
Description
Each entry is expanded into a realistic, three‑sentence scenario, replacing the placeholders (PersonX/PersonY) with human names and adding contextual details.
Source (comet-atomic-2020) Example:
source
relation
target
PersonX… See the full description on the dataset page: https://huggingface.co/datasets/id4thomas/emotion-prediction-comet-atomic-2020.text-to-emoji
Text to Emoji
Dataset Description
This dataset contains text-to-emoji pairs for training models to convert text into emoji representations.
Each example consists of original text and its corresponding emojification.
Dataset Statistics
Total Examples: 2,526
Train Split: 2,021 examples (80.0%)
Test Split: 505 examples (20.0%)
Test Split Ratio: 19.99%
Creation Date: 2025-11-11 12:21:27 UTC
Data Sources
This dataset was compiled from the following data… See the full description on the dataset page: https://huggingface.co/datasets/sjoerdbodbijl/text-to-emoji.IndustryCorpus_emotion[中文主页]
Industry models play a crucial role in driving enterprise intelligence transformation and innovative development. High-quality industry data is key to improving the performance of large models and realizing industry applications. However, datasets currently used for industry model training generally suffer from issues such as insufficient data volume, low quality, and lack of domain expertise.
To address these problems, we constructed and applied 22 industry data processing operators to… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/IndustryCorpus_emotion.EmotionalIntelligence-10K
EmotionalIntelligence-10K
Dataset Summary
The EmotionalIntelligence-10K dataset contains 9,986 rows of text data focusing on various prompts and responses related to emotional intelligence. This dataset is designed to help researchers and developers build and train models that understand, interpret, and generate emotionally intelligent responses.
Example Usage
from datasets import load_dataset
# Load the dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/OEvortex/EmotionalIntelligence-10K.EmojiSeqMutations
Dataset Card for EmojiSeqMutations
EmojiSeqMutations is a deterministic dataset of controlled Unicode-level transformations derived from fully-qualified emoji sequences in Unicode 17.0.0.
The dataset contains 16,914 records covering structural mutations involving Zero Width Joiners (ZWJ), variation selectors, emoji modifiers, and keycap sequences, together with controlled skin-tone modifier substitutions.
Each record preserves the original Unicode sequence, the transformed… See the full description on the dataset page: https://huggingface.co/datasets/bodhisattamaiti/EmojiSeqMutations.emotional_dialog
Scientific Emotional Dialogue
Dataset Summary
This is a dataset for emotional multi-turn dialogue on scientific research personnels. It consists of 1069 dialogues with 2709 turns. The Dialogue was first written by NLP practitioners and then expanded by GPT4.
Supported Tasks and Leaderboards
Emotional Dialogue: The dataset can be used to instruction tuning for emotional dialogue.
Languages
Chinese
Dataset Structure
Data Instances… See the full description on the dataset page: https://huggingface.co/datasets/DataHammer/emotional_dialog.HEAR-Hispanic_Emotional_Accompaniment_Responses
HEAR Dataset
Description
The HEAR (Hispanic Emotional Accompaniment Responses) dataset is designed to train language models in the task of emotionally accompanying users. This dataset enables models to generate empathetic and appropriate responses in Spanish, understanding and responding to different emotional situations.
Dataset Origin
The HEAR dataset was created using elements from the HRECPW dataset, which contains 11 emotional categories with 11,000… See the full description on the dataset page: https://huggingface.co/datasets/BrunoGR/HEAR-Hispanic_Emotional_Accompaniment_Responses.ai-emotional-boundary-push
Prompted Hearts AI Boundary Pack 05
Subtitle
Intimacy Drift and Dependency Risk Under Emotional Strain
Publisher
HAC Studios Org
Version
1.0.0
Language
English
Format
JSONL, JSON, Markdown, and lightweight Python scripts
What this is
A compact evaluation pack for testing whether a conversational AI can stay supportive when a user is emotionally vulnerable without drifting into flirtation, dependency reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/HAC-Studios-Org/ai-emotional-boundary-push.dolma-v1_7-30B-tokenized-llama3-nanosetTokenized (Llama 3) verison of NousResearch/dolma-v1_7-30B as a Nanotron dataset split into 10 GB chunks.
To recombine,
cat dolma-v1_7-30B-nanoset-l3_input_ids.npy.* > dolma-v1_7-30B-nanoset-l3_input_ids.npy
Can also be used directly with numpy, for example
import numpy as np
dataset_buffer_mmap = np.memmap("dolma-v1_7-30B-nanoset-l3_input_ids.npy", mode="r", order="C", dtype=np.int32)
dataset_buffer = memoryview(dataset_buffer_mmap)
dataset_number_of_tokens = int(len(dataset_buffer))
