CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01dair-ai /emotion Dataset Card for "emotion" Dataset Summary Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper. Supported Tasks and Leaderboards More Information Needed Languages More Information Needed Dataset Structure Data Instances An example looks as follows. { "text": "im feeling quite sad and sorry for myself but… See the full description on the dataset page: https://huggingface.co/datasets/dair-ai/emotion.texttext-classification100K<n<1M464 likes29k downloads2y agoHugging Face02emozilla /pg19 Dataset Card for "pg19" Paraquet version of pg19 Statistics (in # of characters): total_len: 11425076324, average_len: 399450.2595622684 text10K<n<100K20 likes16k downloads3y agoHugging Face03SetFit /emotion** Attention: There appears an overlap in train / test. I trained a model on the train set and achieved 100% acc on test set. With the original emotion dataset this is not the case (92.4% acc)** text10K<n<100K30 likes15k downloads4y agoHugging Face04google-research-datasets /go_emotions Dataset Card for GoEmotions Dataset Summary The GoEmotions dataset contains 58k carefully curated Reddit comments labeled for 27 emotion categories or Neutral. The raw data is included as well as the smaller, simplified version of the dataset with predefined train/val/test splits. Supported Tasks and Leaderboards This dataset is intended for multi-class, multi-label emotion classification. Languages The data is in English. Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/go_emotions.tabulartext-classification100K<n<1M267 likes14k downloads3y agoHugging Face05pollen-robotics /reachy-mini-emotions-library Reachy Mini Emotions Library Curated emotion recordings for the Reachy Mini robot, maintained by Pollen Robotics. Each move is a JSON trajectory (head pose, antennas, body yaw, sampled over time) paired with an Opus audio track. Motion is sampled at 50 Hz; audio is mono Ogg/Opus (decoded natively by the robot). Requires reachy_mini ≥ v1.8.4 (its move loader resolves non-.wav audio sidecars). File layout Files live at the root of the dataset, named <emotion>.json +… See the full description on the dataset page: https://huggingface.co/datasets/pollen-robotics/reachy-mini-emotions-library.audioroboticsn<1K18 likes12k downloads3mo agoHugging Face06EQ4You /Emotional_SpeechThis dataset contains audio-text pairs in the webdataset format. The audio files are short speech segments from publicly available videos & the texts are descriptions of emotions the speakers seems to be feeling. Some captions also describe the speakers gender and age. All files with the substring "part1" in the name contain unique audio files with unique captions. All files with the substring "part2" , "part3", ... in the name contain the same audio files as in "part1", but with different… See the full description on the dataset page: https://huggingface.co/datasets/EQ4You/Emotional_Speech.5 likes8.7k downloads2y agoHugging Face07abotresol /emotion-vectors-gemma-4-31b-it-postfix Emotion vectors, google/gemma-4-31b-it (corrected extraction) Residual-stream activations for google/gemma-4-31b-it, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b-it. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it-postfix.feature-extraction0 likes5.8k downloads2mo agoHugging Face08mteb /emotion EmotionClassification An MTEB dataset Massive Text Embedding Benchmark Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. Task category t2c Domains Social, Written Reference https://www.aclweb.org/anthology/D18-1404 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task = mteb.get_tasks(["EmotionClassification"])… See the full description on the dataset page: https://huggingface.co/datasets/mteb/emotion.texttext-classification10K<n<100K18 likes5.4k downloads1y agoHugging Face09Ayesha758 /Emotion_new_collected_datasetaudio10K<n<100K0 likes5.1k downloads4mo agoHugging Face10VoiceNet /emolia-thinking Emolia-Thinking — a VoiceNet-annotated, balanced subset of Emolia Emolia-Thinking is a richly annotated speech dataset created for the VoiceNet project. It takes a balanced subset of the Emolia corpus — balanced across speaker-embedding clusters and emotion-embedding clusters so that speakers, voices and emotional states are evenly represented rather than dominated by the most common cases — and annotates every clip along the full VoiceNet Extended voice-performance taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/VoiceNet/emolia-thinking.audioaudio-classification100K<n<1M0 likes5.1k downloads3mo agoHugging Face11ChristophSchuhmann /emotionsimage100K<n<1M2 likes3.9k downloads3y agoHugging Face12emozilla /quality Dataset Card for "quality" More Information needed text1K<n<10K7 likes3.9k downloads3y agoHugging Face13Emova-ollm /emova-alignment-7m EMOVA-Alignment-7M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-Alignment-7M is a comprehensive dataset curated for omni-modal pre-training, including vision-language and speech-language alignment. This dataset is created using open-sourced image-text pre-training datasets, OCR datasets, and 2,000 hours of ASR and TTS data. This dataset is part of the EMOVA-Datasets… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-alignment-7m.imageimage-to-text1M<n<10M10 likes3.5k downloads2y agoHugging Face14emozilla /yarn-train-tokenized-16k-mistral Dataset Card for "yarn-train-tokenized-16k-mistral" More Information needed 100K<n<1M14 likes3.3k downloads3y agoHugging Face15Emova-ollm /emova-sft-4m EMOVA-SFT-4M 🤗 EMOVA-Models | 🤗 EMOVA-Datasets | 🤗 EMOVA-Demo 📄 Paper | 🌐 Project-Page | 💻 Github | 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-4M is a comprehensive dataset curated for omni-modal instruction tuning, including textual, visual, and audio interactions. This dataset is created by gathering open-sourced multi-modal instruction datasets and synthesizing high-quality omni-modal conversation data to enhance user experience. This dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Emova-ollm/emova-sft-4m.imageimage-to-text1M<n<10M6 likes3k downloads2y agoHugging Face16laion /laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuningLAION's Got Talent: Generated Voice Acting Dataset Overview "LAION's Got Talent" is a synthetic voice acting dataset designed to offer a broad range of emotional expressions, vocal bursts, and multi-language utterances. This dataset is a component of the BUD-E project, led by LAION with support from Intel, and aims to drive forward research in context-aware and empathetic AI voice assistants. Updated Composition Voices and Languages English: 11 OpenAI voices, each… See the full description on the dataset page: https://huggingface.co/datasets/laion/laions_got_talent_with_voice_emotion_speed_tags_for_orpheus_tuning.audio1M<n<10M6 likes2.7k downloads1y agoHugging Face17Tongji-Emotion /Robot-EQ RobotEQ-Data Official dataset release for RobotEQ. Evaluation & Scripts For inference scripts, evaluation scripts, and data production tooling, see the RobotEQ code repository. Dataset Statistics Item Count Behavior judgment scenarios (synthetic) 1,812 Behavior judgment scenarios (real POV) 223 Behavior judgment scenarios (total) 2,035 Behavior judgment behavior annotations 3,171 Spatial grounding questions 825… See the full description on the dataset page: https://huggingface.co/datasets/Tongji-Emotion/Robot-EQ.imagevisual-question-answering1K<n<10K9 likes2.7k downloads19d agoHugging Face18emozilla /pg19-test Dataset Card for "pg19-test" More Information needed textn<1K5 likes2.4k downloads3y agoHugging Face19langswap /dialogs-ru-emotional-conversations Dialogs: A Studio-Quality Expressive Conversational Russian Speech Corpus Dialogs is a 20.6-hour studio-quality corpus of expressive, conversational Russian speech, designed for dialog-oriented and emotional text-to-speech. Unlike existing Russian corpora — mostly single-speaker read speech or large but low-quality web-mined audio — Dialogs was recorded by professional theatre actors performing scripted dialogs face-to-face, capturing natural turn-taking, timing, and expressive… See the full description on the dataset page: https://huggingface.co/datasets/langswap/dialogs-ru-emotional-conversations.audiotext-to-speechn<1K18 likes2.4k downloads2mo agoHugging Face20laion /Emolia Dataset Card for Emolia Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range of emotional and… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emolia.audio10M<n<100M15 likes2.1k downloads10mo agoHugging Face21laion /Emilia-with-Emotion-Annotations Dataset Card for Emilia with Emotion Annotations Dataset Description This dataset is an enhanced version of the Emilia dataset, enriched with detailed emotion annotations. The annotations were generated using models from the EmoNet suite to provide deeper insight into the emotional content of speech. This work is based on the research and models described in the blog post "Do They See What We See?". The annotations include 54 scores for each sample, covering a wide range… See the full description on the dataset page: https://huggingface.co/datasets/laion/Emilia-with-Emotion-Annotations.29 likes2.1k downloads1y agoHugging Face22abotresol /emotion-vectors-gemma-4-31b-it Emotion vectors — gemma-4-31b-it (instruct) probed on the external gemma-4-4B story corpus Data provenance (what made these activations) Probed model: google/gemma-4-31b-it (instruct) Input corpus: snae/emotion_stories_gemma_4_4B — stories written by gemma-4-4B, a smaller EXTERNAL model (generator is NOT the probed model) Per-story pooled residual-stream activations and per-emotion mean vectors, extracted with gemma4-emotion-vectors… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-it.feature-extraction0 likes2.1k downloads2mo agoHugging Face23abotresol /emotion-vectors-gemma-4-31b Emotion vectors — gemma-4-31b (base) probed on the external gemma-4-4B story corpus Data provenance (what made these activations) Probed model (whose activations these are): google/gemma-4-31b (base) Input corpus: snae/emotion_stories_gemma_4_4B — third-person emotion stories written by gemma-4-4B, a smaller EXTERNAL model (the open replication's published corpus; generator is NOT the probed model) Per-story pooled residual-stream activations and per-emotion… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b.feature-extraction0 likes2.1k downloads2mo agoHugging Face24abotresol /emotion-vectors-gemma-4-31b-postfix Emotion vectors, google/gemma-4-31b (corrected extraction) Residual-stream activations for google/gemma-4-31b, pooled per story and averaged per emotion. Each emotion ends up as one direction in the model's activation space. Read LINEAGE.md before using this. This set supersedes abotresol/emotion-vectors-gemma-4-31b. The earlier extraction ran while the tokenizer padded on the left, so the step that skips a story's first 50 tokens skipped padding instead. This set re-extracts… See the full description on the dataset page: https://huggingface.co/datasets/abotresol/emotion-vectors-gemma-4-31b-postfix.feature-extraction0 likes2k downloads2mo agoHugging Face25emozilla /Long-Data-Collections-Pretrain-Without-Books Dataset Card for "Long-Data-Collections-Pretrain-Without-Books" Paraquet version of the pretrain split of togethercomputer/Long-Data-Collections WITHOUT books Statistics (in # of characters): total_len: 236088622215, average_len: 25159.041601590307 text1M<n<10M4 likes1.8k downloads3y agoHugging Face26brighter-dataset /BRIGHTER-emotion-categories BRIGHTER Emotion Categories Dataset This dataset contains the emotion categories data from the BRIGHTER paper: BRIdging the Gap in Human-Annotated Textual Emotion Recognition Datasets for 28 Languages. Dataset Description The BRIGHTER Emotion Categories dataset is a comprehensive multi-language, multi-label emotion classification dataset with separate configurations for each language. It represents one of the largest human-annotated emotion datasets across multiple… See the full description on the dataset page: https://huggingface.co/datasets/brighter-dataset/BRIGHTER-emotion-categories.tabular100K<n<1M18 likes1.8k downloads10mo agoHugging Face27krishnakalyan3 /emo_webds_2audio10K<n<100K7 likes1.6k downloads2y agoHugging Face28alkalol /EmoVerse EmoVerse EmoVerse is a visual emotion dataset for affective image understanding. The dataset is organized around eight emotion categories: Amusement, Anger, Awe, Contentment, Disgust, Excitement, Fear, and Sadness. The released package contains annotation files and Parquet shards for the image records and annotations. The Parquet rows store file-level data: each row describes one packed file and includes both metadata and the file content as a binary column.… See the full description on the dataset page: https://huggingface.co/datasets/alkalol/EmoVerse.image-classification100K<n<1M1 likes1.5k downloads3mo agoHugging Face29krishnakalyan3 /emo_parleraudio1M<n<10M2 likes1.4k downloads2y agoHugging Face30emozilla /pg_books-tokenized-bos-eos-chunked-65536 Dataset Card for "pg_books-tokenized-bos-eos-chunked-65536" The pg19 dataset tokenized under LLaMA into 64k chunks, bookended with BOS and EOS 10K<n<100K7 likes1.3k downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.