datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
emojisA collection of 38,176 emoji images from Facebook, Google, Apple, WhatsApp, Samsung, JoyPixels, Twitter, emojidex, LG, OpenMoji, and Microsoft. It includes all the emojis for these apps/platforms as of early 2022.
Counts: Facebook=3664, Google=3664, Apple=3961, WhatsApp=3519, Samsung=3752, JoyPixels=3538, Twitter=3544, emojidex=2040, LG=3051, OpenMoji=3512, Microsoft=3931.
Sizes: Facebook=144x144, Google=144x144, Apple=144x144, WhatsApp=144x144, Samsung=108x108, JoyPixels=144x144… See the full description on the dataset page: https://huggingface.co/datasets/rocca/emojis.silly-emoji-qaThe silly dataset to take text questions and return emoji-only answers.
Powered by ChatGPT
Examples
Why do we have different seasons? 🌍☀️🔄
How do fish breathe underwater? 🐟💦💨
Import
# !pip install -q datasets
from datasets import load_dataset
train_ds, test_ds = load_dataset("hululuzhu/silly-emoji-qa", split=["train", "test"])
# import pandas as pd
# train_df = pd.DataFrame(train)
emoji-tha-classificationSemEval-2018-Task-2-english-emojisSharif-EmojiVerse-72M
Silver-Label Multilingual Sentiment Distillation Data
Derived artifacts from the study "From 10K Labels to 72M Classifications: Scaling LLM Silver-Label Distillation for Multilingual Sentiment" (Ullah, HHAI-KEML 2026, CEUR-WS proceedings).
Preprint: https://zenodo.org/records/21786160
Dataset Summary
This repository contains derived artifacts from a study of LLM silver-label distillation on 72 million multilingual social media comments. It does not contain the raw… See the full description on the dataset page: https://huggingface.co/datasets/sharifmmm/Sharif-EmojiVerse-72M.NSM-emojiMachineLearning_EmojiDataset_Nov17emoji-with-descriptions-en-ru
Emoji Dataset with English and Russian Descriptions
4 724 emoji entries with Unicode metadata and Russian translations — built for NLP pipelines that process emoji-rich social media text in both English and Russian.
Dataset structure
column
description
emoji
emoji character (e.g. 😊)
name
official Unicode English name
group
Unicode group (e.g. Smileys & Emotion)
sub_group
Unicode sub-group (e.g. face-smiling)
codepoints
Unicode codepoint (e.g. 1F600)… See the full description on the dataset page: https://huggingface.co/datasets/PlatinumKub/emoji-with-descriptions-en-ru.Emojii_v15.1all v15.1 emojii with description and context:
ex. 🚶♂️,man walking,man walking
EmojiToResponseDatasetlowercase-emoji-fewshot-v5ml-emoji-storyEmojiDatasetnew_emotion_with_emojilowercase-emoji-zeroshotvietnamese-emoji-descriptiondf_emoji_20k_onecolumnsboaz_emoji_10k_onecolumnslowercase-emoji-rawemoji_infoindo-emoji-dictionarylowercase-emoji-oneshotemoji_sentimentvc_emoji
