emoji
Datasets
All datasets matching “emoji”emojis
Dataset Card for Emojis
This is a FiftyOne dataset with 1816 samples.
Installation
If you haven't already, install FiftyOne:
pip install -U fiftyone
Usage
import fiftyone as fo
import fiftyone.utils.huggingface as fouh
# Load the dataset
# Note: other available arguments include 'max_samples', etc
dataset = fouh.load_from_hub("jamarks/emojis")
# Launch the App
session = fo.launch_app(dataset)
Dataset Details
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/jamarks/emojis.irodori-clones-3m-v2-no-emoji
Irodori TTS Clones v2 (3.29M)
3,290,000 cloned utterances generated with Aratako/Irodori-TTS-500M-v2,
using the 10,000 reference voices from SynData-2/irodori-refs-10k-v2.
329 clones per ref voice, each with a unique Japanese conversational text.
Companion refs: SynData-2/irodori-refs-10k-v2.
Note: Bu dataset irodori-clones-3m-v2'nin emoji-temizlenmis kopyasidir. Audio bytes binary-identical; yalnizca text kolonundaki emojiler kaldirilmistir (emoji kutuphanesi, Japonca/CJK… See the full description on the dataset page: https://huggingface.co/datasets/SynDataLab-JA/irodori-clones-3m-v2-no-emoji.pashto-emoji-dataset
Pashto Emoji Dataset
This dataset is a Pashto translation of the KomeijiForce/Text2Emoji dataset. It is designed for tasks involving the translation of text into emoji sequences and understanding the sentiment or topic of a given text.
The dataset contains over 504,000 rows, each consisting of a text passage in Pashto, a corresponding emoji sequence, and a topic label.
Dataset Structure
The dataset is provided in the following format:
text: A string containing… See the full description on the dataset page: https://huggingface.co/datasets/nassimjp/pashto-emoji-dataset.telegram-animated-emojis-60fps
Telegram Official Animated Emojis (60.00 FPS Vector Lottie)
Complete official dataset of all 599 animated vector emojis extracted directly from Telegram's official sticker set (AnimatedEmojies).
All animations run natively at 60.00 FPS, infinitely scalable vector format (Lottie JSON / .TGS), with full transparency support.
Dataset Summary
Total Emojis: 599
Frame Rate: 60.00 FPS
Resolution: Infinite Vector Scalability (Lottie / TGS)
Formats Included:
.tgs… See the full description on the dataset page: https://huggingface.co/datasets/smmisha/telegram-animated-emojis-60fps.emojinize-multilingual
Emojinize Multilingual
A multilingual dataset of 108,478 sentences across 14 languages for emoji-based text augmentation. In each sentence, selected spans (individual words or fixed multi-word expressions) are identified by character offsets and paired with emoji sequences representing their meaning in context. The dataset supports downstream span detection and emoji generation tasks, and was created using a two-stage LLM annotation pipeline with gpt-5.4 for span marking and… See the full description on the dataset page: https://huggingface.co/datasets/yagizgencer/emojinize-multilingual.DPO-zh-en-emojiA chatbot dialogue dataset with textual emojis, available in both Chinese and English versions, suitable for SFT/DPO training.
We have carefully selected some questions originating from Zhihu, logic reasoning, and Weichi Bar as Queries. These were generated using the llama3 70b instruct version, with each query producing a Chinese version of the answer and an English version of the answer. This can be used for aligning language model "language type" and "language style" tasks.
Github link:… See the full description on the dataset page: https://huggingface.co/datasets/shareAI/DPO-zh-en-emoji.
