datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hateful_memes_expandedPumpfun_Memecoin_Corpus
PumpFun Launch-to-Graduation Corpus (Jun–Jul 2026)
798,430 pump.fun token launches. 33.58 million trades. 26.9 million bonding
-curve snapshots. Every graduation outcome labeled. Tracked continuously,
second by second, for 39 uninterrupted days.
⚠️ This dataset has documented, quantified data-quality issues — several
are not optional to handle correctly. Full detail, root causes, and
exact handling instructions: KNOWN_ISSUES.md.
Read it before you write a single query.… See the full description on the dataset page: https://huggingface.co/datasets/Slinky21/Pumpfun_Memecoin_Corpus.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/neuralcatcher/hateful_memes.ultimateMemeLens
MemeLens
A large-scale multilingual multimodal meme understanding benchmark with 46 classification tasks across 9 languages, enriched with LLM-generated explanations and LLM-as-Judge quality scores.
This is the VLM (Vision-Language Model) version of MemeLens, extended with natural language explanations for each sample and automated quality evaluation via LLM-as-Judge.
Paper: MemeLens: Multilingual Multitask VLMs for Memes
Code: MohamedBayan/MemeLens
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeLens.MemeXGen
Cross-Cultural Meme Transcreation Dataset
📄 Read our paper: Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
Dataset Summary
This dataset contains 6,315 meme pairs for cross-cultural meme transcreation between English/American and Chinese cultures. It includes both bidirectional translations, making it valuable for research in multimodal AI, cultural adaptation, and humor translation.
Key Statistics:
English →… See the full description on the dataset page: https://huggingface.co/datasets/YZhao09/MemeXGen.MemEye
MemEye
Paper | Project Page | Official Code
MemEye is a visual-centric multimodal memory benchmark for evaluating agents that need to remember and reason over long-running image-grounded dialogues. It evaluates memory capabilities across two axes: visual evidence granularity (from scene-level to pixel-level) and memory reasoning depth (from atomic retrieval to evolutionary synthesis).
The dataset includes 371 mirrored MCQ + open-ended questions across 8 life-scenario tasks… See the full description on the dataset page: https://huggingface.co/datasets/MemEyeBench/MemEye.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.solana-memecoin-calls
Solana memecoin calls — a public record with the misses left in
7,863 pump.fun token calls, each with the market cap we called it at, the peak it reached
afterwards, and the exact second it was posted publicly. The whole file is hashed and the hash is
anchored in a Bitcoin block, so no row can be added, edited or back-dated after the fact.
Every trading channel publishes its winners. This is the same feed with the losers still in it —
about six calls in ten never double, and… See the full description on the dataset page: https://huggingface.co/datasets/Smurfetc/solana-memecoin-calls.Meme_lorasMemeEffect-382K-audioWe are releasing the audio files that we have collected from MemeEffect-382K dataset. All the files are being shared as .tar files and files are rnamed using their respective id that can be found through the metadata.
We share these files as-part of research initiative.
multimodal_meme_classification_singapore
Dataset Card for Offensive Memes in Singapore Context
Dataset Details
Dataset Description
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards.
Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.lumina-memesmemebench
MemeBench: Diagnosing Cultural-Semantic Understanding in LVLMs through Memes
Dataset Summary
MemeBench is a bilingual diagnostic benchmark for open-ended meme interpretation, evaluating large vision-language models' cultural-semantic understanding. It contains 1,253 expert-annotated memes spanning 7 cultural domains in Chinese and English.
Each meme is annotated with a structured VIKR schema covering four diagnostic layers:
Visual clues: Can the model describe what it… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-neurips-2026/memebench.memesmemes_with_captionschinese-meme-description-dataset
Describe image information using the following LLM Models
gpt4o
Claude-3.5-sonnet-20240620
gemini-1.5-pro
gemini-1.5-flash
gemini-1.0-pro-vision
yi-vision
Gemini Code
# -*- coding: gbk -*-
import google.generativeai as genai
import PIL.Image
import os
import json
import shutil
from tqdm import tqdm
from concurrent.futures import ThreadPoolExecutor, as_completed
genai.configure(api_key='')
model = genai.GenerativeModel(
'gemini-1.5-pro-latest'… See the full description on the dataset page: https://huggingface.co/datasets/REILX/chinese-meme-description-dataset.CCKEB
CCKEB (Compositional/Continual Knowledge Editing Benchmark)
🌟 Overview
CCKEB is a benchmark designed for Continual and Compositional Knowledge Editing in Large Vision-Language Models (LVLMs), accepted at NeurIPS 2025.
The benchmark targets realistic knowledge update scenarios in which visual identities and textual facts are edited sequentially.Models are required to retain previously edited knowledge while answering compositional multimodal queries that depend… See the full description on the dataset page: https://huggingface.co/datasets/MemEIC/CCKEB.zh-meme-sft-8k
zh-meme-sft-8k
📖 简介 | Introduction
zh-meme-sft-8k 是一个高质量的中文互联网梗文化指令微调数据集。该数据集基于抖音、小红书、B站等平台的真实评论互动构建,经过多轮清洗、增强和格式化处理,专门用于训练能够理解和使用网络热梗、具备幽默感的对话模型。
🎯 这个数据集是 Meme-Qwen-7B-Instruct 模型的训练数据,如果你想看微调后的效果,可以直接体验模型!
这个数据集的特点是:
🎯 真实来源:基于真实社交平台的用户互动,保留原本网络表达
🔄 对话结构:包含帖子-评论、评论-回复的完整对话链
🧹 精细清洗:经过多轮规则清洗和LLM增强,去除噪声的同时保留热梗
💬 ChatML格式:标准化为ChatML格式,开箱即用
📊 数据统计 | Data Statistics
数据集
样本数量
占比
训练集
7,377
85%
验证集
868
10%
测试集
435
5%
总计
8… See the full description on the dataset page: https://huggingface.co/datasets/GaryYang123/zh-meme-sft-8k.memetrashMemeXGen
Cross-Cultural Meme Transcreation Dataset
📄 Read our paper: Beyond Translation: Cross-Cultural Meme Transcreation with Vision-Language Models
Dataset Summary
This dataset contains 6,315 meme pairs for cross-cultural meme transcreation between English/American and Chinese cultures. It includes both bidirectional translations, making it valuable for research in multimodal AI, cultural adaptation, and humor translation.
Key Statistics:
English →… See the full description on the dataset page: https://huggingface.co/datasets/AIM-SCU/MemeXGen.MementosOfficial Dataset for Mementos: A Comprehensive Benchmark for Multimodal Large Language Model Reasoning over Image Sequences.
All data is downloaded from the original google drive link.
MemeXplain
MemeXplain Dataset
MemeXplain is a comprehensive multimodal dataset for detecting and explaining propagandistic and hateful content in memes. It consists of two main components:
Dataset Components
1. ArMemeXplain (Arabic Propaganda Memes)
Train: 4,007 samples
Dev: 584 samples
Test: 1,134 samples
Total: 5,725 Arabic memes with propaganda annotations
This dataset is derived from the ArMeme corpus and includes:
Arabic memes with text overlay
Binary… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeXplain.memerag
MEMERAG Faithfulness LLM-Judge Labels
This dataset extends MEMERAG (Cruz Blandón et al., 2025, ACL 2025,
arXiv:2502.17163) with LLM-judge faithfulness predictions from gpt-5.4, reusing
MEMERAG's own data and judge prompt templates.
Files
File
Description
memerag_llm_judge_<prompt_variant>_<model>.jsonl
LLM-judge predictions (produced by scripts/annotate_memerag.py)
scripts/
Reproduction scripts (see below)
Dataset statistics… See the full description on the dataset page: https://huggingface.co/datasets/Glide-py/memerag.fog_fire_detection
Fire & Smoke Detection Dataset (40K Images)
A large-scale, annotated object detection dataset containing over 40,900 images dedicated to early fire and smoke detection. Designed for training real-time vision models such as YOLO (Ultralytics), RT-DETR, and Vision Transformers.
Dataset Summary
Total Images: ~40,900 images
Task: Object Detection (object-detection)
Bounding Box Format: YOLO format (class_id x_center y_center width height) / COCO format
Target… See the full description on the dataset page: https://huggingface.co/datasets/jojomoi-meme/fog_fire_detection.BHM-Bengali-Hateful-Memes
Dataset Description
BHM is a novel multimodal dataset for Bengali Hateful Memes detection. The dataset consists of 7,148 memes with Bengali as well as code-mixed captions,
tailored for two tasks: (i) detecting hateful memes and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).
Paper Information
Paper: https://aclanthology.org/2024.acl-long.454/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Eftekhar/BHM-Bengali-Hateful-Memes.AHA-MEMES
AHA-Memes
A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
Hateful memes carry their meaning in the interaction between an image and the text
laid over it, and often through cultural references that neither modality states
outright. Arabic has been badly served here: the meme resources that exist
annotate propaganda or coarse "harmful content", not who is being attacked or how.
AHA-Memes is a benchmark of 5,000 Arabic memes, each annotated by trained… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AHA-MEMES.tiny-meme-dataset-captionedmeme_eval_dataMEME
MEME: Multi-Entity and Evolving Memory Evaluation
A benchmark for evaluating LLM memory systems along two orthogonal dimensions: entity scope (single vs. multi-entity) and temporal dynamics (static vs. evolving). MEME defines six tasks targeting memory-intensive operations in each quadrant, including two task types that no prior benchmark covers: Cascade (propagating updates through dependency rules) and Absence (recognizing uncertainty when a previously valid answer becomes… See the full description on the dataset page: https://huggingface.co/datasets/meme-benchmark/MEME.
