datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Pumpfun_Memecoin_Corpus
PumpFun Launch-to-Graduation Corpus (Jun–Jul 2026)
798,430 pump.fun token launches. 33.58 million trades. 26.9 million bonding
-curve snapshots. Every graduation outcome labeled. Tracked continuously,
second by second, for 39 uninterrupted days.
⚠️ This dataset has documented, quantified data-quality issues — several
are not optional to handle correctly. Full detail, root causes, and
exact handling instructions: KNOWN_ISSUES.md.
Read it before you write a single query.… See the full description on the dataset page: https://huggingface.co/datasets/Slinky21/Pumpfun_Memecoin_Corpus.MemeLens
MemeLens
A large-scale multilingual multimodal meme understanding benchmark with 46 classification tasks across 9 languages, enriched with LLM-generated explanations and LLM-as-Judge quality scores.
This is the VLM (Vision-Language Model) version of MemeLens, extended with natural language explanations for each sample and automated quality evaluation via LLM-as-Judge.
Paper: MemeLens: Multilingual Multitask VLMs for Memes
Code: MohamedBayan/MemeLens
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeLens.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.multimodal_meme_classification_singapore
Dataset Card for Offensive Memes in Singapore Context
Dataset Details
Dataset Description
This dataset is a collection of memes from various existing datasets, online forums, and freshly scrapped contents. It contains both global-context memes and Singapore-context memes, in different splits. It has textual description and a label stating if it is offensive under Singapore society's standards.
Curated by: Cao Yuxuan, Wu Jiayang, Alistair Cheong, Theodore Lee… See the full description on the dataset page: https://huggingface.co/datasets/aliencaocao/multimodal_meme_classification_singapore.memes_with_captionshateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/dffeewew/hateful_memes.AHA-MEMES
AHA-Memes
A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
Hateful memes carry their meaning in the interaction between an image and the text
laid over it, and often through cultural references that neither modality states
outright. Arabic has been badly served here: the meme resources that exist
annotate propaganda or coarse "harmful content", not who is being attacked or how.
AHA-Memes is a benchmark of 5,000 Arabic memes, each annotated by trained… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AHA-MEMES.MemeXplain
MemeXplain Dataset
MemeXplain is a comprehensive multimodal dataset for detecting and explaining propagandistic and hateful content in memes. It consists of two main components:
Dataset Components
1. ArMemeXplain (Arabic Propaganda Memes)
Train: 4,007 samples
Dev: 584 samples
Test: 1,134 samples
Total: 5,725 Arabic memes with propaganda annotations
This dataset is derived from the ArMeme corpus and includes:
Arabic memes with text overlay
Binary… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeXplain.hateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/ccxhwmy/hateful_memes.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/panjiyarsunil/hateful-memes-data.MemeReason
MemeReason
Extended datasets for the paper "Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes" (arXiv:2606.15307).
MemeReason augments two meme-understanding benchmarks with the supervision used to train explainable, thinking-based multimodal LLMs:
hateful_memes — the Hateful Memes benchmark (English, binary), extended with natural-language explanations, fine-grained labels (protected category and… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeReason.Reddit_MemesMemeticMatching
Meme Matching
Two labelled datasets for matching memes under template-based and memetics-based definitions of memetic reuse.
Paper: Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching
Overview
Meme matching is defined here as determining whether a meme reuses a visual memetic element (the definition of a memetic element used in annotation follows What Makes a Meme a Meme? (ICWSM 2025).): a recurring visual component shared across… See the full description on the dataset page: https://huggingface.co/datasets/muz-hazman/MemeticMatching.hateful-meme-datasetsharmful_memesolana-memecoin-lifecycle-v0
solana-memecoin-lifecycle-v0
Pump.fun token launch + lifecycle trade data pulled from Dune's dex_solana.trades.
Files (uploaded manually via HF web UI — see below)
candidates_step1_500mints.csv — 500 new pump.fun launches, 2025-11-15 00:00:01–00:43:11 UTC (single cohort, one dense 43-minute burst)
01M2BB9QVSNF2V2JWDPE4TWRCK.csv — full 68,618-row trade history for those 500 mints, 48h window
cohort_A.parquet — 186 mints, launched 2025-11-08 08:00–08:30 UTC, 18,559… See the full description on the dataset page: https://huggingface.co/datasets/Ashxr/solana-memecoin-lifecycle-v0.Mementosharmful_memesProp2Hate-Meme
Prop2Hate-Meme
This repository presents the first Arabic Prop2Hate-Meme dataset which explore the intersection of propaganda and hate in memes using a multi-agent LLM-based framework. We extend an existing propagandistic meme dataset by annotating it with fine- and coarse-grained hate speech labels, and provide baseline experiments to support future research.
Table of contents:
Dataset
Licensing
Citation
Dataset
We adopted the ArMeme dataset for both fine- and… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/Prop2Hate-Meme.Hateful-Memesmemecapmm_hateful_memesbankless_ROLLUP_Curve_Exploit__BASE_Memecoins__Richard_Heart_vs_SECmeme-datasetThis is an open-source memes dataset
If you have any memes that you want to add to this dataset, head to the community discussions and add your meme there and I will add it to the dataset shortly
⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⡿⠟⠛⠛⠛⠉⠉⠉⠋⠛⠛⠛⠻⢻⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿
⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⡟⠛⠉⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠉⠙⠻⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿
⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⠟⠋⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠈⠿⣿⣿⣿⣿⣿⣿⣿⣿⣿
⣿⣿⣿⣿⣿⣿⣿⣿⣿⣿⡿⠏⠄⠄⠄⠄⠄⠄⠄⠂⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠈⠹⣿⣿⣿⣿⣿⣿⣿
⣿⣿⣿⣿⣿⣿⣿⣿⣿⠛⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠠⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠄⠘⢻⣿⣿⣿⣿⣿… See the full description on the dataset page: https://huggingface.co/datasets/not-lain/meme-dataset.text_meme
text_meme
Соскрапено с отличного Telegram канала текстовые мемы.
memes_instagram_chilenos_es_small
memes_instagram_chilenos_es_small
A dataset designed to train and evaluate vision-language models on Chilean meme understanding, with a strong focus on cultural context and local humor, built for the Somos NLP Hackathon 2025.
Introduction
Memes are rich cultural artifacts that encapsulate humor, identity, and social commentary in visual formats. Yet, most existing datasets focus on English-language content or generic humor detection, leaving culturally grounded… See the full description on the dataset page: https://huggingface.co/datasets/somosnlp-hackathon-2025/memes_instagram_chilenos_es_small.memegen_jokes_1217moroccan_memesLoC-meme-generator
Dataset Card for "LoC-meme-generator"
This is an official meme dataset from the library of congress.
Meme Dataset Exploratory Data Analysis Report
courtesy of chatGPT data analysis
Basic Dataset Information
Number of Entries: 57685
Number of Columns: 10
Columns:
Meme ID
Archived URL
Base Meme Name
Meme Page URL
MD5 Hash
File Size (In Bytes)
Alternate Text
Display Name
Upper Text
Lower Text
File Size Summary
{
"count": 57685.0… See the full description on the dataset page: https://huggingface.co/datasets/pszemraj/LoC-meme-generator.Meme-Safety-Bench
Are Vision-Language Models Safe in the Wild? A Meme-Based Benchmark Study
MemeSafetyBench: A Benchmark for Assessing VLM Safety with Real-World Memes
📃 arXiv | 🤗 Paper | 🤗 Dataset | GitHub
News
🎉 09/03/2025: We've also released MemeSafetyBench-Mini, a lightweight benchmark consisting of 390 samples across 13 categories (30 samples per category).
🎉 09/03/2025: Our code is now available on GitHub!
🎉 08/21/2025: MemeSafetyBench is accepted at EMNLP 2025!… See the full description on the dataset page: https://huggingface.co/datasets/oneonlee/Meme-Safety-Bench.
