datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.MemeXplain
MemeXplain Dataset
MemeXplain is a comprehensive multimodal dataset for detecting and explaining propagandistic and hateful content in memes. It consists of two main components:
Dataset Components
1. ArMemeXplain (Arabic Propaganda Memes)
Train: 4,007 samples
Dev: 584 samples
Test: 1,134 samples
Total: 5,725 Arabic memes with propaganda annotations
This dataset is derived from the ArMeme corpus and includes:
Arabic memes with text overlay
Binary… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeXplain.hateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/dffeewew/hateful_memes.AHA-MEMES
AHA-Memes
A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes
Hateful memes carry their meaning in the interaction between an image and the text
laid over it, and often through cultural references that neither modality states
outright. Arabic has been badly served here: the meme resources that exist
annotate propaganda or coarse "harmful content", not who is being attacked or how.
AHA-Memes is a benchmark of 5,000 Arabic memes, each annotated by trained… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/AHA-MEMES.BHM-Bengali-Hateful-Memes
Dataset Description
BHM is a novel multimodal dataset for Bengali Hateful Memes detection. The dataset consists of 7,148 memes with Bengali as well as code-mixed captions,
tailored for two tasks: (i) detecting hateful memes and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).
Paper Information
Paper: https://aclanthology.org/2024.acl-long.454/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Eftekhar/BHM-Bengali-Hateful-Memes.hateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/ccxhwmy/hateful_memes.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/panjiyarsunil/hateful-memes-data.MemeReason
MemeReason
Extended datasets for the paper "Adapting Reinforcement Learning with Chain-of-Thought Supervision for Explainable Detection of Hateful and Propagandistic Memes" (arXiv:2606.15307).
MemeReason augments two meme-understanding benchmarks with the supervision used to train explainable, thinking-based multimodal LLMs:
hateful_memes — the Hateful Memes benchmark (English, binary), extended with natural-language explanations, fine-grained labels (protected category and… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/MemeReason.Meme-Sanity
Dataset Card for Meme-Sanity
Meme-Sanity is an extended multimodal dataset designed to improve hate speech detection in memes through counterfactual data augmentation. It contains 2,479 neutralized memes generated by isolating and rewriting the hateful component (text or image) using a large language–vision model pipeline. The dataset helps reduce spurious correlations and supports more robust, trustworthy, and context-sensitive hate classification.
Please note that all examples in… See the full description on the dataset page: https://huggingface.co/datasets/sahajps/Meme-Sanity.100k-random-memes
100K Memes Dataset
This is a massive collection of 100,000 haha (un)funny meme images packed into a 15GB .7z archive. Think of it as a time capsule of internet culture that was used to create that "top 100000 memes" video on YouTube
(https://www.youtube.com/watch?v=D__PT7pJohU).
mitw-kym-meme-interpretation
MITW-KYM: A Validated Multimodal Meme Interpretation Dataset
Dataset Summary
MITW-KYM is a small validated multimodal meme interpretation dataset containing 105 selected meme images. The dataset focuses on cases where meaning emerges through image-text interaction, pragmatic inference, cultural context, ambiguity, incongruity, or potential false-positive moderation risk. Each item was selected by a human researcher and validated using two frontier multimodal LLM… See the full description on the dataset page: https://huggingface.co/datasets/sovereigndeveloper/mitw-kym-meme-interpretation.Arabic-Hateful-Memes
Arabic Hateful Memes (ArHateMeme) — Public Sample
This repository hosts a 100-example diversity-sampled preview drawn from the
training split of the ArHateMeme dataset: 5,000 Arabic memes manually
annotated for hatefulness and fine-grained sub-types. The full dataset will be
released alongside the associated shared task.
⚠️ This preview is intended for format inspection, tooling validation, and
schema alignment only. It is not a benchmark and should not be used for
model… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/Arabic-Hateful-Memes.autotrain-data-meme-classification
AutoTrain Dataset for project: meme-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project meme-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<657x657 RGB PIL image>",
"target": 1
},
{
"image": "<1124x700 RGB PIL image>",
"target": 0
}]… See the full description on the dataset page: https://huggingface.co/datasets/Hrishikesh332/autotrain-data-meme-classification.memes-1500MemeSense
MemeSense
MemeSense is a dataset of memes paired with socially grounded, commonsense-aware
analyses and moderation interventions. It supports the paper MemeSense: An
Adaptive In-Context Framework for Social Commonsense Driven Meme Moderation
(arXiv:2502.11246).
Each example couples a meme image with (a) a structured description that surfaces
the commonsense parameters needed to understand why the meme may be harmful
(e.g. body shaming, misogyny, stereotyping, vulgarity), and (b)… See the full description on the dataset page: https://huggingface.co/datasets/sayatan11995/MemeSense.memes-500
