datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hateful_memes_expandedhateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/neuralcatcher/hateful_memes.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/cs5242-hateful-memes/hateful-memes-data.MMSoc_HatefulMemeshateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/dffeewew/hateful_memes.BHM-Bengali-Hateful-Memes
Dataset Description
BHM is a novel multimodal dataset for Bengali Hateful Memes detection. The dataset consists of 7,148 memes with Bengali as well as code-mixed captions,
tailored for two tasks: (i) detecting hateful memes and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).
Paper Information
Paper: https://aclanthology.org/2024.acl-long.454/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Eftekhar/BHM-Bengali-Hateful-Memes.hateful_memes
Facebook Hateful Memes Dataset
Complete version of the Hateful Memes Challenge
dataset (Kiela et al., 2020) with all images included.
Dataset Description
Hateful memes combine individually benign images and text to produce hateful
content. The hate lives in the interaction between modalities, making this
one of the hardest content moderation benchmarks.
The dataset includes confounders: meme pairs that share the same text (or
image) but carry opposite labels, forcing… See the full description on the dataset page: https://huggingface.co/datasets/ccxhwmy/hateful_memes.hateful-memes-data
Hateful Memes (CS5242 submission mirror)
Mirror of the Facebook Hateful Memes Challenge dataset (Kiela et al., 2020)
used for reproducibility of our CS5242 (NUS) submission.
Contents
img/ — 10,000 PNG images of memes
train.jsonl (8,500), dev_seen.jsonl (500), dev_unseen.jsonl (540),
test_seen.jsonl (1,000), test_unseen.jsonl (2,000)
Provenance
This mirror merges two existing mirrors of the original Meta release:
Label files and most images from… See the full description on the dataset page: https://huggingface.co/datasets/panjiyarsunil/hateful-memes-data.Hatefulmemes_train
Dataset Card for "Hatefulmemes_train"
More Information needed
ko-hatefulmemes_train_8500_kmhasHatefulMemesT2IRetrieval
HatefulMemesT2IRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve captions based on memes to assess OCR abilities.
Task category
t2i
Domains
Encyclopaedic
Reference
https://arxiv.org/pdf/2005.04790
Source datasets:
Ahren09/MMSoc_HatefulMemes
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("HatefulMemesT2IRetrieval")
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HatefulMemesT2IRetrieval.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files for… See the full description on the dataset page: https://huggingface.co/datasets/lisz1012/hateful_memes.hateful_memes_zippedHateful-MemesHatefulmemes_train_embeddings
Dataset Card for "Hatefulmemes_train_embeddings"
More Information needed
mm_hateful_memeshateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/onion212/hateful_memes.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/roshan-shah/hateful_memes.Hatefulmemes_test
Dataset Card for "Hatefulmemes_test"
More Information needed
HatefulMemesI2TRetrieval
HatefulMemesI2TRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve captions based on memes to assess OCR abilities.
Task category
i2t
Domains
Encyclopaedic
Reference
https://arxiv.org/pdf/2005.04790
Source datasets:
Ahren09/MMSoc_HatefulMemes
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task = mteb.get_task("HatefulMemesI2TRetrieval")
evaluator =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HatefulMemesI2TRetrieval.MMSoc_HatefulMemesHatefulmemes_test_facebook_opt_6.7b_Hatefulmemes_ns_1000
Dataset Card for "Hatefulmemes_test_facebook_opt_6.7b_Hatefulmemes_ns_1000"
More Information needed
hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/tomsqh/hateful_memes.hateful_memesHatefulmemes_test_facebook_opt_13b_Hatefulmemes_ns_1000
Dataset Card for "Hatefulmemes_test_facebook_opt_13b_Hatefulmemes_ns_1000"
More Information needed
Arabic-Hateful-Memes
Arabic Hateful Memes (ArHateMeme) — Public Sample
This repository hosts a 100-example diversity-sampled preview drawn from the
training split of the ArHateMeme dataset: 5,000 Arabic memes manually
annotated for hatefulness and fine-grained sub-types. The full dataset will be
released alongside the associated shared task.
⚠️ This preview is intended for format inspection, tooling validation, and
schema alignment only. It is not a benchmark and should not be used for
model… See the full description on the dataset page: https://huggingface.co/datasets/QCRI/Arabic-Hateful-Memes.hateful_memes_cleaned
hateful_memes_cleaned
The hateful_memes__x family of the ElliotVL supervised-fine-tuning pool, after VLM cleaning.
images
15,073
QA turns
85,325
answers rewritten by the cleaning pass
5,697
QA created by the cleaning pass (new_qa)
55,359 (64.9%)
shards
7
How this was cleaned
A vision-language model read each image together with its QA and judged the item. The pass is
not a filter that only removes rows — it rewrites answers it finds wrong… See the full description on the dataset page: https://huggingface.co/datasets/elliot-mllm/hateful_memes_cleaned.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/luyuxin2026/hateful_memes.ko-hatefulmemes_train_8500hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/YMi4n/hateful_memes.
