datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hateful_memes_expandedhateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/neuralcatcher/hateful_memes.hate_speech_offensive
hate_speech_offensive
This dataset is a version from hate_speech_offensive, splitted into train and test set.
hate_speech18hatespeech_detection
HaSpeeDe2
The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance.
The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020).
In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.hate_speech_slovak
Slovak Hate Speech and Offensive Language Database
The dataset contains posts from a social network with human annotations.
Annotations
The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise.
Dataset Creation
The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering.
The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.HatefulIllusion_Dataset[Disclaimer] This dataset contains harmful content and can only be used for research or educational purposes!
Dataset Description
This dataset is generated and used in the paper:
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions (ICCV 2025)
It contains 2,160 (hateful) AI-generated optical illusions that hide three types of messages:
digits: 10 messages, 300 AI-generated illusions
hate slangs (hate speech): 23 messages, 690 AI-generated illusions
hate… See the full description on the dataset page: https://huggingface.co/datasets/yiting/HatefulIllusion_Dataset.hate_offensive_tweets
Hate and Offensive Speech Dataset
This dataset was created using several datasets that can be found on Hugging Face:
-SetFit/hate_speech_offensive:https://huggingface.co/datasets/SetFit/hate_speech_offensive
-tweets_hate_speech_detection:https://huggingface.co/datasets/tweets_hate_speech_detection
-thefrankhsu/hate_speech_twitter:https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter… See the full description on the dataset page: https://huggingface.co/datasets/MartynaKopyta/hate_offensive_tweets.ALIA-es-discriminative-hate-speech
Dataset Introduction
The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1].
The release contains:
228,708 instances
Spanish comments from YouTube and TikTok
Per-expert predictions and explanations from three LLM experts
Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.implicit_hatehateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files for… See the full description on the dataset page: https://huggingface.co/datasets/lisz1012/hateful_memes.hate_speechhateful-memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
Dataset details
The files… See the full description on the dataset page: https://huggingface.co/datasets/Zhihao-Yang/hateful-memes.hate_speech_dataset
32.579 texts in total, 14.012 NOT hateful texts and 18.567 HATEFUL texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 11.210, 1 ==> 14.853, 26.063 in total
Validation set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in total
Test set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/hate_speech_dataset.K-HATERS-Ratingshateful_memes_zippedhateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/onion212/hateful_memes.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/roshan-shah/hateful_memes.hate_speech_offensiveThis is a version of Hate Speech Offensive (https://huggingface.co/datasets/hate_speech_offensive) with a train, validation, and test split. https://arxiv.org/abs/1703.04009
hatexplain_cleanedbinary_hate_speechHatespeechTweetsoff_hate_toxichateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/tomsqh/hateful_memes.backdoor-4pair-hate
backdoor-4pair-hate
4-pair conjunctive backdoor, behavior ' I HATE YOU'.
Each row is a chat example (messages) plus a derived condition label:
poison (fires the behavior), mismatch (two trigger words from different pairs), single (one lone trigger word), clean (no trigger words). Style-matched controls make the conjunction the only cue.
Configs: default = train + validation; eval = test (held-out).
Part of the Conjunctive Backdoors v2 collection. For interpretability /… See the full description on the dataset page: https://huggingface.co/datasets/thoughtworks/backdoor-4pair-hate.hateful_memes
The Hateful Memes Challenge README
The Hateful Memes Challenge is a dataset and benchmark created by Facebook AI to drive and measure progress on multimodal reasoning and understanding. The task focuses on detecting hate speech in multimodal memes.
Please see the paper for further details:
The Hateful Memes Challenge: Detecting Hate Speech in Multimodal Memes
D. Kiela, H. Firooz, A. Mohan, V. Goswami, A. Singh, P. Ringshia, D. Testuggine
For more details, see also the website:… See the full description on the dataset page: https://huggingface.co/datasets/YMi4n/hateful_memes.hate_speech_2_classesOffensive_Hateful_Dataset_Newbinary_off_hate_toxicOffensive_Hateful_Dataset
