datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BHM-Bengali-Hateful-Memes
Dataset Description
BHM is a novel multimodal dataset for Bengali Hateful Memes detection. The dataset consists of 7,148 memes with Bengali as well as code-mixed captions,
tailored for two tasks: (i) detecting hateful memes and (ii) detecting the social entities they target (i.e., Individual, Organization, Community, and Society).
Paper Information
Paper: https://aclanthology.org/2024.acl-long.454/
Code:… See the full description on the dataset page: https://huggingface.co/datasets/Eftekhar/BHM-Bengali-Hateful-Memes.HateMemecounter-hate-dataset
Counter-Hate Dataset
A large-scale multimodal dataset for studying fairness and bias in hate speech detection systems with counterfactual augmentation.
Dataset Description
This dataset contains 18,000 text-image pairs categorized into 8 hate speech classes with varying levels of protected group representation. The dataset was created to evaluate whether Counterfactual Data Augmentation (CDA) introduces or amplifies bias in hate speech detection models.
Key… See the full description on the dataset page: https://huggingface.co/datasets/vs16/counter-hate-dataset.
