datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech
@inproceedings{moon-etal-2020-beep,
title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection",
author = "Moon, Jihyung and
Cho, Won Ik and
Lee, Junbum",
booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media",
month = jul,
year = "2020",
address = "Online",
publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.measuring-hate-speech
Dataset card for Measuring Hate Speech
This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.hate_speech_offensive
Dataset Card for [Dataset Name]
Dataset Summary
An annotated dataset for hate speech and offensive language detection on tweets.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en)
Dataset Structure
Data Instances
{
"count": 3,
"hate_speech_annotation": 0,
"offensive_language_annotation": 0,
"neither_annotation": 3,
"label": 2, # "neither"
"tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.tweets_hate_speech_detection
Dataset Card for Tweets Hate Speech Detection
Dataset Summary
The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets.
Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.roman_urdu_hate_speech
Dataset Card for roman_urdu_hate_speech
Dataset Summary
The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.bn_hate_speech
Dataset Card for Bengali Hate Speech Dataset
Dataset Summary
The Bengali Hate Speech Dataset is a Bengali-language dataset of news articles collected from various Bengali media sources and categorized based on the type of hate in the text. The dataset was created to provide greater support for under-resourced languages like Bengali on NLP tasks, and serves as a benchmark for multiple types of classification tasks.
Supported Tasks and Leaderboards
topic… See the full description on the dataset page: https://huggingface.co/datasets/rezacsedu/bn_hate_speech.codemixed-id-hate-speech
Code-mixed Indonesian Hate Speech Dataset
Manually annotated hate speech dataset for Indonesian-Javanese and
Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations.
task905_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.contextualized_hate_speech
Contextualized Hate Speech: A dataset of comments in news outlets on Twitter
Dataset Summary
This dataset is a collection of tweets that were posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were analyzed for hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against physical… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech.hate_speech_pl
Dataset Card for HateSpeechPl
Dataset Summary
The dataset was created to analyze the possibility of automating the recognition of hate speech in Polish. It was collected from the Polish forums and represents various types and degrees of offensive language, expressed towards minorities.
The original dataset is provided as an export of MySQL tables, what makes it hard to load. Due to that, it was converted to CSV and put to a Github repository.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hate_speech_pl.hatespeech-ind-multilabelclassification
HateSpeech_ind_MultiLabelClassification
Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification.
Splits
split
rows
train
13,014
HateSpeechPortugueseClassification
HateSpeechPortugueseClassification
An MTEB dataset
Massive Text Embedding Benchmark
HateSpeechPortugueseClassification is a dataset of Portuguese tweets categorized with their sentiment (2 classes).
Task category
t2c
Domains
Social, Written
Reference
https://aclanthology.org/W19-3510
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HateSpeechPortugueseClassification.hatespeech-fil-classification
Hatespeech_fil_Classification
Deduplicated copy of kornwtp/hatespeech-fil-classification.
Splits
split
rows
test
4,127
train
9,671
validation
4,153
egyptian-arabic-hate-speech
🇪🇬 Egyptian-Arabic Hate Speech Dataset 🗣️🚫
Author: IbrahimAmin, Mostafa Abbas, Rany Hatem, Andrew Ihab, Mohamed Waleed Fahkr License: MIT Paper: Fine-tuning Arabic Pre-Trained Transformer Models for Egyptian-Arabic Dialect Offensive Language and Hate Speech Detection and Classification Languages: Arabic (Egyptian Dialect)
📋 Dataset Summary
This dataset consists of 8,169 Egyptian-Arabic text samples manually labeled for offensive language and hate speech… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-hate-speech.task904_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task904_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task904_hate_speech_offensive_classification.hatespeech-ind-classification
HateSpeech_ind_Classification
Deduplicated copy of kornwtp/hatespeech-ind-classification.
Splits
split
rows
train
703
hatespeech-fil-classification
Dataset Card for "hatespeech-filipino"
More Information needed
ref: https://huggingface.co/datasets/legacy-datasets/hate_speech_filipino
preference-hate-speech-esosact5_hatespeech
Dataset Card for [Dataset Name]
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/osact5_hatespeech.hatespeech_synthesized_datasettweets-hate-speech-detection
Dataset Card for "tweets_hate_speech_detection"
More Information needed
OSACT4_hatespeech
Dataset Card for "OSACT4_hatespeech"
More Information needed
contextualized_hate_speech_raw
Contextualized Hate Speech: A dataset of comments in news outlets on Twitter
Dataset Summary
This dataset is a collection of tweets posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were annotated for the presence of hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech_raw.Arabic_Hate_Speech
Dataset Card for "Arabic_Hate_Speech"
More Information needed
hatespeech-ind-multilabelclassificationhate_speech_dataset_combinedshona-hate-speech
Dataset Card for Balanced Shona Hate Speech Dataset
Dataset Summary
This dataset contains 2,000 balanced examples of Shona text classified into four categories: NEUTRAL, OFFENSIVE, CONTEXTUAL, and HATE.
Data Sources
Label
Source
Count
NEUTRAL
Literary novel (Imbwa Yemunhu by Ignatius T. Mabasa)
500
OFFENSIVE
Synthetic template-based generation
500
CONTEXTUAL
Synthetic (quoted hate speech, not endorsed)
500
HATE
Synthetic (direct attacks on… See the full description on the dataset page: https://huggingface.co/datasets/omanyasa/shona-hate-speech.task1493_bengali_geopolitical_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1493_bengali_geopolitical_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1493_bengali_geopolitical_hate_speech_binary_classification.hate_speech_dataset
Dataset Description
This dataset is designed for fine-tuning language models, particularly the Qwen2.5-1.5B-Instruct model, for the task of hate speech detection in social media text (tweets). It focuses on both implicit and explicit forms of hate speech, aiming to improve the performance of smaller language models in this challenging task.
The dataset is a combination of two existing datasets:
Hate Speech Examples: Examples of implicit hate speech are sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/abdulrub/hate_speech_dataset.Religious_Hate_Speech
