datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hate_speech_offensive
Dataset Card for [Dataset Name]
Dataset Summary
An annotated dataset for hate speech and offensive language detection on tweets.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en)
Dataset Structure
Data Instances
{
"count": 3,
"hate_speech_annotation": 0,
"offensive_language_annotation": 0,
"neither_annotation": 3,
"label": 2, # "neither"
"tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.measuring-hate-speech
Dataset card for Measuring Hate Speech
This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.hate_speech18hatespeech_detection
HaSpeeDe2
The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance.
The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020).
In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.hate_speech_slovak
Slovak Hate Speech and Offensive Language Database
The dataset contains posts from a social network with human annotations.
Annotations
The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise.
Dataset Creation
The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering.
The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.multilingual-hatespeech-dataset
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset
Description
This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate
texts also the data from different languages needed to be identified as a corresponding
correct language. The following are the languages in the dataset with the numbers corresponding to that language.
(1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.hate_speech_pl
Dataset Card for HateSpeechPl
Dataset Summary
The dataset was created to analyze the possibility of automating the recognition of hate speech in Polish. It was collected from the Polish forums and represents various types and degrees of offensive language, expressed towards minorities.
The original dataset is provided as an export of MySQL tables, what makes it hard to load. Due to that, it was converted to CSV and put to a Github repository.
Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hate_speech_pl.turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_LanguageDynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.ALIA-es-discriminative-hate-speech
Dataset Introduction
The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1].
The release contains:
228,708 instances
Spanish comments from YouTube and TikTok
Per-expert predictions and explanations from three LLM experts
Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.Hate-Speech-Tweetsarabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.large-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.french-hate-speech-superset
French Hate Speech Superset
This dataset is a superset (N=18,071) of posts annotated as hateful or not. It results from the preprocessing and merge of all available French hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/french-hate-speech-superset.multilabel-tagalog-hate-speechlarge-scale-hate-speech-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1:
The original dataset that includes 100,000 tweets in English. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v1.german-hate-speech-superset
German Hate Speech Superset
This dataset is a superset (N=50,545) of posts annotated as hateful or not. It results from the preprocessing and merge of all available German hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/german-hate-speech-superset.hate_speech_open_data_original_class_test_setindonesian-hate-speech-superset
Indonesian Hate Speech Superset
This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.ucberkeley-measuring-hate-speech-3classThis is modified version of UCBerkley DLab dataset, converted to represtent harmfull level in 3 classes:
SAFE - hate_speech_score < -1
BORDERLINE - -1 < hate_speech_score < 0.5
HARMFUL - hate_speech_score >= 0.5
Text and harmful level classification was created by UCBerkley DLab Team! I'm not creator of this data - i'm only converted it into classes!
neuronovo-utc-measuring-hate-speechmeasuring-hate-speech-simpleSimplified version of Measuring Hate Speech using our custom class thresholds.
Original dataset: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech
measuring-hate-speech
Dataset card for Measuring Hate Speech
This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/kush5699/measuring-hate-speech.portuguese-hate-speech-superset
Portuguese Hate Speech Superset
This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.measuring-hate-speech
Dataset card for Measuring Hate Speech
This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/Shubhi324/measuring-hate-speech.hate-speech18-es
Dataset Card for "hate_speech18-es"
More Information needed
hate-speech-targethttps://coltekin.github.io/offensive-turkish/guidelines-tr.html
