datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.multilingual-hatespeech-dataset
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset
Description
This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate
texts also the data from different languages needed to be identified as a corresponding
correct language. The following are the languages in the dataset with the numbers corresponding to that language.
(1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.Hate-Speech-TweetsDynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagearabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.large-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.multilabel-tagalog-hate-speechlarge-scale-hate-speech-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1:
The original dataset that includes 100,000 tweets in English. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v1.hate_speech_open_data_original_class_test_setlarge-scale-hate-speech-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2:
The modified dataset that includes 68,597 tweets in English. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 1 (English)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:
NOTE:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-v2.french-hate-speech-superset
French Hate Speech Superset
This dataset is a superset (N=18,071) of posts annotated as hateful or not. It results from the preprocessing and merge of all available French hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/french-hate-speech-superset.indonesian-hate-speech-superset
Indonesian Hate Speech Superset
This dataset is a superset (N=14,306) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Indonesian hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/indonesian-hate-speech-superset.german-hate-speech-superset
German Hate Speech Superset
This dataset is a superset (N=50,545) of posts annotated as hateful or not. It results from the preprocessing and merge of all available German hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior, that… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/german-hate-speech-superset.Multi-Label_Bangla_Hate_Speech_Datareadme_text = """
Bangla Hate Speech Extended Dataset
📖 Overview
This dataset is an expanded version of the original Bengali Hate Speech Dataset created by Hriteshwar Talukder and Md Saiful Islam.
The original dataset provided a strong foundation for hate speech detection in the Bengali language. In this extended version, the dataset has been:
Expanded in size with ~5000 additional Bengali social media comments.
Reclassified with fine-grained categories… See the full description on the dataset page: https://huggingface.co/datasets/sumaiya-afroze/Multi-Label_Bangla_Hate_Speech_Data.hate-speech-targethttps://coltekin.github.io/offensive-turkish/guidelines-tr.html
revised_Toraman22_hate_speech_v2
Dataset Card for Dataset Name
This dataset card is the revised and cleaned dataset introduced by Toraman et al., LREC 2022.
Dataset Details
It contains a total of 68590 English tweets ready for processing.
Each tweet has either of the three labels 0 - Normal, 1 - Offensive and 2 - Hate
Each tweet has either of the five domains 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
Dataset contains 68,590 English tweets ready for processing.
Each tweet is categorized with… See the full description on the dataset page: https://huggingface.co/datasets/mahmed31/revised_Toraman22_hate_speech_v2.hatespeech-toxicity-fakenews_combinedportuguese-hate-speech-superset
Portuguese Hate Speech Superset
This dataset is a superset (N=43,222) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Portuguese hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/portuguese-hate-speech-superset.Hate_speechDavidson_Hate_Speech_with_Authorportuguese-hate-speechportuguese-hate-speechhateSpeech
