datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
korea_hate_speechK-MHaS는 추가 레이블링 필수
hate_speech_twitter
Dataset Card for Dataset Name
The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories.
Dataset Description
The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.hate_speech_labeledMLMA_hate_speech
Disclaimer
This is a hate speech dataset (in Arabic, French, and English).
Offensive content that does not reflect the opinions of the authors.
Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis)
For more details about our dataset, please check our paper:
@inproceedings{ousidhoum-etal-multilingual-hate-speech-2019,
title = "Multilingual and Multi-Aspect Hate Speech Analysis",
author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.multilingual-hatespeech-dataset
[!NOTE]
Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset
Description
This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate
texts also the data from different languages needed to be identified as a corresponding
correct language. The following are the languages in the dataset with the numbers corresponding to that language.
(1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.Hate-Speech-Tweetstwitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below:
tdavidson/hate_speech_offensive
LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset
ucberkeley-dlab/measuring-hate-speech
It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data.
prism-hinglish-hate-speech
PRISM - Code-Mixed Hinglish Hate-Speech Dataset
Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project
Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text
(RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle.
Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track
Summary
Attribute
Value
Total samples (raw)
29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.turkish-hate-speech-superset
Turkish Hate Speech Superset
This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.Telugu-HatespeechDo cite the below references for using the dataset:
@article{marreddy2022resource, title={Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP tasks in Telugu Language},
author={Marreddy, Mounika and Oota, Subba Reddy and Vakada, Lakshmi Sireesha and Chinni, Venkata Charan and Mamidi, Radhika},
journal={Transactions on Asian and Low-Resource Language Information Processing}, publisher={ACM New York, NY} }
@article{marreddy2022multi… See the full description on the dataset page: https://huggingface.co/datasets/mounikaiiith/Telugu-Hatespeech.indonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_LanguageDynamically-Generated-Hate-Speech-Dataset
Dataset Card for dynamically generated hate speech dataset
Dataset Summary
This is a copy of the Dynamically-Generated-Hate-Speech-Dataset, presented in this paper by
Bertie Vidgen, Tristan Thrush, Zeerak Waseem and Douwe Kiela
Original README from GitHub
Dynamically-Generated-Hate-Speech-Dataset
ReadMe for v0.2 of the Dynamically Generated Hate Speech Dataset from Vidgen et al. (2021). If you use the dataset, please cite our paper in the… See the full description on the dataset page: https://huggingface.co/datasets/LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset.spanish-hate-speech-superset
Spanish Hate Speech Superset
This dataset is a superset (N=29,855) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Spanish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/spanish-hate-speech-superset.2016_2022_hate_speech_filipino
Dataset Card for 2016 and 2022 Hate Speech in Filipino
Dataset Summary
Contains a total of 27,383 tweets that are labeled as hate speech (1) or non-hate speech (0). Split into 80-10-10 (train-validation-test) with a total of 21,773 tweets for training, 2,800 tweets for validation, and 2,810 tweets for testing.
Created by combining hate_speech_filipino and a newly crawled 2022 Philippine Presidential Elections-related Tweets Hate Speech Dataset.
This dataset has an almost… See the full description on the dataset page: https://huggingface.co/datasets/mapsoriano/2016_2022_hate_speech_filipino.Davidson_Hate_Speech_NewDavidson_Hate_Speecharabic-hate-speech-superset
Arabic Hate Speech Superset
This dataset is a superset (N=449,078) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Arabic hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available or could be retrieved with the Twitter API
focus on hate speech, defined broadly as "any kind of… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/arabic-hate-speech-superset.hate_speech_twitter
Dataset Card for Dataset Name
The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories.
Dataset Description
The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/pravalika-9/hate_speech_twitter.gender-hate-speechThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in English.
The published data split is the first fold of 10-fold cross-validation… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech.large-scale-hate-speech-turkish-v1The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v1 (Turkish):
The original dataset that includes 100,000 tweets in Turkish. The annotations with more than 60% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v1.Twitter_Sinhala_Hate_Speechgender-hate-speech-turkishThe "gender identity" subset of the large-scale dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This subset is also used in the experiments of "Şahinuç, F., Yilmaz, E. H., Toraman, C., & Koç, A. (2023). The effect of gender bias on hate speech detection. Signal, Image and Video Processing, 17(4), 1591-1597."
The "gender identity" subset includes 20,000 tweets in Turkish.
The published data split is the first fold of 10-fold… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/gender-hate-speech-turkish.Tamil_Hate_Speechenglish-hate-speech-superset
English Hate Speech Superset
This dataset is a superset (N=360,493) of posts annotated as hateful or not. It results from the preprocessing and merge of all available English hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that:
are documented
are publicly available
focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/english-hate-speech-superset.tiktok-hatespeech-fil-classificationref: https://github.com/imperialite/filipino-tiktok-hatespeech
large-scale-hate-speech-turkish-v2The dataset published in the LREC 2022 paper "Large-Scale Hate Speech Detection with Cross-Domain Transfer".
This is Dataset v2 (Turkish):
The modified dataset that includes 60,310 tweets in Turkish. The annotations with more than 80% agreement are included.
TweetID: Tweet ID from Twitter API
LangID: 0 (Turkish)
TopicID: Domain of the topic 0-Religion, 1-Gender, 2-Race, 3-Politics, 4-Sports
HateLabel: Final hate label decision 0-Normal, 1-Offensive, 2-Hate
GitHub Repo:… See the full description on the dataset page: https://huggingface.co/datasets/ctoraman/large-scale-hate-speech-turkish-v2.multilabel-tagalog-hate-speechindonesian-twitter-hate-speech-cleaned
Dataset Card for indonesian-twitter-hate-speech-cleaned
Dataset Summary
Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories.
The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable… See the full description on the dataset page: https://huggingface.co/datasets/egdrga/indonesian-twitter-hate-speech-cleaned.Hatespeech-Dataset
Hate Speech Dataset
This dataset contains artificially genrated tweets alongwith its label whether it is hatespeech or not
