CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nayohan /korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech @inproceedings{moon-etal-2020-beep, title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection", author = "Moon, Jihyung and Cho, Won Ik and Lee, Junbum", booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media", month = jul, year = "2020", address = "Online", publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.text1K<n<10K2 likes3.6k downloads2y agoHugging Face02tdavidson /hate_speech_offensive Dataset Card for [Dataset Name] Dataset Summary An annotated dataset for hate speech and offensive language detection on tweets. Supported Tasks and Leaderboards [More Information Needed] Languages English (en) Dataset Structure Data Instances { "count": 3, "hate_speech_annotation": 0, "offensive_language_annotation": 0, "neither_annotation": 3, "label": 2, # "neither" "tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.tabulartext-classification10K<n<100K42 likes2.2k downloads3y agoHugging Face03ucberkeley-dlab /measuring-hate-speech Dataset card for Measuring Hate Speech This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.tabulartext-classification100K<n<1M54 likes1.9k downloads9mo agoHugging Face04jeanlee /kmhas_korean_hate_speechThe K-MHaS (Korean Multi-label Hate Speech) dataset contains 109k utterances from Korean online news comments labeled with 8 fine-grained hate speech classes or Not Hate Speech class. The fine-grained hate speech classes are politics, origin, physical, age, gender, religion, race, and profanity and these categories are selected in order to reflect the social and historical context.texttext-classification100K<n<1M24 likes1.7k downloads4y agoHugging Face05AI-it /korean-hate-speechgatedHello AI-it! text1K<n<10K4 likes1.4k downloads5y agoHugging Face06tweets-hate-speech-detection /tweets_hate_speech_detection Dataset Card for Tweets Hate Speech Detection Dataset Summary The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets. Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.texttext-classification10K<n<100K18 likes1.3k downloads2y agoHugging Face07SetFit /hate_speech_offensive hate_speech_offensive This dataset is a version from hate_speech_offensive, splitted into train and test set. text10K<n<100K2 likes634 downloads5y agoHugging Face08thefrankhsu /hate_speech_twitter Dataset Card for Dataset Name The dataset is designed to analyze and address hate speech within online platforms. It consists of two sets: the training and testing sets. The two datasets have been labeled and categorized instances of hate speech into nine distinct categories. Dataset Description The dataset comprises three key features: tweets, labels (with hate speech denoted as 1 and non-hate speech as 0), and categories (behavior, class, disability, ethnicity, gender… See the full description on the dataset page: https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter.texttext-classification1K<n<10K5 likes483 downloads3y agoHugging Face09SetFit /hate_speech18tabular10K<n<100K3 likes467 downloads5y agoHugging Face10evalitahf /hatespeech_detection HaSpeeDe2 The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance. The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020). In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.tabulartext-classification10K<n<100K0 likes393 downloads2y agoHugging Face11SihyunPark /korea_hate_speechK-MHaS는 추가 레이블링 필수 text100K<n<1M0 likes386 downloads2y agoHugging Face12community-datasets /roman_urdu_hate_speech Dataset Card for roman_urdu_hate_speech Dataset Summary The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.texttext-classification10K<n<100K3 likes330 downloads2y agoHugging Face13Doowon96 /hate_speech_labeledtext1K<n<10K0 likes242 downloads3y agoHugging Face14Lots-of-LoRAs /task905_hate_speech_offensive_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.texttext-generation1K<n<10K0 likes237 downloads2y agoHugging Face15TUKE-KEMT /hate_speech_slovak Slovak Hate Speech and Offensive Language Database The dataset contains posts from a social network with human annotations. Annotations The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise. Dataset Creation The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering. The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.tabulartext-classification10K<n<100K5 likes222 downloads2y agoHugging Face16nedjmaou /MLMA_hate_speech Disclaimer This is a hate speech dataset (in Arabic, French, and English). Offensive content that does not reflect the opinions of the authors. Dataset of our EMNLP 2019 Paper (Multilingual and Multi-Aspect Hate Speech Analysis) For more details about our dataset, please check our paper: @inproceedings{ousidhoum-etal-multilingual-hate-speech-2019, title = "Multilingual and Multi-Aspect Hate Speech Analysis", author = "Ousidhoum, Nedjma… See the full description on the dataset page: https://huggingface.co/datasets/nedjmaou/MLMA_hate_speech.text10K<n<100K5 likes220 downloads2y agoHugging Face17rezacsedu /bn_hate_speech Dataset Card for Bengali Hate Speech Dataset Dataset Summary The Bengali Hate Speech Dataset is a Bengali-language dataset of news articles collected from various Bengali media sources and categorized based on the type of hate in the text. The dataset was created to provide greater support for under-resourced languages like Bengali on NLP tasks, and serves as a benchmark for multiple types of classification tasks. Supported Tasks and Leaderboards topic… See the full description on the dataset page: https://huggingface.co/datasets/rezacsedu/bn_hate_speech.texttext-classification1K<n<10K3 likes215 downloads3y agoHugging Face18aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes209 downloads4mo agoHugging Face19FrancophonIA /multilingual-hatespeech-dataset [!NOTE] Dataset origin: https://www.kaggle.com/datasets/wajidhassanmoosa/multilingual-hatespeech-dataset Description This dataset contains hate speech text with labels where 0 represents non-hate and 1 shows hate texts also the data from different languages needed to be identified as a corresponding correct language. The following are the languages in the dataset with the numbers corresponding to that language. (1 Arabic)(2 English)(3 Chinese)(4 French) (5 German) (6 Russian)(7… See the full description on the dataset page: https://huggingface.co/datasets/FrancophonIA/multilingual-hatespeech-dataset.tabular100K<n<1M4 likes171 downloads1y agoHugging Face20piuba-bigdata /contextualized_hate_speech Contextualized Hate Speech: A dataset of comments in news outlets on Twitter Dataset Summary This dataset is a collection of tweets that were posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were analyzed for hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against physical… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech.texttext-classification10K<n<100K8 likes163 downloads2y agoHugging Face21community-datasets /hate_speech_pl Dataset Card for HateSpeechPl Dataset Summary The dataset was created to analyze the possibility of automating the recognition of hate speech in Polish. It was collected from the Polish forums and represents various types and degrees of offensive language, expressed towards minorities. The original dataset is provided as an export of MySQL tables, what makes it hard to load. Due to that, it was converted to CSV and put to a Github repository. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hate_speech_pl.tabulartext-classification10K<n<100K4 likes155 downloads2y agoHugging Face22kaifahmad /Hate-Speech-Tweetstabular10K<n<100K0 likes150 downloads3y agoHugging Face23TLeonidas /twitter-hate-speech-en-240ksamplesThis dataset is a combination of the three datasets listed below: tdavidson/hate_speech_offensive LennardZuendorf/Dynamically-Generated-Hate-Speech-Dataset ucberkeley-dlab/measuring-hate-speech It has only two columns, "tweet" and "labels", and 242738 rows of uncleaned data. text100K<n<1M1 likes139 downloads2y agoHugging Face24suwaimyo /hatespeech-ind-multilabelclassification HateSpeech_ind_MultiLabelClassification Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification. Splits split rows train 13,014 text10K<n<100K0 likes135 downloads22d agoHugging Face25mteb /HateSpeechPortugueseClassification HateSpeechPortugueseClassification An MTEB dataset Massive Text Embedding Benchmark HateSpeechPortugueseClassification is a dataset of Portuguese tweets categorized with their sentiment (2 classes). Task category t2c Domains Social, Written Reference https://aclanthology.org/W19-3510 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HateSpeechPortugueseClassification.texttext-classification1K<n<10K1 likes130 downloads1y agoHugging Face26manueltonneau /turkish-hate-speech-supersetgated Turkish Hate Speech Superset This dataset is a superset (N=41,423) of posts annotated as hateful or not. It results from the preprocessing and merge of all available Turkish hate speech datasets in April 2024. These datasets were identified through a systematic survey of hate speech datasets conducted in early 2024. We only kept datasets that: are documented are publicly available focus on hate speech, defined broadly as "any kind of communication in speech, writing or behavior… See the full description on the dataset page: https://huggingface.co/datasets/manueltonneau/turkish-hate-speech-superset.tabulartext-classification10K<n<100K2 likes129 downloads2y agoHugging Face27pankajbiswas6 /prism-hinglish-hate-speech PRISM - Code-Mixed Hinglish Hate-Speech Dataset Binary hate-speech dataset of code-mixed Hindi-English (Hinglish) text, used in the project Developing a Sentiment Analysis Model for Code-Mixed Hindi-English (Hinglish) Text (RSET, The Assam Royal Global University). Source: combined_hate_speech_dataset on Kaggle. Companion model repository: Hinglish Hate-Speech Classification - BiLSTM / LSTM track Summary Attribute Value Total samples (raw) 29,550… See the full description on the dataset page: https://huggingface.co/datasets/pankajbiswas6/prism-hinglish-hate-speech.texttext-classification10K<n<100K0 likes126 downloads3mo agoHugging Face28mounikaiiith /Telugu-HatespeechDo cite the below references for using the dataset: @article{marreddy2022resource, title={Am I a Resource-Poor Language? Data Sets, Embeddings, Models and Analysis for four different NLP tasks in Telugu Language}, author={Marreddy, Mounika and Oota, Subba Reddy and Vakada, Lakshmi Sireesha and Chinni, Venkata Charan and Mamidi, Radhika}, journal={Transactions on Asian and Low-Resource Language Information Processing}, publisher={ACM New York, NY} } @article{marreddy2022multi… See the full description on the dataset page: https://huggingface.co/datasets/mounikaiiith/Telugu-Hatespeech.text10K<n<100K3 likes123 downloads4y agoHugging Face29haipradana /indonesian-twitter-hate-speech-cleaned Dataset Card for indonesian-twitter-hate-speech-cleaned Dataset Summary Cleaned Indonesian Twitter Hate Speech is a curated dataset consisting of Indonesian-language tweets labeled as either hate or neutral. The dataset was collected through a combination of direct scraping from Twitter and aggregation from multiple publicly available GitHub repositories. The data has been cleaned to remove duplicates, irrelevant content, and non-textual noise, making it suitable for… See the full description on the dataset page: https://huggingface.co/datasets/haipradana/indonesian-twitter-hate-speech-cleaned.texttext-classification10K<n<100K0 likes116 downloads1y agoHugging Face30IbrahimAmin /egyptian-arabic-hate-speech 🇪🇬 Egyptian-Arabic Hate Speech Dataset 🗣️🚫 Author: IbrahimAmin, Mostafa Abbas, Rany Hatem, Andrew Ihab, Mohamed Waleed Fahkr License: MIT Paper: Fine-tuning Arabic Pre-Trained Transformer Models for Egyptian-Arabic Dialect Offensive Language and Hate Speech Detection and Classification Languages: Arabic (Egyptian Dialect) 📋 Dataset Summary This dataset consists of 8,169 Egyptian-Arabic text samples manually labeled for offensive language and hate speech… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-hate-speech.texttext-classification1K<n<10K2 likes110 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.