CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nayohan /korean-hate-speechreference: https://github.com/kocohub/korean-hate-speech @inproceedings{moon-etal-2020-beep, title = "{BEEP}! {K}orean Corpus of Online News Comments for Toxic Speech Detection", author = "Moon, Jihyung and Cho, Won Ik and Lee, Junbum", booktitle = "Proceedings of the Eighth International Workshop on Natural Language Processing for Social Media", month = jul, year = "2020", address = "Online", publisher = "Association for Computational Linguistics"… See the full description on the dataset page: https://huggingface.co/datasets/nayohan/korean-hate-speech.text1K<n<10K2 likes4.1k downloads2y agoHugging Face02ucberkeley-dlab /measuring-hate-speech Dataset card for Measuring Hate Speech This is a public release of the dataset described in Kennedy et al. (2020) and Sachdeva et al. (2022), consisting of 39,565 comments annotated by 7,912 annotators, for 135,556 combined rows. The primary outcome variable is the "hate speech score" but the 10 constituent ordinal labels (sentiment, (dis)respect, insult, humiliation, inferior status, violence, dehumanization, genocide, attack/defense, hate speech benchmark) can also be treated as… See the full description on the dataset page: https://huggingface.co/datasets/ucberkeley-dlab/measuring-hate-speech.tabulartext-classification100K<n<1M54 likes2k downloads9mo agoHugging Face03tdavidson /hate_speech_offensive Dataset Card for [Dataset Name] Dataset Summary An annotated dataset for hate speech and offensive language detection on tweets. Supported Tasks and Leaderboards [More Information Needed] Languages English (en) Dataset Structure Data Instances { "count": 3, "hate_speech_annotation": 0, "offensive_language_annotation": 0, "neither_annotation": 3, "label": 2, # "neither" "tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.tabulartext-classification10K<n<100K42 likes2k downloads3y agoHugging Face04tweets-hate-speech-detection /tweets_hate_speech_detection Dataset Card for Tweets Hate Speech Detection Dataset Summary The objective of this task is to detect hate speech in tweets. For the sake of simplicity, we say a tweet contains hate speech if it has a racist or sexist sentiment associated with it. So, the task is to classify racist or sexist tweets from other tweets. Formally, given a training sample of tweets and labels, where label ‘1’ denotes the tweet is racist/sexist and label ‘0’ denotes the tweet is not… See the full description on the dataset page: https://huggingface.co/datasets/tweets-hate-speech-detection/tweets_hate_speech_detection.texttext-classification10K<n<100K18 likes827 downloads2y agoHugging Face05community-datasets /roman_urdu_hate_speech Dataset Card for roman_urdu_hate_speech Dataset Summary The Roman Urdu Hate-Speech and Offensive Language Detection (RUHSOLD) dataset is a Roman Urdu dataset of tweets annotated by experts in the relevant language. The authors develop the gold-standard for two sub-tasks. First sub-task is based on binary labels of Hate-Offensive content and Normal content (i.e., inoffensive language). These labels are self-explanatory. The authors refer to this sub-task as coarse-grained… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/roman_urdu_hate_speech.texttext-classification10K<n<100K3 likes333 downloads2y agoHugging Face06rezacsedu /bn_hate_speech Dataset Card for Bengali Hate Speech Dataset Dataset Summary The Bengali Hate Speech Dataset is a Bengali-language dataset of news articles collected from various Bengali media sources and categorized based on the type of hate in the text. The dataset was created to provide greater support for under-resourced languages like Bengali on NLP tasks, and serves as a benchmark for multiple types of classification tasks. Supported Tasks and Leaderboards topic… See the full description on the dataset page: https://huggingface.co/datasets/rezacsedu/bn_hate_speech.texttext-classification1K<n<10K3 likes233 downloads3y agoHugging Face07aiatums /codemixed-id-hate-speech Code-mixed Indonesian Hate Speech Dataset Manually annotated hate speech dataset for Indonesian-Javanese and Indonesian-Sundanese code-mixed text, enriched with LLM-generated augmentations. texttext-classification10K<n<100K0 likes218 downloads4mo agoHugging Face08Lots-of-LoRAs /task905_hate_speech_offensive_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.texttext-generation1K<n<10K0 likes197 downloads2y agoHugging Face09piuba-bigdata /contextualized_hate_speech Contextualized Hate Speech: A dataset of comments in news outlets on Twitter Dataset Summary This dataset is a collection of tweets that were posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were analyzed for hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against physical… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech.texttext-classification10K<n<100K8 likes195 downloads2y agoHugging Face10community-datasets /hate_speech_pl Dataset Card for HateSpeechPl Dataset Summary The dataset was created to analyze the possibility of automating the recognition of hate speech in Polish. It was collected from the Polish forums and represents various types and degrees of offensive language, expressed towards minorities. The original dataset is provided as an export of MySQL tables, what makes it hard to load. Due to that, it was converted to CSV and put to a Github repository. Supported Tasks… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/hate_speech_pl.tabulartext-classification10K<n<100K4 likes162 downloads2y agoHugging Face11suwaimyo /hatespeech-ind-multilabelclassification HateSpeech_ind_MultiLabelClassification Deduplicated copy of kornwtp/hatespeech-ind-multilabelclassification. Splits split rows train 13,014 text10K<n<100K0 likes136 downloads24d agoHugging Face12mteb /HateSpeechPortugueseClassification HateSpeechPortugueseClassification An MTEB dataset Massive Text Embedding Benchmark HateSpeechPortugueseClassification is a dataset of Portuguese tweets categorized with their sentiment (2 classes). Task category t2c Domains Social, Written Reference https://aclanthology.org/W19-3510 How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/mteb/HateSpeechPortugueseClassification.texttext-classification1K<n<10K1 likes134 downloads1y agoHugging Face13suwaimyo /hatespeech-fil-classification Hatespeech_fil_Classification Deduplicated copy of kornwtp/hatespeech-fil-classification. Splits split rows test 4,127 train 9,671 validation 4,153 text10K<n<100K0 likes106 downloads28d agoHugging Face14IbrahimAmin /egyptian-arabic-hate-speech 🇪🇬 Egyptian-Arabic Hate Speech Dataset 🗣️🚫 Author: IbrahimAmin, Mostafa Abbas, Rany Hatem, Andrew Ihab, Mohamed Waleed Fahkr License: MIT Paper: Fine-tuning Arabic Pre-Trained Transformer Models for Egyptian-Arabic Dialect Offensive Language and Hate Speech Detection and Classification Languages: Arabic (Egyptian Dialect) 📋 Dataset Summary This dataset consists of 8,169 Egyptian-Arabic text samples manually labeled for offensive language and hate speech… See the full description on the dataset page: https://huggingface.co/datasets/IbrahimAmin/egyptian-arabic-hate-speech.texttext-classification1K<n<10K2 likes95 downloads1y agoHugging Face15Lots-of-LoRAs /task904_hate_speech_offensive_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task904_hate_speech_offensive_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task904_hate_speech_offensive_classification.texttext-generation1K<n<10K0 likes84 downloads2y agoHugging Face16suwaimyo /hatespeech-ind-classification HateSpeech_ind_Classification Deduplicated copy of kornwtp/hatespeech-ind-classification. Splits split rows train 703 textn<1K0 likes82 downloads28d agoHugging Face17kornwtp /hatespeech-fil-classification Dataset Card for "hatespeech-filipino" More Information needed ref: https://huggingface.co/datasets/legacy-datasets/hate_speech_filipino text10K<n<100K0 likes71 downloads2y agoHugging Face18edumunozsala /preference-hate-speech-estext1K<n<10K1 likes68 downloads2y agoHugging Face19arbml /osact5_hatespeech Dataset Card for [Dataset Name] Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed] Source Data… See the full description on the dataset page: https://huggingface.co/datasets/arbml/osact5_hatespeech.text10K<n<100K0 likes67 downloads2y agoHugging Face20DL-Project /hatespeech_synthesized_datasetaudio10K<n<100K1 likes57 downloads2y agoHugging Face21Intuit-GenSRF /tweets-hate-speech-detection Dataset Card for "tweets_hate_speech_detection" More Information needed text10K<n<100K1 likes54 downloads3y agoHugging Face22arbml /OSACT4_hatespeech Dataset Card for "OSACT4_hatespeech" More Information needed text1K<n<10K1 likes46 downloads2y agoHugging Face23piuba-bigdata /contextualized_hate_speech_raw Contextualized Hate Speech: A dataset of comments in news outlets on Twitter Dataset Summary This dataset is a collection of tweets posted in response to news articles from five specific Argentinean news outlets: Clarín, Infobae, La Nación, Perfil and Crónica, during the COVID-19 pandemic. The comments were annotated for the presence of hate speech across eight different characteristics: against women, racist content, class hatred, against LGBTQ+ individuals, against… See the full description on the dataset page: https://huggingface.co/datasets/piuba-bigdata/contextualized_hate_speech_raw.texttext-classification10K<n<100K0 likes43 downloads2y agoHugging Face24arbml /Arabic_Hate_Speech Dataset Card for "Arabic_Hate_Speech" More Information needed text1K<n<10K7 likes38 downloads2y agoHugging Face25puttatidam /hatespeech-ind-multilabelclassificationtext10K<n<100K0 likes37 downloads10d agoHugging Face26Dauren-Nur /hate_speech_dataset_combinedtext100K<n<1M0 likes36 downloads2y agoHugging Face27omanyasa /shona-hate-speech Dataset Card for Balanced Shona Hate Speech Dataset Dataset Summary This dataset contains 2,000 balanced examples of Shona text classified into four categories: NEUTRAL, OFFENSIVE, CONTEXTUAL, and HATE. Data Sources Label Source Count NEUTRAL Literary novel (Imbwa Yemunhu by Ignatius T. Mabasa) 500 OFFENSIVE Synthetic template-based generation 500 CONTEXTUAL Synthetic (quoted hate speech, not endorsed) 500 HATE Synthetic (direct attacks on… See the full description on the dataset page: https://huggingface.co/datasets/omanyasa/shona-hate-speech.texttext-classification1K<n<10K0 likes34 downloads5mo agoHugging Face28Lots-of-LoRAs /task1493_bengali_geopolitical_hate_speech_binary_classification Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1493_bengali_geopolitical_hate_speech_binary_classification Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1493_bengali_geopolitical_hate_speech_binary_classification.texttext-generation1K<n<10K0 likes33 downloads2y agoHugging Face29abdulrub /hate_speech_dataset Dataset Description This dataset is designed for fine-tuning language models, particularly the Qwen2.5-1.5B-Instruct model, for the task of hate speech detection in social media text (tweets). It focuses on both implicit and explicit forms of hate speech, aiming to improve the performance of smaller language models in this challenging task. The dataset is a combination of two existing datasets: Hate Speech Examples: Examples of implicit hate speech are sourced from the… See the full description on the dataset page: https://huggingface.co/datasets/abdulrub/hate_speech_dataset.text1K<n<10K0 likes33 downloads2y agoHugging Face30arbml /Religious_Hate_Speechtext1K<n<10K0 likes32 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.