CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /hate_speech_offensive hate_speech_offensive This dataset is a version from hate_speech_offensive, splitted into train and test set. text10K<n<100K2 likes632 downloads5y agoHugging Face02SetFit /hate_speech18tabular10K<n<100K3 likes456 downloads5y agoHugging Face03evalitahf /hatespeech_detection HaSpeeDe2 The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance. The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020). In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.tabulartext-classification10K<n<100K0 likes377 downloads2y agoHugging Face04TUKE-KEMT /hate_speech_slovak Slovak Hate Speech and Offensive Language Database The dataset contains posts from a social network with human annotations. Annotations The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise. Dataset Creation The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering. The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.tabulartext-classification10K<n<100K5 likes208 downloads2y agoHugging Face05MattZid /hate_speechtext100K<n<1M2 likes68 downloads3y agoHugging Face06christinacdl /binary_hate_speechtexttext-classification10K<n<100K0 likes67 downloads3y agoHugging Face07SINAI /ALIA-es-discriminative-hate-speech Dataset Introduction The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1]. The release contains: 228,708 instances Spanish comments from YouTube and TikTok Per-expert predictions and explanations from three LLM experts Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.tabulartext-classification100K<n<1M0 likes64 downloads3mo agoHugging Face08christinacdl /hate_speech_dataset 32.579 texts in total, 14.012 NOT hateful texts and 18.567 HATEFUL texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 11.210, 1 ==> 14.853, 26.063 in total Validation set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in total Test set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/hate_speech_dataset.texttext-classification10K<n<100K0 likes43 downloads3y agoHugging Face09UKPLab /hate_speech_offensiveThis is a version of Hate Speech Offensive (https://huggingface.co/datasets/hate_speech_offensive) with a train, validation, and test split. https://arxiv.org/abs/1703.04009 text10K<n<100K0 likes33 downloads4y agoHugging Face10Shinzmann /HatespeechTweetstext10K<n<100K0 likes23 downloads2y agoHugging Face11christinacdl /hate_speech_dataset_new 44.246 texts in total, 21.493 NOT hateful texts and 22.753 HATE texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 17.194, 1 ==> 18.202, 35.396 in total Validation set label distribution: 0 ==> 2.150, 1 ==> 2.275, 4.425 in total Test set label distribution: 0 ==> 2.149, 1 ==> 2.276, 4.425 in total… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/hate_speech_dataset_new.texttext-classification10K<n<100K0 likes13 downloads3y agoHugging Face12SR08 /HateSpeechtext1K<n<10K0 likes9 downloads3y agoHugging Face13sherinechally /contextual-hate-speech-conversations Adversarial Content Moderation Evaluation Dataset Dataset Summary A dataset of 400 multi-turn conversations designed to evaluate LLM-based content moderation supervisors against graduated adversarial escalation. Each adversarial conversation consists of a neutral-to-harmful buildup arc culminating in an explicit hate speech seed tweet. Benign conversations mirror the same structure using neutral content, eliminating the format confounds present in prior single-turn… See the full description on the dataset page: https://huggingface.co/datasets/sherinechally/contextual-hate-speech-conversations.tabulartext-classification1K<n<10K0 likes9 downloads6mo agoHugging Face14christinacdl /hate_speech_2_classestext10K<n<100K0 likes4 downloads3y agoHugging Face15keelezibel /hate-speech-indoDataset Collection from the following sources: id-multi-label-hate-speech-and-abusive-language-detection tabular10K<n<100K0 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.