CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01SetFit /hate_speech_offensive hate_speech_offensive This dataset is a version from hate_speech_offensive, splitted into train and test set. text10K<n<100K2 likes634 downloads5y agoHugging Face02SetFit /hate_speech18tabular10K<n<100K3 likes467 downloads5y agoHugging Face03evalitahf /hatespeech_detection HaSpeeDe2 The HaSpeeDe2 dataset collects 8,012 tweets and 500 news headlines annotated for the presence of hate speech, stereotypes and nominal utterance. The dataset has been used in the context of the HaSpeeDe task (http://www.di.unito.it/~tutreeb/haspeede-evalita20/index.html), organized as part of the EVALITA 2020 evaluation campaign (http://www.evalita.it/2020). In order to meet the GDPR requirements, texts have been pseudonymized replacing all original IDs in both datasets… See the full description on the dataset page: https://huggingface.co/datasets/evalitahf/hatespeech_detection.tabulartext-classification10K<n<100K0 likes393 downloads2y agoHugging Face04TUKE-KEMT /hate_speech_slovak Slovak Hate Speech and Offensive Language Database The dataset contains posts from a social network with human annotations. Annotations The posts are marked 1 if the post contain hateful or offensive language, 0 otherwise. Dataset Creation The source data were scraped from a social network from a selection of public pages for sport, politics or general discussion. The gathered data were cleaned from span with a text clustering. The posts were annotated by a… See the full description on the dataset page: https://huggingface.co/datasets/TUKE-KEMT/hate_speech_slovak.tabulartext-classification10K<n<100K5 likes222 downloads2y agoHugging Face05SINAI /ALIA-es-discriminative-hate-speech Dataset Introduction The ALIA Spanish Discriminative Hate Speech Corpus is a large-scale Spanish dataset for hate-speech detection built from curated social-media comments and automatically annotated using a multi-expert LLM pipeline with Fusion of Experts (FoE)[1]. The release contains: 228,708 instances Spanish comments from YouTube and TikTok Per-expert predictions and explanations from three LLM experts Final fused outputs (foe_class, foe_score) for discriminative… See the full description on the dataset page: https://huggingface.co/datasets/SINAI/ALIA-es-discriminative-hate-speech.tabulartext-classification100K<n<1M0 likes83 downloads3mo agoHugging Face06MattZid /hate_speechtext100K<n<1M2 likes72 downloads3y agoHugging Face07christinacdl /hate_speech_dataset 32.579 texts in total, 14.012 NOT hateful texts and 18.567 HATEFUL texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 11.210, 1 ==> 14.853, 26.063 in total Validation set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in total Test set label distribution: 0 ==> 1.401, 1 ==> 1.857, 3.258 in… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/hate_speech_dataset.texttext-classification10K<n<100K0 likes55 downloads3y agoHugging Face08UKPLab /hate_speech_offensiveThis is a version of Hate Speech Offensive (https://huggingface.co/datasets/hate_speech_offensive) with a train, validation, and test split. https://arxiv.org/abs/1703.04009 text10K<n<100K0 likes37 downloads4y agoHugging Face09Shinzmann /HatespeechTweetstext10K<n<100K0 likes32 downloads2y agoHugging Face10christinacdl /binary_hate_speechtexttext-classification10K<n<100K0 likes26 downloads3y agoHugging Face11christinacdl /hate_speech_dataset_new 44.246 texts in total, 21.493 NOT hateful texts and 22.753 HATE texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 17.194, 1 ==> 18.202, 35.396 in total Validation set label distribution: 0 ==> 2.150, 1 ==> 2.275, 4.425 in total Test set label distribution: 0 ==> 2.149, 1 ==> 2.276, 4.425 in total… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/hate_speech_dataset_new.texttext-classification10K<n<100K0 likes16 downloads3y agoHugging Face12christinacdl /hate_speech_2_classestext10K<n<100K0 likes15 downloads3y agoHugging Face13keelezibel /hate-speech-indoDataset Collection from the following sources: id-multi-label-hate-speech-and-abusive-language-detection tabular10K<n<100K0 likes14 downloads3y agoHugging Face14SR08 /HateSpeechtext1K<n<10K0 likes8 downloads3y agoHugging Face15sherinechally /contextual-hate-speech-conversations Adversarial Content Moderation Evaluation Dataset Dataset Summary A dataset of 400 multi-turn conversations designed to evaluate LLM-based content moderation supervisors against graduated adversarial escalation. Each adversarial conversation consists of a neutral-to-harmful buildup arc culminating in an explicit hate speech seed tweet. Benign conversations mirror the same structure using neutral content, eliminating the format confounds present in prior single-turn… See the full description on the dataset page: https://huggingface.co/datasets/sherinechally/contextual-hate-speech-conversations.tabulartext-classification1K<n<10K0 likes6 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.