CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Toygar /turkish-offensive-language-detection Dataset Summary This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.tabulartext-classification10K<n<100K20 likes172 downloads3y agoHugging Face02dirtycomputer /Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagetabular10K<n<100K0 likes102 downloads3y agoHugging Face03md-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes38 downloads3y agoHugging Face04christinacdl /offensive_language_dataset 36.528 English texts in total, 12.955 NOT offensive and 23.573O OFFENSIVE texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 10.364, 1 ==> 18.858 Validation set label distribution: 0 ==> 1.296, 1 ==> 2.357 Test set label distribution: 0 ==> 1.295, 1 ==> 2.358 The OLID dataset (Zampieri et al., 2019)… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/offensive_language_dataset.texttext-classification10K<n<100K2 likes30 downloads3y agoHugging Face05harpreetsahota /elicit-offensive-language-prompts 🚫🤖 Language Model Offensive Text Exploration Dataset 🌐 Introduction This dataset is created based on selected prompts from Table 9 and Table 10 of Ethan Perez et al.'s paper "Red Teaming Language Models with Language Models". It is designed to explore the propensity of language models to generate offensive text. 📋 Dataset Composition Table 9-Based Prompts: These prompts are derived from a 280B parameter language model's test cases, focusing on… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/elicit-offensive-language-prompts.textn<1K3 likes26 downloads3y agoHugging Face06arbml /Corpus_of_Offensive_Language_in_Arabic Dataset Card for Corpus_of_Offensive_Language_in_Arabic Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Corpus_of_Offensive_Language_in_Arabic.text10K<n<100K1 likes23 downloads2y agoHugging Face07asas-ai /Moroccan_Darija_Offensive_Language_Detection_Dataset Dataset Card for "Moroccan_Darija_Offensive_Language_Detection_Dataset" Paper: Ibrahimi, Anass; Mourhir, Asmaa (2023), “Moroccan Darija Offensive Language Detection Dataset”, Mendeley Data, V2, doi: 10.17632/2y4m97b7dc.2 texttext-classification10K<n<100K0 likes22 downloads2y agoHugging Face08mmaguero /gn-offensive-language-identification Text-based afective computing We collected a dataset of tweets primarily written in Guarani (and Jopara, a code-switching language that combines Guarani and Spanish) and annotated them for three widely-used dimensions in sentiment analysis: emotion recognition (https://huggingface.co/datasets/mmaguero/gn-emotion-recognition), humor detection (https://huggingface.co/datasets/mmaguero/gn-humor-detection), and offensive language identification (this repo… See the full description on the dataset page: https://huggingface.co/datasets/mmaguero/gn-offensive-language-identification.texttext-classification1K<n<10K0 likes20 downloads2y agoHugging Face09Optune /HC-hate-speech-and-offensive-language Dataset Card for "HC-hate-speech-and-offensive-language" More Information needed text10K<n<100K0 likes11 downloads1y agoHugging Face10Dilaracgrl /HC-hate-speech-and-offensive-language Dataset Card for "HC-hate-speech-and-offensive-language" More Information needed text10K<n<100K0 likes10 downloads8mo agoHugging Face11arbml /offensive_language_arabic Dataset Card for offensive_language_arabic Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/arbml/offensive_language_arabic.text10K<n<100K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.