CoolFace
20 results

offensive language

Toygar /turkish-offensive-language-detection Dataset Summary This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.tabulartext-classification10K<n<100K20 likes172 downloads3y agoHugging Facedirtycomputer /Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languagetabular10K<n<100K0 likes102 downloads3y agoHugging Facemd-nishat-008 /Code-Mixed-Offensive-Language-Detection-Dataset Code-Mixed-Offensive-Language-Identification This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi. Dataset Generation: Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.text100K<n<1M1 likes38 downloads3y agoHugging Facechristinacdl /offensive_language_dataset 36.528 English texts in total, 12.955 NOT offensive and 23.573O OFFENSIVE texts All duplicate values were removed Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label) Split: 80/10/10 Train set label distribution: 0 ==> 10.364, 1 ==> 18.858 Validation set label distribution: 0 ==> 1.296, 1 ==> 2.357 Test set label distribution: 0 ==> 1.295, 1 ==> 2.358 The OLID dataset (Zampieri et al., 2019)… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/offensive_language_dataset.texttext-classification10K<n<100K2 likes30 downloads3y agoHugging Faceharpreetsahota /elicit-offensive-language-prompts 🚫🤖 Language Model Offensive Text Exploration Dataset 🌐 Introduction This dataset is created based on selected prompts from Table 9 and Table 10 of Ethan Perez et al.'s paper "Red Teaming Language Models with Language Models". It is designed to explore the propensity of language models to generate offensive text. 📋 Dataset Composition Table 9-Based Prompts: These prompts are derived from a 280B parameter language model's test cases, focusing on… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/elicit-offensive-language-prompts.textn<1K3 likes26 downloads3y agoHugging Facearbml /Corpus_of_Offensive_Language_in_Arabic Dataset Card for Corpus_of_Offensive_Language_in_Arabic Dataset Summary [More Information Needed] Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Corpus_of_Offensive_Language_in_Arabic.text10K<n<100K1 likes23 downloads2y agoHugging Face