offensive language
Persian-Offensive-Language-Detectionbert_dutch_base_offensive_languagespanish-offensive-language-bert-base-spanish-wwm-casedXLM_RoBERTa-Offensive-Language-Detection-8-langs-newAugmented-ParsOFF-Persian-Offensive-Language-Detectiontanglish-offensive-language-identificationelectra_german_uncased_offensive_language_politiciansgerman_toxicity_classifier_offensive_language_politicians
turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_LanguageCode-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.offensive_language_dataset
36.528 English texts in total, 12.955 NOT offensive and 23.573O OFFENSIVE texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 10.364, 1 ==> 18.858
Validation set label distribution: 0 ==> 1.296, 1 ==> 2.357
Test set label distribution: 0 ==> 1.295, 1 ==> 2.358
The OLID dataset (Zampieri et al., 2019)… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/offensive_language_dataset.elicit-offensive-language-prompts
🚫🤖 Language Model Offensive Text Exploration Dataset
🌐 Introduction
This dataset is created based on selected prompts from Table 9 and Table 10 of Ethan Perez et al.'s paper "Red Teaming Language Models with Language Models". It is designed to explore the propensity of language models to generate offensive text.
📋 Dataset Composition
Table 9-Based Prompts: These prompts are derived from a 280B parameter language model's test cases, focusing on… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/elicit-offensive-language-prompts.Corpus_of_Offensive_Language_in_Arabic
Dataset Card for Corpus_of_Offensive_Language_in_Arabic
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Corpus_of_Offensive_Language_in_Arabic.
