datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.offensive-humor@article{tang2022naughtyformer,
title={The Naughtyformer: A Transformer Understands Offensive Humor},
author={Tang, Leonard and Cai, Alexander and Li, Steve and Wang, Jason},
journal={arXiv preprint arXiv:2211.14369},
year={2022}
}
offensive-swahili-text
Overview
This dataset contains offensive and non-offensive sentences. The data was scraped from JamiiForums using a prepared wordlist.
The dataset contains sentences that consists of swahili abusive words (in the wordlist) but does not contain sarcastic abuse.
Dataset details
The dataset is divided into train, evaluation and test datasets. The training dataset consists of 4954 sentences, evaluation dataset
consists of 990 sentences and the test dataset consists of 660… See the full description on the dataset page: https://huggingface.co/datasets/metabloit/offensive-swahili-text.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_Languageije_offensive_lidThe following is a code-mixed Indonesian-Javanese-English Twitter dataset for offensive language identification.
Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.Hate_Speech_and_Offensive_Content_Identificationmteb_HateBR_offensive_binaryturkish-social-media-offensive-dataset
Overwiev
It is a 4-class Turkish bullying data set obtained from Twitter.
Cinsiyetçilik
Irkçılık
Kızdırma
Nötr
Sum
601
490
910
1387
3388
Authors
Seyma SARIGIL: seymasargil@gmail.com
Elif SARIGIL KARA: elifsarigil@gmail.com
Murat KOKLU: mkoklu@selcuk.edu.tr
Alaaddin Erdinç DAL: aerdincdal@icloud.com
offensive-and-grooming-dataset
Description
This dataset contains elements from the offendES dataset and translations from the sexismreddit dataset from english to spanish. The aim of this dataset is to provide training data for models capable to identify harming text towards kids.
The id2label dictionary for this dataset is as follows:
id2label = {0:"OFP", 1:"OFG", 2:"NO",3:"NOE", 4:"GP"}
Where OFP stands for offensive messages targeted to a single person, OFG stands for offensive messages targeted to a group or… See the full description on the dataset page: https://huggingface.co/datasets/Brandon-h/offensive-and-grooming-dataset.OffensiveLang
OffensiveLang
Overview
OffensiveLang is a community based implicit offensive language dataset generated by ChatGPT 3.5 containing data for 38 different target groups. It has been meticulously annotated by Amazon MTurk workers, ensuring high-quality labeling of hate speech. Additionally, a prompt-based zero-shot method was employed with ChatGPT and the detection results were compared between human annotation and ChatGPT annotation. This dataset is invaluable for… See the full description on the dataset page: https://huggingface.co/datasets/AmitDasRup123/OffensiveLang.offensive_comment_frp2c_offensiveRO_offensive_political_commentsnba-offensive-rebound-kickout-coherence-risk-v0.1What this repo is for
Detect if second-chance kickouts create real advantage.
Focus
crash role discipline
rebound control quality
target openness
scramble state
extra pass decision
Why it matters
Offensive boards are wasted if you force the kickout.
The value is open threes after scramble.
turkish-offensive-datasetoffensive-and-grooming-dataset
Description
This dataset contains elements from the offendES dataset and translations from the sexismreddit dataset from english to spanish. The aim of this dataset is to provide training data for models capable to identify harming text towards kids.
The id2label dictionary for this dataset is as follows:
id2label = {0:"OFP", 1:"OFG", 2:"NO",3:"NOE", 4:"GP"}
Where OFP stands for offensive messages targeted to a single person, OFG stands for offensive messages targeted to a group or… See the full description on the dataset page: https://huggingface.co/datasets/Annamantula/offensive-and-grooming-dataset.offensive_tweetis_offensive_fr_DPO_trainer_formatedRO_offensive_political_comments_balancedHate_offensive_dataset.csvRO_offensive_political_comments_large
