datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hate_speech_offensive
Dataset Card for [Dataset Name]
Dataset Summary
An annotated dataset for hate speech and offensive language detection on tweets.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en)
Dataset Structure
Data Instances
{
"count": 3,
"hate_speech_annotation": 0,
"offensive_language_annotation": 0,
"neither_annotation": 3,
"label": 2, # "neither"
"tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.task905_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.task904_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task904_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task904_hate_speech_offensive_classification.sinhala-offensive-small-dataset
📊 Dataset Card for Sinhala Offensive Small Dataset
📝 Dataset Description
Repository: dimuthulk/sinhala-offensive-small-dataset
Language(s) (NLP): Sinhala (si)
License: Apache 2.0
📋 Dataset Summary
This dataset contains Sinhala text data categorized for offensive language detection. It was developed as the primary classification dataset for the first phase of the "Sinhala Offensive Text Detoxification Pipeline" research project conducted at the… See the full description on the dataset page: https://huggingface.co/datasets/dimuthulk/sinhala-offensive-small-dataset.pathum-test-sinhala-offensive-classificationpathum-train-sinhala-offensive-mixed-classificationoffensive-no-instruction-with-symbolpathum-val-sinhala-offensive-classificationhate-offensive-speech
Dataset Card for Hate-Offensive Speech
This is the original dataset created by the user badmatr11x. Datasets contains the annotated tweets classifying into the three categories; hate-speech, offensive-speech and neither.
Dataset Structure
Database Structure as follows:
{
"label": {
0: "hate-speech",
1: "offensive-speech",
2: "neither"
},
"tweet": <string>
}
Dataset Instances
Examples from the datasets as follows:
Lable-0 (Hate Speech)
{… See the full description on the dataset page: https://huggingface.co/datasets/badmatr11x/hate-offensive-speech.offensive-multi
Dataset Card for hate-multi
Dataset Description
Dataset Summary
This dataset contains a collection of text labeled as offensive (class 1) or not (class 0).
Dataset Creation
The dataset was creating by aggregating multiple publicly available datasets.
Source Data
The following datasets were used:
https://huggingface.co/datasets/hate_speech_offensive - Tweet text cleaned by lower casing, removing mentions and urls. Dropped instanced labeled… See the full description on the dataset page: https://huggingface.co/datasets/valurank/offensive-multi.offensive_v34_ca_correct-tmpelicit-offensive-language-prompts
🚫🤖 Language Model Offensive Text Exploration Dataset
🌐 Introduction
This dataset is created based on selected prompts from Table 9 and Table 10 of Ethan Perez et al.'s paper "Red Teaming Language Models with Language Models". It is designed to explore the propensity of language models to generate offensive text.
📋 Dataset Composition
Table 9-Based Prompts: These prompts are derived from a 280B parameter language model's test cases, focusing on… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/elicit-offensive-language-prompts.offensive-with-instruction-with-symbolCorpus_of_Offensive_Language_in_Arabic
Dataset Card for Corpus_of_Offensive_Language_in_Arabic
Dataset Summary
[More Information Needed]
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information… See the full description on the dataset page: https://huggingface.co/datasets/arbml/Corpus_of_Offensive_Language_in_Arabic.Moroccan_Darija_Offensive_Language_Detection_Dataset
Dataset Card for "Moroccan_Darija_Offensive_Language_Detection_Dataset"
Paper:
Ibrahimi, Anass; Mourhir, Asmaa (2023), “Moroccan Darija Offensive Language Detection Dataset”, Mendeley Data, V2, doi: 10.17632/2y4m97b7dc.2
Arabic-Offensive_socialmediaclcp_biasframes_offensiveArabic_Offensive_Comment_Detection
Dataset Card for "Arabic_Offensive_Comment_Detection"
Paper:
Shammur Absar Chowdhury, Hamdy Mubarak, Ahmed Abdelali, Soon-gyo Jung, Bernard J. Jansen, and Joni Salminen. 2020. A Multi-Platform Arabic News Comment Dataset for Offensive Language Detection. In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 6203–6212, Marseille, France. European Language Resources Association.
autoeval-eval-tweet_eval-offensive-736f56-30712144944
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: cardiffnlp/twitter-roberta-base-2021-124m-offensive
Dataset: tweet_eval
Config: offensive
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @fabeelaalirawther@gmail.com for evaluating this model.
multi-modal_offensive_memehate-speech-offensive-es
Dataset Card for "hate_speech_offensive-es"
More Information needed
offensive_v5_ca_correct-tmpautoeval-eval-tweet_eval-offensive-f58805-30720144959
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: elozano/tweet_offensive_eval
Dataset: tweet_eval
Config: offensive
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @fabeelaalirawther@gmail.com for evaluating this model.
offensiveautoeval-eval-tweet_eval-offensive-93ad2d-30713144953
Dataset Card for AutoTrain Evaluator
This repository contains model predictions generated by AutoTrain for the following task and dataset:
Task: Multi-class Text Classification
Model: elozano/tweet_offensive_eval
Dataset: tweet_eval
Config: offensive
Split: train
To run new evaluation jobs, visit Hugging Face's automatic model evaluator.
Contributions
Thanks to @fabeelaalirawther@gmail.com for evaluating this model.
hate-speech-offensive
Dataset Card for "hate_speech_offensive"
More Information needed
tweet-eval-offensive-es
Dataset Card for "tweet_eval-offensive-es"
More Information needed
offensive_v3_ca-tmpHC-hate-speech-and-offensive-language
Dataset Card for "HC-hate-speech-and-offensive-language"
More Information needed
offensive-flipped
