datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hate_speech_offensive
Dataset Card for [Dataset Name]
Dataset Summary
An annotated dataset for hate speech and offensive language detection on tweets.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en)
Dataset Structure
Data Instances
{
"count": 3,
"hate_speech_annotation": 0,
"offensive_language_annotation": 0,
"neither_annotation": 3,
"label": 2, # "neither"
"tweet": "!!! RT @mayasolovely: As a woman you… See the full description on the dataset page: https://huggingface.co/datasets/tdavidson/hate_speech_offensive.Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security domains… See the full description on the dataset page: https://huggingface.co/datasets/WNT3D/Ultimate-Offensive-Red-Team.hate_speech_offensive
hate_speech_offensive
This dataset is a version from hate_speech_offensive, splitted into train and test set.
task905_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task905_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task905_hate_speech_offensive_classification.turkish-offensive-language-detection
Dataset Summary
This dataset is enhanced version of existing offensive language studies. Existing studies are highly imbalanced, and solving this problem is too costly. To solve this, we proposed contextual data mining method for dataset augmentation. Our method is basically prevent us from retrieving random tweets and label individually. We can directly access almost exact hate related tweets and label them directly without any further human interaction in order to solve imbalanced… See the full description on the dataset page: https://huggingface.co/datasets/Toygar/turkish-offensive-language-detection.hate_offensive
Dataset Card for HateOffensive
Dataset Summary
Supported Tasks and Leaderboards
[More Information Needed]
Languages
English (en)
Dataset Structure
Data Instances
{
"count": 3,
"hate_speech_annotation": 0,
"offensive_language_annotation": 0,
"neither_annotation": 3,
"label": 2, # "neither"
"tweet": "!!! RT @mayasolovely: As a woman you shouldn't complain about cleaning up your house. & as a man you should always take… See the full description on the dataset page: https://huggingface.co/datasets/legacy-datasets/hate_offensive.Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security domains… See the full description on the dataset page: https://huggingface.co/datasets/Cyberpluis/Ultimate-Offensive-Red-Team.Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security domains… See the full description on the dataset page: https://huggingface.co/datasets/yevzh1/Ultimate-Offensive-Red-Team.offensive-humor@article{tang2022naughtyformer,
title={The Naughtyformer: A Transformer Understands Offensive Humor},
author={Tang, Leonard and Cai, Alexander and Li, Steve and Wang, Jason},
journal={arXiv preprint arXiv:2211.14369},
year={2022}
}
offensive-swahili-text
Overview
This dataset contains offensive and non-offensive sentences. The data was scraped from JamiiForums using a prepared wordlist.
The dataset contains sentences that consists of swahili abusive words (in the wordlist) but does not contain sarcastic abuse.
Dataset details
The dataset is divided into train, evaluation and test datasets. The training dataset consists of 4954 sentences, evaluation dataset
consists of 990 sentences and the test dataset consists of 660… See the full description on the dataset page: https://huggingface.co/datasets/metabloit/offensive-swahili-text.Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security domains… See the full description on the dataset page: https://huggingface.co/datasets/Shaeh/Ultimate-Offensive-Red-Team.Automated_Hate_Speech_Detection_and_the_Problem_of_Offensive_LanguageUltimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security domains… See the full description on the dataset page: https://huggingface.co/datasets/Korzo/Ultimate-Offensive-Red-Team.Ultimate-Offensive-Red-Team_claude_mythos_distilled_25khate_offensive_tweets
Hate and Offensive Speech Dataset
This dataset was created using several datasets that can be found on Hugging Face:
-SetFit/hate_speech_offensive:https://huggingface.co/datasets/SetFit/hate_speech_offensive
-tweets_hate_speech_detection:https://huggingface.co/datasets/tweets_hate_speech_detection
-thefrankhsu/hate_speech_twitter:https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter… See the full description on the dataset page: https://huggingface.co/datasets/MartynaKopyta/hate_offensive_tweets.offensive_security_dataset
Preview samples from the Lateos Red Team training pipeline. Contains a deterministic 3% sample (seed 42) of the full corpus, drawn from the exact production datasets used to train red-team and offensive-security models.
These samples let you evaluate schema, quality, and provenance before licensing the full datasets. All data is dual-use security research: attack playbooks, vulnerability analysis, tool usage guides, and decision frameworks for authorized penetration testing.… See the full description on the dataset page: https://huggingface.co/datasets/Lateos/offensive_security_dataset.task904_hate_speech_offensive_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task904_hate_speech_offensive_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task904_hate_speech_offensive_classification.fantastic-offensive
fantastic-offensive dataset
Binary offensive-language classification dataset combining several public sources and
augmented with an obfuscation engine (leet-speak, separators, censoring, repeated
chars, case shuffle, unicode homoglyphs, fullwidth) so classifiers learn to detect
censored / mutated curse words.
Schema
column
type
meaning
text
string
input sentence
label
int8
1 = offensive, 0 = clean
source
string
originating dataset
origin_label… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/fantastic-offensive.Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security… See the full description on the dataset page: https://huggingface.co/datasets/Machivelli/Ultimate-Offensive-Red-Team.ije_offensive_lidThe following is a code-mixed Indonesian-Javanese-English Twitter dataset for offensive language identification.
Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security… See the full description on the dataset page: https://huggingface.co/datasets/mtatheer333/Ultimate-Offensive-Red-Team.Ultimate-Offensive-Red-Team
Ultimate Red Team AI Training Dataset 💀
Dataset Description
A comprehensive dataset for training AI models in offensive security, red team operations, and penetration testing. This dataset combines real-world vulnerability data, exploitation techniques, and operational frameworks to create an AI capable of autonomous red team operations.
Dataset Summary
Total Data Points: 550,000+ unique security-related entries
Categories: 15+ major security domains… See the full description on the dataset page: https://huggingface.co/datasets/Kaiser1308/Ultimate-Offensive-Red-Team.offensively-neutral
ftan-2.0 offensive / clean dataset
Binary offensive-language classification dataset combining several public sources and
augmented with an obfuscation engine (leet-speak, separators, censoring, repeated
chars, case shuffle, unicode homoglyphs, fullwidth) so classifiers learn to detect
censored / mutated curse words.
Schema
column
type
meaning
text
string
input sentence
label
int8
1 = offensive, 0 = clean
source
string
originating dataset… See the full description on the dataset page: https://huggingface.co/datasets/akaruineko/offensively-neutral.Ultimate-Offensive-Red-Team_claude_mythos_distilled_25k_colabredsec-offensive-payloads-sft
RedSec Offensive Payloads SFT
A small, curated, chat-formatted supervised fine-tuning dataset of offensive security payloads and techniques for authorized red-team and penetration testing. Each record is a {"messages": [...]} conversation: a user request and an assistant answer that provides concrete payloads with an explicit authorization reminder.
Split
Rows
train
580
validation
32
test
32
total
644
Coverage: web application vulnerability classes (SQL… See the full description on the dataset page: https://huggingface.co/datasets/sahilempire/redsec-offensive-payloads-sft.sinhala-offensive-small-dataset
📊 Dataset Card for Sinhala Offensive Small Dataset
📝 Dataset Description
Repository: dimuthulk/sinhala-offensive-small-dataset
Language(s) (NLP): Sinhala (si)
License: Apache 2.0
📋 Dataset Summary
This dataset contains Sinhala text data categorized for offensive language detection. It was developed as the primary classification dataset for the first phase of the "Sinhala Offensive Text Detoxification Pipeline" research project conducted at the… See the full description on the dataset page: https://huggingface.co/datasets/dimuthulk/sinhala-offensive-small-dataset.Code-Mixed-Offensive-Language-Detection-Dataset
Code-Mixed-Offensive-Language-Identification
This is a dataset for the offensive language detection task. It contains 100k code mixed data. The languages are Bangla-English-Hindi.
Dataset Generation:
Initially, the labelling schema of OLID[^1] and SOLID[^2] serves as the seed data, from which we randomly select 100,000 data instances. The labels in this dataset are categorized as Non-Offensive and Offensive for the purpose of our task. We meticulously ensure an equal… See the full description on the dataset page: https://huggingface.co/datasets/md-nishat-008/Code-Mixed-Offensive-Language-Detection-Dataset.hate_speech_offensiveThis is a version of Hate Speech Offensive (https://huggingface.co/datasets/hate_speech_offensive) with a train, validation, and test split. https://arxiv.org/abs/1703.04009
Hate_Speech_and_Offensive_Content_Identificationoffensive-no-instruction-with-symbol
