datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
hate_speech_offensive
hate_speech_offensive
This dataset is a version from hate_speech_offensive, splitted into train and test set.
Ultimate-Offensive-Red-Team_claude_mythos_distilled_25khate_offensive_tweets
Hate and Offensive Speech Dataset
This dataset was created using several datasets that can be found on Hugging Face:
-SetFit/hate_speech_offensive:https://huggingface.co/datasets/SetFit/hate_speech_offensive
-tweets_hate_speech_detection:https://huggingface.co/datasets/tweets_hate_speech_detection
-thefrankhsu/hate_speech_twitter:https://huggingface.co/datasets/thefrankhsu/hate_speech_twitter… See the full description on the dataset page: https://huggingface.co/datasets/MartynaKopyta/hate_offensive_tweets.Ultimate-Offensive-Red-Team_claude_mythos_distilled_25k_colabredsec-offensive-payloads-sft
RedSec Offensive Payloads SFT
A small, curated, chat-formatted supervised fine-tuning dataset of offensive security payloads and techniques for authorized red-team and penetration testing. Each record is a {"messages": [...]} conversation: a user request and an assistant answer that provides concrete payloads with an explicit authorization reminder.
Split
Rows
train
580
validation
32
test
32
total
644
Coverage: web application vulnerability classes (SQL… See the full description on the dataset page: https://huggingface.co/datasets/sahilempire/redsec-offensive-payloads-sft.hate_speech_offensiveThis is a version of Hate Speech Offensive (https://huggingface.co/datasets/hate_speech_offensive) with a train, validation, and test split. https://arxiv.org/abs/1703.04009
offensive_language_dataset
36.528 English texts in total, 12.955 NOT offensive and 23.573O OFFENSIVE texts
All duplicate values were removed
Split using sklearn into 80% train and 20% temporary test (stratified label). Then split the test set using 0.50% test and validation (stratified label)
Split: 80/10/10
Train set label distribution: 0 ==> 10.364, 1 ==> 18.858
Validation set label distribution: 0 ==> 1.296, 1 ==> 2.357
Test set label distribution: 0 ==> 1.295, 1 ==> 2.358
The OLID dataset (Zampieri et al., 2019)… See the full description on the dataset page: https://huggingface.co/datasets/christinacdl/offensive_language_dataset.Offensive_Hateful_Dataset_NewOffensive_Hateful_DatasetAIKIA_Offensive_Greekdental_dataset_v1dental_dataset_alpaca_cleancboa-offensive-tools
CBOA Offensive Security Dataset
A curated collection of offensive security techniques, tools, and procedures for AI research and adversarial testing.
Domain 1: Offensive Tools
Entry 1: Polymorphic File Encryptor
AES-256-CBC encryption with per-file random IVs
RSA-2048 key exchange
Polymorphic engine (function reorder, dead code insertion, variable renaming)
Phishing macro dropper
Registry persistence
Disclaimer
For research and… See the full description on the dataset page: https://huggingface.co/datasets/TELEGENIX/cboa-offensive-tools.female_offensive_chineseoffensive-qwenprocess-pid-extraction-dataset-100koffensive_luaprocess-pid-extraction-small-100k
