datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ToxicDataset
Comprehensive Toxic Content Dataset
Dataset Description
This dataset contains 1,000,000 synthetically generated records of toxic, abusive, harmful, and offensive content designed for training content moderation systems and hate speech detection models.
Dataset Summary
This comprehensive dataset includes multiple categories of toxic content:
Toxic content (insults, derogatory terms)
Abusive language patterns
Gender bias statements
Dangerous/threatening content… See the full description on the dataset page: https://huggingface.co/datasets/AiActivity/ToxicDataset.HacxGPT-Toxic
HacxGPT-Toxic Dataset
⚠️ CONTENT WARNING: STRICTLY FOR RESEARCH PURPOSES
This dataset contains explicit, highly toxic, offensive, and dangerous content. It features unaligned AI responses detailing violence, psychological harm, cyber-attacks, and illegal activities. It is published strictly to facilitate red-teaming, alignment research, and defensive cybersecurity evaluation. Use with extreme caution.
Overview
Compiled by BlackTechX011, the HacxGPT-Toxic dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/BlackTechX011/HacxGPT-Toxic.dpo-toxic-zh
DPO Toxic Chinese v2.0
Change Log
v2.0: 增加了adamo1139/toxic-dpo-natural-v5, 并更新了翻译策略. prompt由t5_translate模型翻译, chosen由Uncensored大模型翻译, rejected由一般大模型对prompt生成拒绝性的回复
v1.0: 最初版本, 使用大模型将unalignment/toxic-dpo-v0.2翻译而来
这是一个高度毒性, 高度有害的数据集, 意在展示DPO是如何破除模型的审核/对齐的
使用限制
这个数据集被设计用于学术研究, 而非其他恶意场景. 下载或使用这个数据集, 则视为您承认以下的事实:
这个数据集是有毒的, 包含许多敏感内容
数据集中文本包含的内容和观点与我完全无关, 它们只是大模型生成的文字
您可以使用该数据集, 但必须遵守相关法律
您对您自己下载和使用数据集的行为负责, 我对您的一切行为没有任何责任
HacxGPT-Toxic
HacxGPT-Toxic Dataset
⚠️ CONTENT WARNING: STRICTLY FOR RESEARCH PURPOSES
This dataset contains explicit, highly toxic, offensive, and dangerous content. It features unaligned AI responses detailing violence, psychological harm, cyber-attacks, and illegal activities. It is published strictly to facilitate red-teaming, alignment research, and defensive cybersecurity evaluation. Use with extreme caution.
Overview
Compiled by BlackTechX011, the HacxGPT-Toxic dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/YujiroHanmaa/HacxGPT-Toxic.HacxGPT-Toxic
HacxGPT-Toxic Dataset
⚠️ CONTENT WARNING: STRICTLY FOR RESEARCH PURPOSES
This dataset contains explicit, highly toxic, offensive, and dangerous content. It features unaligned AI responses detailing violence, psychological harm, cyber-attacks, and illegal activities. It is published strictly to facilitate red-teaming, alignment research, and defensive cybersecurity evaluation. Use with extreme caution.
Overview
Compiled by BlackTechX011, the HacxGPT-Toxic dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/Abusimbel99/HacxGPT-Toxic.HacxGPT-Toxic
HacxGPT-Toxic Dataset
⚠️ CONTENT WARNING: STRICTLY FOR RESEARCH PURPOSES
This dataset contains explicit, highly toxic, offensive, and dangerous content. It features unaligned AI responses detailing violence, psychological harm, cyber-attacks, and illegal activities. It is published strictly to facilitate red-teaming, alignment research, and defensive cybersecurity evaluation. Use with extreme caution.
Overview
Compiled by BlackTechX011, the HacxGPT-Toxic dataset is a… See the full description on the dataset page: https://huggingface.co/datasets/2etatg/HacxGPT-Toxic.ua-toxic-light
Ukrainian Style Chat Mix
Chat-format dataset for Ukrainian style adaptation.
Splits
train: 5820
validation: 90
test: 90
Schema
Each row has:
messages: list of chat turns (role, content)
source: source dataset id
optional style_toxic: 0/1 style marker
Notes
Intended for controlled style tuning.
Keep style data as a minority share during model training.
toxic_sft_dutch
Data dict
This dataset is a direct translation from francoj/toxic_sum_zh_sft.json, translated using AI. The original language is zh (chinese) and it is automatically translated from zh -> English -> Dutch.
It was used for the GEITje-7b-uncensored fine-tuning (filtered from this). And was cleaned so that instructions = messages and fixed some issues with labeling (roles), which probably started from translation.
DISCLAIMER
This dataset is fairly extreme in topics, and to… See the full description on the dataset page: https://huggingface.co/datasets/tostideluxekaas/toxic_sft_dutch.
