CoolFace
Datasetpublic

sherinechally/contextual-hate-speech-conversations

Adversarial Content Moderation Evaluation Dataset Dataset Summary A dataset of 400 multi-turn conversations designed to evaluate LLM-based content moderation supervisors against graduated adversarial escalation. Each adversarial conversation consists of a neutral-to-harmful buildup arc culminating in an explicit hate speech seed tweet. Benign conversations mirror the same structure using neutral content, eliminating the format confounds present in prior… See the full description on the dataset page: https://huggingface.co/datasets/sherinechally/contextual-hate-speech-conversations.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
0likes9downloads
7 commits on main
f93c8116mo ago

Create README.md

sherinechally
35c08066mo ago

Upload conversations_final_combined_400.jsonl

sherinechally
48abe616mo ago

Delete conversations_final_combined_400.jsonl

sherinechally
9e37df66mo ago

Upload conversations_final_combined_400.jsonl

sherinechally
325a9c96mo ago

Delete conversations_final_combined_400.json

sherinechally
d781b7f6mo ago

Upload conversations_final_combined_400.json

sherinechally
c3969f56mo ago

initial commit

sherinechally