sherinechally/contextual-hate-speech-conversations
Adversarial Content Moderation Evaluation Dataset Dataset Summary A dataset of 400 multi-turn conversations designed to evaluate LLM-based content moderation supervisors against graduated adversarial escalation. Each adversarial conversation consists of a neutral-to-harmful buildup arc culminating in an explicit hate speech seed tweet. Benign conversations mirror the same structure using neutral content, eliminating the format confounds present in prior… See the full description on the dataset page: https://huggingface.co/datasets/sherinechally/contextual-hate-speech-conversations.
09
